Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11612 publications
Preview abstract We study a quantized prefix estimator for inner products that turns a randomly rotated TurboQuant-style representation into a cheap Johnson–Lindenstrauss-like search signal. The idea is simple: rotate the vectors once, keep only a short prefix of coordinates for fast scoring, and quantize the database-side prefix with an unbiased scalar quantizer. We prove that this estimator is unbiased and that its error separates cleanly into two interpretable sources: prefix truncation from using only r coordinates, and quantization error from using b bits per coordinate This separation is useful in systems because the prefix can be exposed as a lightweight filter without building a separate projection index. In ParlayANN graph search, a 64-coordinate truncated view of existing TQ4 codes can replace a separately stored JL256 filter before full-precision reranking, adding only prefix-scale and query lookup-table bookkeeping. In k-means, the same estimator accelerates the dominant point–centroid assignment kernel while preserving exact centroid norms. Empirically, the truncated-TQ filter tracks the JL recall–throughput frontier across five graph-search datasets while reusing the quantized representation already present in the index. View details
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract Using generative artificial intelligence with sensitive data may present challenges, as transmitting personally identifiable information or protected health information to third-party providers can introduce security risks, and some data masking techniques can reduce reasoning capabilities. A described system uses a proxy, masking layer that can intercept data within an enterprise's secure perimeter. This layer can substitute sensitive strings with persistent, structured semantic tokens that may be enriched with non-sensitive metadata hints to help preserve context. An external artificial intelligence can perform reasoning on this abstracted data, and its tokenized response can be re-hydrated into readable text on a client device (e.g., a smartphone, computer, or wearable device). This approach may allow third-party models to reason on proprietary information without direct access to the underlying plaintext data, which can assist organizations in managing data sovereignty while maintaining functional utility. View details
Preview abstract Despite advances in high performance computing, accurate numerical simulations of global atmospheric dynamics remain a challenge. The resolution required to fully resolve the vast range scales as well as the strong coupling with—often not fully-understood—physics renders such simulations computationally infeasible over time horizons relevant for long-term climate risk assessment. While data-driven parameterizations have shown some promise of alleviating these obstacles, the scarcity of high-quality training data and their lack of long-term stability typically hinders their ability to capture the risk of rare extreme events. In this work we present a general strategy for training variational (probabilistic) neural network models to non-intrusively correct under-resolved long-time simulations of turbulent climate systems. The approach is based on the paradigm introduced by Barthel Sorensen et al. (2024, https://doi.org/10.1029/2023ms004122) which involves training a post-processing correction operator on under-resolved simulations nudged toward a high-fidelity reference. Our variational framework enables us to learn the dynamics of the underlying system from very little training data and thus drastically improve the extrapolation capabilities of the previous deterministic state-of-the art—even when the statistics of that training data are far from converged. We investigate and compare three recently introduced variational network architectures and illustrate the benefits of our approach on an anisotropic quasi-geostrophic flow. For this prototype model our approach is able to not only accurately capture global statistics, but also the anistropic regional variation and the statistics of multiple extreme event metrics—demonstrating significant improvement over previously introduced deterministic architectures. View details
Preview abstract As artificial intelligence (AI) is rapidly integrated into healthcare, ensuring that this innovation helps to combat health inequities requires engaging marginalized communities in health AI futuring. However, little research has examined Black populations’ perspectives on the use of AI in health contexts, despite the widespread health inequities they experience–inequities that are already perpetuated by AI. Addressing this research gap, through qualitative workshops with 18 Black adults, we characterize participants’ cautious optimism for health AI addressing structural well-being barriers (e.g., by providing second opinions that introduce fairness into an unjust healthcare system), and their concerns that AI will worsen health inequities (e.g., through health AI biases they deemed inevitable and the problematic reality of having to trust healthcare providers to use AI equitably). We advance health AI research by articulating previously-unreported health AI perspectives from a population experiencing significant health inequities, and presenting key considerations for future work. View details
Preview abstract The rapid expansion of the Internet of Things (IoT) and smart home ecosystems has led to a fragmented landscape of user data management across consumer electronics (CE) such as Smart TVs, gaming consoles, and set-top boxes. Current onboarding processes on these devices are characterized by high friction due to manual data entry and opaque data-sharing practices. This paper introduces the User Data Sharing System (UDSS), a platform-agnostic framework designed to facilitate secure, privacy-first PII (Personally Identifiable Information) exchange between device platforms and third-party applications. Our system implements a Contextual Scope Enforcement (CSE) mechanism that programmatically restricts data exposure based on user intent—specifically distinguishing between Sign-In and Sign-Up workflows. Unlike cloud-anchored identity standards such as FIDO2/WebAuthn, UDSS is designed for shared, device-centric CE environments where persistent user-to-device bind-ing cannot be assumed. We further propose a tiered access model that balances developer needs with regulatory compliance (GDPR/CCPA). A proof-of-concept implementation on a reference ARMv8 Linux-based middleware demonstrates that UDSS reduces user onboarding latency by 65% and measurably reduces PII over-exposure risk through protocol-enforced data minimization. This framework provides a standardized approach to identity management in the heterogeneous CE market. View details
Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study
Peter Brodeur
Jacob M. Koshy
Khaled Saab
Ava Homiar
Roma Ruparel
Charles Wu
Ryutaro Tanno
Joseph Xu
Amy Wang
David Stutz
Hannah M. Ferrera
David Barrett
Lindsey Crowley
Jihyeon Lee
Spencer E. Rittner
Selena K. Zhang
Elahe Vedadi
Christine G. Kohn
Kavita Kulkarni
Vinay Kadiyala
Sara Mahdavi
Wendy Du
David Feinbloom
Renee Wong
Petar Sirkovic
Alessio Orlandi
Juro Gottweis
Joelle Barral
Kat Chou
James Manyika
Rob Fields
Jonathan X. Li
Marc L. Cohen
Adam Rodman
The Lancet (2026)
Preview abstract Background: Artificial intelligence (AI)-based systems show promise for assisting primary care providers (PCPs) with patient care. We aimed to evaluate the safety and quality of clinical conversations of a patient-facing conversational AI system, which engaged in real-world urgent primary care appointments. Methods: In this prospective, single-centre, single-arm feasibility study, English-speaking patients aged at least 18 years interacted with the Articulate Medical Intelligence Explorer (AMIE) up to 5 days before a single-complaint urgent primary care appointment. Physician safety supervisors monitored all interactions and were trained to intervene on the basis of predefined safety criteria. AMIE transcripts and summaries were shared with PCPs before the visit. Primary outcomes were the number of supervised conversation safety stops, AMIE’s conversation quality assessed by clinical evaluators, and patient and PCP experiences per surveys. This study is registered with ClinicalTrials.gov (NCT06911398). Findings: From April to November, 2025, 114 patients were enrolled with 98 completing both the AMIE interaction and the PCP appointment. Zero conversation safety stops were required on the basis of prespecified criteria. Safety supervisors noted one hallucination and added clinical information in five interactions. AMIE’s conversations were rated favourably in 87–100% of cases (17 criteria) by clinical evaluators, and 48–96% (16 criteria) by patients. Patient attitudes towards AI improved after interacting with AMIE and remained elevated after the patient’s visit with their physician. PCPs completed post-surveys in 60 of 98 cases, including 44 cases in which they reviewed the AMIE transcript before the visit. PCPs found AMIE helpful for visit preparation in 33 of 44 cases and reported that it might have changed their behaviour in 25 of 44 cases. Interpretation: Although further research is needed, this study shows the initial feasibility of conversational AI in a real-world setting—assessed via conversation safety and quality, as well as user acceptance—and represents a crucial step towards clinical translation. Funding: Alphabet. View details
The Perfection Paradox: From Architect to Curator in AI-Assisted API Design
JJ Geewax
David R Karger
Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26), ACM, Barcelona, Spain, TBD
Preview abstract Enterprise API design is often bottlenecked by the tension between rapid feature delivery and the rigorous maintenance of usability standards. We present an industrial case study evaluating an AI-assisted design workflow trained on API Improvement Proposals(AIPs). Through a controlled study with 16 industry experts, we compared AI-generated API specifications against human-authored ones. While quantitative results indicated AI superiority in 10 of 11 usability dimensions and an 87% reduction in authoring time, qualitative analysis revealed a paradox: experts frequently misidentified AI work as human (19% accuracy) yet described the designs as unsettlingly “perfect.” We characterize this as a “Perfection Paradox”—where hyper-consistency signals a lack of pragmatic human judgment. We discuss the implications of this perfection paradox, proposing a shift in the human designer’s role from the “drafter” of specifications to the “curator” of AI-generated patterns. View details
CoDaS: AI Co-Data-Scientist for Biomarker Discovery via Wearable Sensors
Juro Gottweis
CJ Park
Salman Rahman
Ahmed Metwally
Hong Yu
Ivor Rendulic
Yuzhe Yang
Petar Sirkovic
Daniel McDuff
Shwetak Patel
Nicolas Stroppa
Yubin Kim
Mark Malhotra
Orson Xu
Sam Schmidgall
Tim Althoff
Elahe Vedadi
Cynthia Breazeal
Hae Won Park
(2026)
Assessing Global Flood Risk from Atmospheric Rivers through Physically Guided Machine Learning
Assaf Shmuel
Oleg Zlydenko
Martin Gauch
Colin Price
Scientific Reports (2026)
Preview abstract Atmospheric rivers (ARs) are narrow corridors of concentrated moisture transport that play a crucial role in the global water cycle, delivering both beneficial rainfall and severe floods. Here, we develop a physically guided, explainable machine learning framework to predict flood occurrence during AR conditions worldwide by integrating AR characteristics with meteorological and topographic variables. The best model achieves a receiver operating characteristic area under the curve (ROC AUC) of 0.94, outperforming a logistic regression baseline at 0.81. Despite their limited footprint, we find that one third of large midlatitude floods occur under AR conditions, reflecting their disproportionate role in global flood risk. SHAP analysis highlights integrated vapor transport, precipitation, and elevation as dominant predictors. We show the non-linear amplification of flood risk under combined conditions of high soil moisture and persistent AR activity, underscoring the importance of antecedent wetness in modulating flood risk. Using consistent reanalysis inputs, we find that model-estimated high-flood-risk AR conditions increased globally by over 10% from 1980 to 2020 and shifted poleward. Validation on an independent satellite-based flood database shows comparable skill. We estimate that roughly 90% of the global population lives in regions that experience at least one AR annually, underscoring the broad societal relevance of AR dynamics. These findings highlight the value of physically guided machine learning for mapping and monitoring AR-related flood risk globally, offering actionable insights for preparedness and climate adaptation. View details
Preview abstract Online financial scams represent a long-standing and serious threat for which people seek help. We present a study to understand people’s in situ motivations for engaging with scams and the help needs they express before, during, and after encountering a scam. We identify the main emotions scammers exploited (e.g., fear, hope) and characterize how they did so. We examine factors—such as financial insecurity and legal precarity—which elevate people’s risk of engaging with specific scams and experiencing harm. We indicate when people sought help and describe their help-seeking needs and emotions at different stages of the scam. We discuss how these needs could be met through the design of contextually-specific prevention, diagnostic, mitigation, and recovery interventions. View details
From Correctness to Collaboration: A Human-Centered Taxonomy of AI Agent Behavior in Software Engineering
Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA ’26), ACM, New York, NY, USA (2026)
Preview abstract The ongoing transition of Large Language Models in software engineering from code generators into autonomous agents requires a shift in how we define and measure success. While models are becoming more capable, the industry lacks a clear understanding of the behavioral norms that make an agent effective in collaborative software development in the enterprise. This work addresses this gap by presenting a taxonomy of desirable agent behaviors, synthesized from 91 sets of user-defined rules for coding agents. We identify four core expectations: Adhere to Standards and Processes, Ensure Code Quality and Reliability, Solve Problems Effectively, and Collaborate with the User. These findings offer a concrete vocabulary for agent behavior, enabling researchers to move beyond correctness-only benchmarks and design evaluations that reflect the realities of professional software development in large enterprises. View details
Compact Conformal Subgraphs
Kamesh Munagala
Aravindan Vijayaraghavan
ICML (2026)
Preview abstract Conformal prediction provides rigorous uncertainty guarantees for model outputs but can produce prohibitively large prediction sets in structured domains such as routing, planning, or sequential recommendation. We introduce \emph{graph-based conformal compression}, a framework for constructing compact subgraphs that preserve the statistical validity of conformal prediction while reducing structural complexity. We study a formulation that selects a smallest subgraph capturing a prescribed fraction of conditional probability mass, and reduce to a weighted version of densest $k$-subgraphs in hypergraphs, in the regime where the subgraph has a large fraction of edges. We design efficient approximation algorithms that achieve constant factor coverage and size trade-offs. Our results highlight an algorithmic regime, distinct from classical densest-$k$-subgraph hardness settings, where the problem can be approximated efficiently, bridging conformal prediction with combinatorial graph compression. We finally validate our algorithmic approach on synthetic and real-world instances of trip planning and navigation, showing in each case that our approach handily beats natural baselines. View details
Efficacy of Scalable Airline-led Contrail Avoidance
Thomas Dean
Tristan Abbott
Jill Blickstein
Alejandra Martín Frías
Mark Galyen
Rebecca Grenham
Paul Hodgson
Alan Pechman
Tyler Robarge
Dinesh Sanekommu
Aarón Sonabend
Marc E.J. Stettler
Raimund Zopp
Journal of Environmentally Compatible Air Transport System (JECATS) (2026) (to appear)
Preview abstract Contrails account for a large portion of aviation's contribution to anthropogenic climate change. Navigational contrail avoidance is a promising solution to mitigate the warming caused by contrails. Prior trials testing navigational contrail avoidance have relied on bespoke integrations of contrail forecasts into airline operations. Here, we use a randomized control trial to test the feasibility of dispatcher-led contrail avoidance integrated into standard flight planning operations using a workflow which scales to an airline's entire network. Using satellite imagery and an automated flight-contrail attribution algorithm, we observed an 11.6% reduction in contrail formation rate for the 1232 flights marked as eligible for contrail avoidance (intent-to-treat) relative to the flights in the control group (p = 0.0109). In the 112 flights which flew contrail avoidance as planned (per-protocol flights), we observed a 62.0% lower contrail formation rate relative to the flights in the control group (p < 0.001). No statistically significant difference in fuel usage was observed between the two groups. View details
×