Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11583 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Towards A Human-in-the-Loop Framework for Reliable Patch Evaluation using an LLM-as-a-Judge
Renyao Wei
José Cambronero
AI-SQE '26: Proceedings of the 1st International Workshop on AI for Software Quality Evaluation - Judgment, Metrics, Benchmarks, and Beyond, ACM (Association for Computing Machinery), New York, NY, USA (2026), pp. 19 - 28
Preview abstract Reliable evaluation is crucial for advancing Automated Program Repair (APR), but prevailing benchmarks that rely on execution-based evaluation methods (pass@k) often fail to capture the patch quality required for real-world adoption. This creates a significant gap between automated metrics and true patch validity (valid@k), a discrepancy observed across several state-of-the-art techniques. To develop a scalable solution for measuring valid@k, we first study the human evaluation process itself. While manual assessment can determine validity, we find it suffers from poor inter-rater reliability (Fleiss' Kappa k=0.307). Our foundational insight is that this inconsistency is largely resolved when evaluators use a shared, high-quality rubric, which significantly improves agreement. Building on this finding, we propose an LLM-as-a-Judge framework that operationalizes rubric-guided evaluation at scale. Our method employs a human-in-the-loop workflow where an LLM first generates a candidate rubric for a given bug, which a human expert then reviews and refines into a "golden" evaluation standard. This golden rubric is then used by an LLM judge to assess the validity of candidate patches. In an evaluation on 48 bugs and 115 patches, our LLM judge demonstrates substantial agreement with the consensus of human developers. This work contributes a scalable and reliable methodology for approximating valid@k, providing a much-needed high-fidelity signal for measuring true progress in the field of automated program repair. View details
Toward a test of medical AI superintelligence
Ethan Goh
David Wu
Chase Walton
Liam McCoy
Anastasia Perez
Laura Wegner
Fateme Nateghi Haredasht
Luyang Luo
Kathleen Lacar
Thomas Buckley
Austin Schoeffler
Peter Brodeur
Kameron C. Black
John Havlik
John Rumsfeld
Daniel Lopez-martinez
Paxton Maeder-York
Karan Singhal
David Gunning
Bon Ku
Haider Warraich
Shantanu Nundy
Vishnu Ravi
Arnold Milstein
Jason Hom
Kevin Schulman
Pranav Rajpurkar
Arjun Manrai
Robert Wachter, MD
Eric Topol
Eric horvitz
Adam Rodman
Jonathan Chen
Nature Medicine (2026)
Preview abstract Researchers urgently need a rigorous, task-based framework to define and measure medical AI ‘superintelligence’, because existing benchmarks are misleading and insufficient. View details
Progressive Photorealistic Simplification
Adi Rosenthal
Yedid Hoshen
Arik Shamir
2026
Preview abstract Existing image simplification techniques often rely on Non-Photorealistic Rendering (NPR), transforming photographs into stylized sketches, cartoons, or paintings. While effective at reducing visual complexity, such approaches typically sacrifice photographic realism. In this work, we explore a complementary direction: simplifying images while preserving their photorealistic appearance. We introduce progressive semantic image simplification, a framework that iteratively reduces scene complexity by removing and inpainting elements in a controlled manner. At each step, the resulting image remains a plausible natural photograph. Our method combines semantic understanding with generative editing, leveraging Vision-Language Models (VLMs) to identify and prioritize elements for removal, and a learned verifier to ensure photorealism and coherence throughout the process. This is implemented via an iterative \emph{Select–Remove–Verify} pipeline that produces high-quality simplification trajectories. To improve efficiency, we further distill this process into an image-to-video generation model that directly predicts coherent simplification sequences from a single input image. Beyond generating cleaner and more focused compositions, our approach enables applications such as content-aware decluttering, semantic layer decomposition, and interactive editing. More broadly, our work suggests that simplification through structured content removal can serve as a practical mechanism for guiding visual interpretation within the photorealistic domain, complementing traditional abstraction methods. View details
Preview abstract Systems for escalating interactions from automated agents to human agents can create inefficiencies, for example, by transferring unstructured transcripts. An intermediary system can employ a generative artificial intelligence synthesis engine to process the context of an automated interaction upon an escalation trigger. The engine may analyze the dialogue transcript, user metadata, and the automated agent's internal state to perform semantic abstraction, diagnose potential failure points, and infer a possible resolution. The system can then generate a structured briefing for the human agent, which could include a concise summary, a failure diagnosis, or a recommended next action presented as an interactive element. This process may facilitate a more efficient handoff and contribute to an improved escalation workflow by providing the human agent with synthesized, contextual information. View details
Nudging Developers Toward Privacy: Evaluating the Impact of Personalized App Review Reports
Omer Akgul
Michelle L. Mazurek
USENIX Symposium on Usable Privacy and Security (SOUPS) (2026)
Preview abstract Mobile application developers often struggle to create accurate privacy notices or implement robust privacy practices due to limited expertise or resources. While users share unsolicited privacy feedback in app reviews, and prior research has characterized this privacy feedback, uncovering developer reactions to this feedback remains unexplored. This study explores whether personalized privacy review reports---summarizing real user feedback for a developer's own app---can effectively nudge them toward planning privacy improvements. We surveyed 42 app developers, presenting them with reports containing privacy themes, temporal trends, peer benchmarks, and emotion distributions derived from their apps' reviews. Our findings indicate that these privacy report interventions proved highly effective, with 76% (32 of 42) of participants finding at least one section of the report useful. Furthermore, exposure to the report increased the participants' intent to pursue privacy-relevant actions -- such as reorganizing the UI, enhancing privacy communications, or adding/removing features -- with 69% (29 of 42) of participants indicating an increased intent to do so. Almost all developers expressed a desire to receive such privacy reports periodically or on demand. These results indicate that making this style of report broadly available across the industry could foster a more privacy-conscious mobile ecosystem. View details
Preview abstract Global shared service centers are critical to modern enterprise operations but struggle to provide consistent, timely support across linguistic boundaries. This paper introduces the Glossary-Grounded Universal Queue (GGUQ), a socio-technical framework designed to bridge the gap between the operational goal of a unified global service queue and the reality of a multilingual workforce. The GGUQ is a real-time, workflow-embedded communication architecture that leverages Large Language Models (LLMs) to provide high-fidelity, two-way translation directly within an agent's enterprise platform. The framework's key innovation is a "glossary-grounded" approach, where translation prompts are programmatically injected with a curated repository of enterprise-specific terminology. This ensures a level of contextual and terminological integrity unachievable by generic machine translation tools. By detailing the GGUQ's three-pillar architecture—Dynamic Translation, Glossary-Grounded Integrity, and Resilient Operations—we propose a new model for computer-mediated communication in global enterprises. This framework aims to move beyond federated, language-siloed support models to enable a true "follow-the-sun" operational capability, promoting both organizational efficiency and a more inclusive employee experience. View details
Preview abstract This paper provides an overview of Google's TPUs across five generations, from TPU v2 to Ironwood, highlighting their evolution as scalable, resilient, and sustainable supercomputers for AI training. It details the TPU’s stable architecture and microarchitecture, which has surprisingly easily accommodated the rapidly changing deep neural network workloads, such as the rise of Transformers. Key advancements over eight years include 10x increase in HBM capacity and bandwidth per node, a 100x increase in peak node performance, and a 3600x increase in supercomputer performance. The paper also discusses the role of optical circuit switches and built-in self test in enhancing resilience, how TPU’s carbon footprint was reduced by improving embodied carbon emissions per floating point operation and a 30x gain in performance per Watt. It concludes by identifying six features that may well characterize the successful AI accelerators of this decade. View details
GroupDPO: Memory-Efficient Group-Wise Direct Preference Optimization
Jixuan Leng
Hsiang-Fu Yu
Vinod Raman
Inderjit Dhillon
The 2026 Conference on Empirical Methods in Natural Language Processing
Preview abstract Preference optimization is widely used to align Large Language Models (LLMs) with preference feedback. However, most existing methods train on a single positive-negative pair per prompt, discarding additional supervision available in preference datasets that typically contain multiple candidate responses. Motivated by this limitation, recent work explores group-wise preference optimization, which jointly contrasts multiple responses for the same prompt, but its empirical behavior and scalability remain underexplored due to the memory overhead of group-coupled objectives. In this work, we present a unified empirical and systems study of group-wise preference optimization and develop a memory-efficient implementation for group-coupled objectives. By instantiating first-order linearization with objective-specific per-response coefficients, our implementation preserves first-order gradients while decoupling samples during backpropagation, substantially reducing peak memory usage and enabling scalable training with larger groups. Across offline and online settings, we show that leveraging multiple responses consistently outperforms single-pair training. Furthermore, incorporating a negative log-likelihood (NLL) term on positive responses is critical for both performance gains and training stability. View details
Identifying Hearing Difficulty Moments in Conversational Audio
Jack Collins
Adrian Buzea
Chris Collier
Alejandro Ballesta Rosen
Julian Maclaren
Kelly Miles
Simon Carlile
Trends in Hearing (2026)
Preview abstract Individuals regularly experience Hearing Difficulty Moments in everyday conversation. Identifying Hearing Difficulty Moments has particular significance in the field of hearing assistive technology where timely interventions are key for real-time hearing assistance. In this article, we propose and compare machine learning solutions for the temporal detection of segments containing Hearing Difficulty Moments in conversational audio. We show that audio language models, through their multimodal reasoning capabilities, can achieve state-of-the-art results for this task, significantly outperforming a simple automatic speech recognition (ASR) hotword heuristic and a more conventional fine-tuning approach with Wav2Vec, an audio-only input architecture that is state-of-the-art for ASR. View details
Leveraging LLMs to understand public perception of earthquake early warnings: A case study of the 2025 M6.2 Türkiye Earthquake
Hanjing Wang
Patrick Robertson
Richard Allen
Alexei Barski
Robert Bosch
Nivetha Thiruverahan
Youngmin Cho
Tajinder Gadh
Steve Malkos
Boone Spooner
Greg Wimpey
Marc Stogaitis
Scientific Reports, 16 (2026), pp. 22226
Preview abstract Android's Earthquake Alert (AEA) system provided life-saving early warnings to millions during the M6.2 Marmara Ereğlisi, Türkiye earthquake on April 23, 2025. This event, the largest in the region in 25 years, served as a critical real-world test for smartphone-based Earthquake Early Warning (EEW) systems. The AEA system successfully delivered alerts to users with high precision, offering over a minute of warning before the strongest shaking reached urban areas. This study leveraged Large Language Models (LLMs) to analyze 560 public social media posts from the X platform, extracting 42 distinct attributes related to user experience and behavior. Statistical analyses revealed significant relationships, notably a strong correlation between user trust and alert timeliness. Key findings indicate that alerts received before shaking, clear information, and audible notifications significantly increased the likelihood of active protective responses and enhanced perceived usefulness and future trust in the system. Conversely, annoyance and perceived inaccuracy led to inaction and diminished trust. The study highlights the profound impact of smartphone-based EEW in mitigating seismic risk, providing actionable insights for optimizing alert design, public education campaigns, and future behavioral research to improve the effectiveness of such systems in seismically active regions. View details
Preview abstract Superconducting qubits are a leading platform for realizing fault-tolerant quantum computers. Current generations demonstrate fast, high fidelity quantum gates and readout on the order of hundreds of nanoseconds, while maintaining coherence times exceeding one hundred microseconds. Achieving this state-of-the-art performance requires a tight co-design, balancing fundamental physics, microwave engineering, and semiconductor fabrication. Readout designs, in particular, benefit from this multidisciplinary approach. In this talk, we discuss the current challenges for readout in superconducting quantum processors from the perspective of Google Quantum AI. We examine the intersection of device physics and microwave engineering constraints, illustrating how optimizing both is essential for scaling next-generation quantum systems. View details
Preview abstract Human-Computer Interaction research and design pedagogy rely on idealized process models, such as the Double Diamond, to describe how user experiences are designed. These models assume an orderly, linear design process that, while easy to understand, fails to capture the iterative and collaborative reality of professional practice. A few qualitative studies have successfully captured this complexity -- still, they often suffer from retrospective narrative smoothing and lack systemic scale. To understand how design unfolds in real products, we analyzed historical snapshots of 102 Figma files from a multi-national technology company and investigated the true trajectories of the design process at scale. Our analysis reveals that while the established process models might be applicable, the operational details are highly non-linear. Rather than a straight line from ideation toward completion, design advances are repeatedly reset to the ideation stage as feedback is received. We argue that by treating design files as operational telemetry, the industry can move beyond abstract frameworks to build practices and collaborative tools that support the non-linear realities of modern product development. View details
Preview abstract As the ECMAScript specification evolves, industrial-scale JavaScript compilers face the challenge of supporting modern language syntax while maintaining compatibility for diverse execution environments. Traditionally, compilers solve this by running transpilation passes in a monolithic pipeline, where the transpilation passes are chosen to execute strictly based on a target language level. This results in significant computational waste, as compilers perform expensive Abstract Syntax Tree (AST) traversals to lower features that may not exist in the actual input source code. We present a static analysis improvement that conditionally executes transpiler passes based on accurately tracking and dynamically maintaining the exact set of language features seen in the compilation unit throughout the transpilation process. It is implemented in the production Google Closure Compiler. By populating and maintaining a FeatureSet at every JavaScript script-level, it dynamically skips running the unnecessary lowering passes. We detail the architectural safeguards - including strategic pass ordering and dynamic validation of the transpiled code for feature-correctness. Evaluation of this improvement on large-scale production applications produced a considerable reduction in compilation time and saved compute and memory usage. View details
×