Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11604 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract We study a quantized prefix estimator for inner products that turns a randomly rotated TurboQuant-style representation into a cheap Johnson–Lindenstrauss-like search signal. The idea is simple: rotate the vectors once, keep only a short prefix of coordinates for fast scoring, and quantize the database-side prefix with an unbiased scalar quantizer. We prove that this estimator is unbiased and that its error separates cleanly into two interpretable sources: prefix truncation from using only r coordinates, and quantization error from using b bits per coordinate This separation is useful in systems because the prefix can be exposed as a lightweight filter without building a separate projection index. In ParlayANN graph search, a 64-coordinate truncated view of existing TQ4 codes can replace a separately stored JL256 filter before full-precision reranking, adding only prefix-scale and query lookup-table bookkeeping. In k-means, the same estimator accelerates the dominant point–centroid assignment kernel while preserving exact centroid norms. Empirically, the truncated-TQ filter tracks the JL recall–throughput frontier across five graph-search datasets while reusing the quantized representation already present in the index. View details
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract The promise of tailored agent behavior is undermined by a critical explainability challenge: it is difficult to assess how closely and consistently the agent follows user-defined rules. As Large Language Models (LLMs) transition from static assistants to autonomous agents, developers have pioneered markdown-based rule files (e.g., GEMINI.md, CIDER_AGENT.md) to steer agent behavior and mitigate a "organizational context gap" that emerges when general-purpose models lack the "organizational context" necessary for contextually relevant results. This paper presents a qualitative study of 12 Google software developers (n=12) to investigate the authoring and efficacy of these agent rules. Our findings reveal that while rules are intended as technical steering mechanisms, they function as a "Black Box" of validation, where 12/12 participants rely on anecdotal "vibe checks" due to a profound lack of formal evaluation and explainability frameworks. We identify this opacity as a systemic Attribution Gap, which prevents developers from discerning whether a successful outcome was the result of deliberate logic or "pure luck." Paradoxically, these files serve a dual role as "Living Documentation," bridging technical instruction for AI with sociotechnical onboarding for humans. We argue for a transition toward library-level governance and rigorous traceability to transform agent customization from an ad-hoc craft into a human-centered science by revealing the internal "seams" of rule interpretation. View details
Preview abstract Large language models (LLMs) have shown promise in assisting cybersecurity tasks, yet existing approaches struggle with automatic vulnerability discovery and exploitation due to limited interaction, weak execution grounding, and a lack of experience reuse. We propose Code-RedTeam, a security-aware multi-agent framework designed to mirror real-world red-teaming workflows by integrating security-domain knowledge, code-aware analysis, execution-grounded iterative reasoning, and long-term memory. Code-RedTeam decomposes vulnerability analysis into coordinated discovery and exploitation stages, enabling agents to plan, execute, validate, and refine actions based on real execution feedback while learning from prior trajectories. Extensive evaluations on challenging security benchmarks demonstrate that Code-RedTeam consistently outperforms strong baselines across diverse backbone models, achieving over 60% attack success rate in vulnerability exploitation and up to 10% absolute improvement in vulnerability detection. Ablation and iteration studies further confirm the critical role of execution feedback, structured interaction, and memory for building robust and generalizable cybersecurity agents. View details
Preview abstract This paper demonstrates that artificial intelligence can accelerate mathematical discovery by autonomously solving an open problem in theoretical physics. We present a neuro-symbolic system, combining the Gemini Deep Think large language model with a systematic Tree Search (TS) framework and automated numerical feedback, that successfully derived novel, exact analytical solutions for the power spectrum of gravitational radiation emitted by cosmic strings. Specifically, the agent evaluated the core integral for arbitrary loop geometries, directly improving upon recent AI-assisted attempts that only yielded partial asymptotic solutions. To substantiate our methodological claims regarding AI-accelerated discovery and to ensure transparency, we detail system prompts, search constraints, and intermittent feedback loops that guided the model. The agent identified a suite of 6 different analytical methods, the most elegant of which expands the kernel in Gegenbauer polynomials to naturally absorb the integrand's singularities. The methods lead to an asymptotic result for at large that both agrees with numerical results and also connects to the continuous Feynman parameterization of Quantum Field Theory. We detail both the algorithmic methodology that enabled this discovery and the resulting mathematical derivations. View details
Preview abstract Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and dynamic environments, but this also introduces severe security risks. In particular, indirect prompt injection attacks can compromise agents through malicious instructions hidden in external sources such as web pages, emails, and retrieved documents. Existing defenses are largely reactive, while current automated red-teaming methods mainly optimize attack success rather than systematically uncovering hidden vulnerabilities within the agent pipeline. In this work, we propose PI-Hunter, an automated agentic red-teaming framework that shifts the focus from attack optimization to vulnerability exposure. By combining static attack-surface analysis, source-aware seeding, trajectory evaluation, and feedback-guided exploration, PI-Hunter proactively discovers vulnerable ingestion paths and localizes how malicious instructions propagate through agent reasoning. Extensive experiments across multiple benchmarks, agent architectures, attacks, and defenses show that \method~substantially improves vulnerability exposure and attack-surface coverage compared with existing automated red-teaming baselines, while remaining effective even under strong prompt injection defenses. View details
Preview abstract In "Elephants, Goldfish and the New Golden Age of Software Engineering," the author discusses how AI is changing knowledge work, especially software development. Written from the perspective of April 2026, the article points out that while AI speeds up coding, it can also quickly generate a lot of mistakes and messy code if it isn't carefully managed by human oversight and clear processes. The paper outlines a practical approach to working with AI, broken down into three main sections: Using AI as a Tool, Not a Toy: The author notes that people often get poor results by asking AI to do everything in a single prompt. Instead, users should have back-and-forth conversations with AI to question assumptions, set clear grading rules, and guide the research. The main point is that humans must still provide the final judgment; AI is simply a way to speed up and record that thinking. The Elephant-Goldfish Model: As AI creates more code than humans can easily read, written design documents become more important than the code itself. To keep AI on track, the author suggests a two-part method: * The Elephant: A long chat session where the human and AI discuss ideas and write a detailed design document *before* any code is written. This session holds all of the project's background information and decisions. * The Goldfish: A brand-new AI chat session with no memory. The human asks this "goldfish" to read the design document. If the goldfish cannot understand the plan based only on that document, the document needs more details. * Only after the design document is clear enough for the goldfish to understand does the human ask the AI to write the code based on those strict instructions. * Managing AI and the Future of Work: The author expects that regular employees will soon act like managers, overseeing multiple AI helpers. Because of this, workers need to learn basic management skills, like how to delegate tasks and set clear boundaries. Also, since AI will handle routine chores, humans will need to practice focusing for longer periods to do deeper, harder thinking. Ultimately, a worker's value will come from their planning and decision-making skills, rather than their ability to type code. View details
Preview abstract Superconducting qubits are a leading platform for realizing fault-tolerant quantum computers. Current generations demonstrate fast, high fidelity quantum gates and readout on the order of hundreds of nanoseconds, while maintaining coherence times exceeding one hundred microseconds. Achieving this state-of-the-art performance requires a tight co-design, balancing fundamental physics, microwave engineering, and semiconductor fabrication. Readout designs, in particular, benefit from this multidisciplinary approach. In this talk, we discuss the current challenges for readout in superconducting quantum processors from the perspective of Google Quantum AI. We examine the intersection of device physics and microwave engineering constraints, illustrating how optimizing both is essential for scaling next-generation quantum systems. View details
Identifying Hearing Difficulty Moments in Conversational Audio
Jack Collins
Adrian Buzea
Chris Collier
Alejandro Ballesta Rosen
Julian Maclaren
Kelly Miles
Simon Carlile
Trends in Hearing (2026)
Preview abstract Individuals regularly experience Hearing Difficulty Moments in everyday conversation. Identifying Hearing Difficulty Moments has particular significance in the field of hearing assistive technology where timely interventions are key for real-time hearing assistance. In this article, we propose and compare machine learning solutions for the temporal detection of segments containing Hearing Difficulty Moments in conversational audio. We show that audio language models, through their multimodal reasoning capabilities, can achieve state-of-the-art results for this task, significantly outperforming a simple automatic speech recognition (ASR) hotword heuristic and a more conventional fine-tuning approach with Wav2Vec, an audio-only input architecture that is state-of-the-art for ASR. View details
Preview abstract We introduce ALPS (Activation-based Length Prediction for Scheduling), a method for predicting LLM generation length from prefill activations before any tokens are generated. Unlike existing approaches that require model fine-tuning or complex entropy-weighted pooling, ALPS uses a simple linear probe on the last-token activation at intermediate layers. We discover that generation length is encoded in prefill representations: a ridge regression probe achieves R-squared > 0.85 across three model families. Validation across Llama-3.1-8B, Gemma-2-9B, and Qwen-2.5-7B demonstrates: (1) intermediate layers generally perform well, with some architectural variation; (2) simple last-token extraction outperforms complex pooling strategies; (3) activations improve substantially over surface-feature baselines (24 percentage points over input length plus lexical features). The best models achieve R-squared = 0.943 (Gemma), R-squared = 0.880 (Llama), and R-squared = 0.857 (Qwen) with MAE of 38-80 tokens. All test prompts terminated naturally (100% EOS), eliminating truncation confounds. While our evaluation uses 200 curated prompts—sufficient for demonstrating the phenomenon but requiring broader validation—cross-validation confirms generalization beyond training data. ALPS enables practical applications including budget-constrained inference, request scheduling, and resource allocation. The probe adds negligible overhead (~16KB direction vector, single dot product), making ALPS practical for production deployment. View details
Looking to the brain to improve energy efficiency of AI
Taro Toyoizumi
Hakwan Lau
Michał Klincewicz
Seng Bum Michael Yoo
Megan Peters
taylor.w.webb@gmail.com
Current Biology (2026)
Preview abstract Modern artificial intelligence (AI) systems have achieved remarkable capabilities, but at an extraordinary energy cost. Training and running large-scale models can consume vast resources, posing environmental, economic, and social challenges. In contrast, biological brains perform lifelong learning, adaptive control, and flexible reasoning using orders of magnitude less energy for learning and adaptation over a lifetime. What accounts for this difference -- and how can it guide future AI development? In this article, we identify key biological principles that support energy-efficient capacities in biological brains, and consider how they might inform the design of more sustainable artificial systems. We organize our analysis around three domains: architectural constraints, signaling strategies, and learning algorithms. In each domain, we discuss concrete observations from biology -- from cell to circuit to cognitive level -- and describe how current and emerging AI systems mirror or diverge from these motifs. One striking feature of biological energy optimization is often overlooked: that brains are remarkably stable in their energy usage across heterogeneous modes, suggesting they may minimize energy needs during active environmental processing through maximizing the utility of “rest-like” background processes. Overall, rather than advocating for biomimicry for its own sake, we argue for biologically informed engineering. Understanding how natural systems minimize energetic cost while maximizing flexibility may help us build AI that is not only powerful, but also efficient, equitable, and environmentally responsible. View details
AIRS: Scaling Live Inference in Resource Constrained Environments
Xiaohao Yang
Tuan Do
Chelsea Chen
Harshvardhan GM
(2026)
Preview abstract Advancements in large language models (LLMs) have made them increasingly useful for complex reasoning tasks which previously required domain experts. One such task is quality evaluation of query responses produced by a search engine. Evaluation generates metrics necessary to study the quality, impact, and usefulness of product changes and features. Typically, to compute evaluation metrics, human experts are asked to rate various attributes of search responses. This process is generally quite expensive and requires several days to complete. As an alternative, LLMs are now being used to perform rating tasks with lower costs and latency. In addition, many new metrics are being developed to evaluate Google's new AI-based offerings, which require ratings too. As a result, there is much higher demand for LLM rating prediction tasks in comparison with the allocated TPU (Tensor Processing Unit) budget. A larger portion of the company's TPU resources are reserved for serving live user traffic. In this paper, we present the AI Rater Service (AIRS), an inference pipeline that employs several software engineering techniques to generate AI ratings with high reliability and low latency. AIRS maximizes LLM inference throughput by optimizing TPU resource utilization across various evaluation workflows, while minimizing latency for higher priority tasks. View details
Preview abstract Every team building Large Language Models (LLMs) faces a core challenge: offline benchmarks show performance gains and user engagement rises post deployment, but isolating cause from effect remains difficult. Simultaneous marketing, media coverage, and seasonal demand obscure whether model updates truly drive engagement gains. This paper presents a novel causal estimation approach that leverages non uniform quality improvements across capabilities within a single model version. Because capabilities improve unevenly (e.g., strong gains in coding versus modest gains in writing), users experience varied quality depending on their task distribution. This variation in experienced quality provides causal signal for estimation. We validate this approach on synthetic data with known ground truth. The raw estimator recovers 81% to 88% of the true effect, with the remainder lost to measurement error in usage estimates. Applying an errors in variables disattenuation adjustment (using a test retest reliability ratio mean correlation of 0.861) corrects the estimate to 1.017 (bootstrapped 95% CI: 0.915 to 1.109). By contrast, naive methods fail significantly, recovering only 63% without version controls and 62% using real time rather than frozen usage patterns. Permutation tests confirm the framework distinguishes true causal effects from noise. Sensitivity analysis indicates recovery improves monotonically from 78% to 87% as pre period length increases from 3 to 7 weeks, highlighting a clear tradeoff between sample duration and estimator precision. Seed sensitivity tests further confirm stability across random draws. View details
Preview abstract Social scientists rely on hypothesis testing to support their research conclusions, but our standard procedures are designed for testing one hypothesis rather than adjudicating between rival possibilities. We develop a new framework, “classification testing”, as an alternative. Instead of selecting one hypothesis to test, a researcher conducting a classification test decides what qualitative distinctions (“classes”) are most substantively relevant; the test either assigns the estimand to a class with error control similar to that of a conventional hypothesis test, or declares the result inconclusive. We argue that classification testing is superior to current practice not just when the objective is to adjudicate between rival possibilities but also when there is one research hypothesis to be tested, because classification testing exposes that hypothesis to refutation. We illustrate the framework by applying it to a well-known media experiment and offer an R package to aid in implementation. View details
×