Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11587 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract This piece analyzes how hardware product managers can navigate the financial pressure of AI-related costs and trade tariffs. It presents a practical framework for deploying the Joint Development Model (JDM) as a product development strategy to balance innovation with cost efficiency, enabling sustainable growth for next-generation AI-native consumer devices. View details
Preview abstract In this paper we introduce \emph{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. We establish a high probability regret bound of $\widetilde{O}\left(H^4\sqrt{T|S||A|\log(|\mathcal{F}||\mathcal{P}|)}\right),$ where $S$ and $A$ denote the state and action spaces, $H$ the horizon length, $T$ the number of episodes, and $\mathcal{F}, \mathcal{P}$ the function classes used to approximate the losses and dynamics, respectively. This result improves the dependence on $|S|$ and $|A|$ compared to the previously known state-of-the-art bound of \citet*{DBLP:conf/nips/QianHS24}. Our analysis introduces sophisticated confidence bounds for stochastic policies that eliminate restrictive assumptions required in prior work. Our results demonstrate that simple policy optimization over optimistic model approximations can achieve better, near-optimal regret bound for CMDPs. View details
Preview abstract In some multi-stage software build pipelines, downstream compiler errors may be reported against ephemeral, machine-generated intermediate artifacts rather than original, human-written source code, which can make remediation challenging. A system and method may address this by intercepting a downstream error, mapping its location back to the original source file, and programmatically injecting a dormant suppression tag into the original source code. During a subsequent build, an intermediate transpiler can propagate this tag into a newly generated intermediate artifact. In the intermediate file, the tag may become active and be recognized by the downstream compiler as a directive to suppress the specific error. This approach can facilitate an automated remediation process for certain build failures that avoids direct modification of ephemeral files and uses the original source code as a record for suppression. View details
Preview abstract Scaling test-time computation improves performance across different tasks on large language models (LLMs), yet mainstream scaling methods remain challenging for tool-augmented LLM agents. Sequential scaling tends to yield shallow tool use and under-exploration, whereas parallel scaling inflates cost through repeated tool calls. The dual costs of tokens and tool calls further complicate cost accounting and hinder fair comparison. In this work, we study the test-time scaling of widely used, tool-reliant search agents under resource constraints, analyzing performance with a unified cost metric that incorporates both tokens and tool calls. To this end, we propose Cost-effective Agent Test-Time Scaling (CATS), a budget-aware framework designed to support more cost-effective scaling by guiding resource allocation between sequential and parallel exploration. Experiments across search-intensive benchmarks show that CATS produces more favorable scaling curves, attaining higher accuracy with fewer tool calls and lower overall cost. Our work introduces a cost-conscious design for agent test-time scaling and contributes empirical insights that enable a more transparent and principled understanding of scaling in tool-augmented agents. View details
Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting
David Autor
Tanya Rodchenko
Joshua Martin
Zanna Iscenko
Scott Strand
David Pearl
Melissa Ferere
NBER (2026)
Preview abstract Whether AI assistance builds or erodes professional expertise is unsettled. In a pre-registered three-month randomized controlled trial, we gave 133 practicing patent lawyers at eleven U.S. intellectual property law firms access to a custom AI drafting assistant and measured both their performance while using AI and their professional judgment afterward without it. All work was scored by blinded expert patent attorneys. Paralleling findings from other white-collar domains, AI access raised the quality of work delivered on benchmark patent drafting tasks at 10 days (0.34 SD, p = 0.03) and 90 days (0.38 SD, p = 0.01), with larger gains among junior lawyers. After three months, all subjects redlined an existing patent application without AI, a core task of patent practice requiring expert judgment. Treated lawyers outperformed controls by 0.32 SD (p = 0.04), but this advantage was concentrated entirely among senior lawyers (0.45 SD, p = 0.02). Junior lawyers showed no average gain; their scores instead bifurcated, with sharply fewer mediocre scores offset by more poor and more good ones. The largest gains from AI thus accrued to the lawyers who retained the least. Foundational expertise may be a prerequisite for extracting durable skill from AI-assisted practice. View details
Visual Planning: Let’s Think Only with Images
Han Zhou
Caiqi Zhang
Anna Korhonen
Chengzu Li
Yi Xu
Ivan Vulic
International Conference on Learning Representations (ICLR) (2026)
Preview abstract Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have significantly enhanced machine reasoning across diverse tasks. However, these models predominantly rely on language as the medium for both expressing and structuring reasoning, even when visual information is present. In this work, we argue that language may not always be the most natural or effective modality for reasoning, particularly in tasks involving spatial, geometric, or physical dynamics. Motivated by this, we propose a new paradigm, Visual Planning, which enables planning through purely visual representations, independent of textual mediation. In this paradigm, planning is executed via sequences of images that encode step-by-step inference in the visual domain, akin to how humans sketch or visualize future actions. We then introduce a novel two-stage reinforcement learning framework empowered by GRPO for post-training large vision models, resulting in substantial improvements in planning accuracy and generalization across both seen and novel scenarios, validated in representative visual navigation tasks, FrozenLake and Maze. Our results establish Visual Planning as a viable and promising alternative to language-based reasoning, opening new avenues for tasks that benefit from intuitive, image-based inference. View details
Preview abstract We study online linear optimization with matrix variables constrained by the operator norm, a setting where the geometry renders designing data-dependent and efficient adaptive algorithms challenging. The best-known adaptive regret bounds are achieved by Shampoo-like methods, but they require solving a costly quadratic projection subproblem. To address this, we extend the gradient-based prediction scheme to adaptive matrix online learning and cast algorithm design as constructing a family of smoothed potentials for the nuclear norm. We define a notion of admissibility for such smoothings and prove any admissible smoothing yields a regret bound matching the best-known guarantees of one-sided Shampoo. We instantiate this framework with two efficient methods that avoid quadratic projections. The first is an adaptive Follow-the-Perturbed-Leader (FTPL) method using Gaussian stochastic smoothing. The second is Follow-the-Augmented-Matrix-Leader (FAML), which uses a deterministic hyperbolic smoothing in an augmented matrix space. By analyzing the admissibility of these smoothings, we show both methods admit closed-form updates and match one-sided Shampoo’s regret up to a constant factor, while significantly reducing computational cost. Lastly, using the online-to nonconvex conversion, we derive two matrix-based optimizers, Pion (from FTPL) and Leon (from FAML). We prove convergence guarantees for these methods in nonsmooth nonconvex settings, a guarantee that the popular Muon optimizer lacks. View details
Preview abstract Large-scale cloud-native systems generate continuous streams of operational alerts across distributed microservice architectures. On-call engineers must manually triage these alerts by correlating signals from heterogeneous observability tools, a process that is time-consuming, cognitively demanding, and prone to error. Despite advances in monitoring and anomaly detection, incident triage remains largely manual. This paper presents a declarative, large language model (LLM)–driven multi-agent approach to automating incident triage and Service Level Objective (SLO) monitoring. The proposed design constrains agent behavior using domain-expertauthored investigation workflows, enabling deterministic execution and reproducibility while preserving operational safety. The framework integrates a unified tool execution layer for interacting with diverse observability systems and an enhanced retrieval-augmented generation (RAG) pipeline optimized for operational knowledge retrieval. The approach has been evaluated in a production cloud environment spanning multiple microservices and geographic regions. Results show reductions in high-severity incident triage time from approximately 30 minutes to under 5 minutes, alert acknowledgement latency from minutes to seconds, and service onboarding effort from weeks to days. These findings suggest that constrained multi-agent systems can substantially reduce on-call cognitive load while maintaining reliability and human oversight. View details
CAST: Modeling Visual State Transitions for Consistent Video Retrieval
Yanqing Liu
Yingcheng Liu
Fanghong Dong
Budianto Budianto
Cihang Xie
Yan Jiao
2026
Preview abstract As video content creation shifts towards long-form narratives, retrieving and composing short clips into coherent storylines becomes a critical challenge. Standard retrieval formulations, however, perform context-agnostic retrieval, prioritizing local semantic alignment while neglecting procedural state and identity consistency across the narrative flow. To address this, we introduce the task of Consistent Video Retrieval (CVR) and establish a benchmark designed to diagnose such inconsistencies via semantic hard negatives. We propose CAST (Context-Aware State Transition), a lightweight adapter that models procedural progression as state-conditioned transitions. Conditioned on visual history, CAST predicts a gated residual vector ($\Delta$) to selectively update the state embedding, ensuring procedural coherence while preserving identity. Extensive experiments demonstrate that CAST significantly outperforms standard retrieval baselines on our CVR benchmark. Furthermore, we show its potential as a plug-and-play consistency verifier, guiding black-box generation models (e.g., Veo) toward coherent video continuations within long-form narratives. View details
OVERVIEW OF THE BLOCK-PARTITIONING FRAMEWORK IN AV2
Chi Yo Tsai
Yue Chen
Jayasingam Adhuran
Liang Zhao
2026
Preview abstract Block partitioning framework is a core component in any modern video coding standard, as it directly determines the block size used for predictions and transforms. Flexible block partitioning plays a crucial role in the compression efficiency of these standards. This paper provides a technical overview of the block partitioning framework in AV2 video codec, developed by Alliance for Open Media. Partitioning scheme for both coding blocks and transform blocks has been redesigned in AV2. Coding block partitioning is fully recursive with newly designed partitioning options. Also, newly introduced Semi-Decoupled Partitioning (SDP) option provides additional flexibility by allowing luma and chroma components to have decoupled coding block partition trees. On the other hand, the transform block partitioning has been redesigned to use single-level partitioning with more partitioning options. In this paper, we provide a technical overview of the block partitioning framework in AV2 and also provide tool-off test results for several block partitioning aspects. View details
Nudging Developers Toward Privacy: Evaluating the Impact of Personalized App Review Reports
Omer Akgul
Michelle L. Mazurek
USENIX Symposium on Usable Privacy and Security (SOUPS) (2026)
Preview abstract Mobile application developers often struggle to create accurate privacy notices or implement robust privacy practices due to limited expertise or resources. While users share unsolicited privacy feedback in app reviews, and prior research has characterized this privacy feedback, uncovering developer reactions to this feedback remains unexplored. This study explores whether personalized privacy review reports---summarizing real user feedback for a developer's own app---can effectively nudge them toward planning privacy improvements. We surveyed 42 app developers, presenting them with reports containing privacy themes, temporal trends, peer benchmarks, and emotion distributions derived from their apps' reviews. Our findings indicate that these privacy report interventions proved highly effective, with 76% (32 of 42) of participants finding at least one section of the report useful. Furthermore, exposure to the report increased the participants' intent to pursue privacy-relevant actions -- such as reorganizing the UI, enhancing privacy communications, or adding/removing features -- with 69% (29 of 42) of participants indicating an increased intent to do so. Almost all developers expressed a desire to receive such privacy reports periodically or on demand. These results indicate that making this style of report broadly available across the industry could foster a more privacy-conscious mobile ecosystem. View details
Preview abstract Optimizing large-language model (LLM) training and serving on large-sacle distributed systems with hundreds and thousands of accelerators is always a challenging task due to the fast evloving LLMs, strong domain expertise required, and various optimization goals from different worklaods. Existing methods rely on either handcrafted optimization performed by human experts, which is tedious and time-consuming or resource-intensive black-box searches, which lack the extensibility to keep pace with evolving models and hardware. To address this, we introduce PROMPTS, a novel multi-agent framework that complements traditional search methods with expert-informed reasoning. It automates the diagnosis of performance bottlenecks by synthesizing profiler data and leverages a knowledge base to propose optimized sharding configurations with detailed justifications. Across eight real-world production workloads, PROMPTS demonstrated remarkable efficiency and accuracy, delivering performance improvements of up to 434%. These workloads spanned diverse model architectures, hardware platforms, computational scales, and various stages of the machine learning lifecycle (pre-training, serving, and post-training). In every case, the configuration adopted by human engineers was identified within the agent's top three proposals from a single invocation. Furthermore, the agent's top-ranked recommendation was the one ultimately adopted in 87.5% of cases, showcasing its ability to not only find optimized solutions, but also to correctly prioritize them. Our work establishes PROMPTS as a scalable, extensible, and explainable methodology for AI-assisted performance engineering in large-scale ML systems. View details
GenAI on Google Cloud: Enterprise Generative AI Systems and AI Agents
Ayo Adedeji
Lavi Nigam
Stephanie Gervasi
O'Reilly Media, Inc. (2026)
Preview abstract In today's AI landscape, success depends not just on prompting large language models but on orchestrating them into intelligent systems that are scalable, compliant, and cost-effective. GenAI on Google Cloud is your hands-on guide to bridging that gap. Whether you're an ML engineer or an enterprise leader, this book offers a practical game plan for taking agentic systems from prototype to production. Written by practitioners with deep experience in AgentOps, data engineering, and GenAI infrastructure, this guide takes you through real-world workflows from data prep and deployment to orchestration and integration. With concrete examples, field-tested frameworks, and honest insights, you'll learn how to build agentic systems that deliver measurable business value. > Bridge the production gap that stalls 90% of vertical AI initiatives using systematic deployment frameworks > Navigate AgentOps complexities through practical guidance on orchestration, evaluation, and responsible AI practices > Build robust multimodal systems for text, images, and video using proven agent architectures > Optimize for scale with strategies for cost management, performance tuning, and production monitoring View details
×