Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11583 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract Image-based sexual abuse (IBSA) refers to the creation, sharing, or threats to share intimate images, whether real or synthetic. Advances in artificial intelligence (AI) now enable the generation of sexually explicit images from non-explicit source material, such as professional headshots and social media profiles. In this study, we found that across three regions (Australia, the United Kingdom, and the United States), 15% of a representative sample reported experiences of IBSA involving digitally altered images, including 6.9% who reported AI-generated IBSA (AI-IBSA) victimization specifically. Higher rates of victimization were observed among respondents under 35, LGBTQ+, and BIPOC participants, people with disabilities, and men. While victimization rates were higher among men, women were more likely to report a broader range of harms and more negative emotional and social impacts. The most common responses to victimization were platform-based actions, including blocking (69%) and reporting (58%). These findings highlight AI-IBSA as a growing and under-recognized form of harm, with important implications for prevention, platform governance, and support for victim-survivors. View details
Learning Conditional Averages
Marco Bressan
Nataly Brukhim
Nicolo Cesa-Bianchi
Emmanuel Esposito
Shay Moran
Maximilian Thiessen
COLT (2026)
Preview abstract We introduce the problem of learning \emph{conditional averages} in the PAC framework. The learner receives a sample labeled by an unknown target concept from a known concept class, as in standard PAC learning. However, instead of learning the target concept itself, the goal is to predict, for each instance, the average label over its \emph{neighborhood}---an arbitrary subset of points that contains the instance. In the degenerate case where all neighborhoods are singletons, the problem reduces exactly to classic PAC learning. More generally, it extends PAC learning to a setting that captures learning tasks arising in several domains, including explainability, fairness, and recommendation systems. %including explainability, fairness, and recommendation systems. Our main contribution is a complete characterization of when conditional averages are learnable, together with sample complexity bounds that are tight up to logarithmic factors. The characterization hinges on the joint finiteness of two novel combinatorial parameters, which depend on both the concept class and the neighborhood system, and are closely related to the independence number of the associated neighborhood graph. View details
Analyzing Bytes: Pre-Disassembly Static Binary Analysis
Soumyakant Priyadarshan
ChenCheng Jiang
R. Sekar
Proceedings of the ACM on Programming Languages, Association for Computing Machinery (2026), pp. 1127-1151
Preview abstract Binary code analysis plays a central role in numerous applications in software security, performance optimization, reverse engineering, and so on. Existing techniques need to first disassemble binaries into functions in assembly code before an analysis can be performed. However, disassembly and function identification have proven to be major challenges for complex variable-length instruction sets such as the x86. A recent trend has been to use static analysis to improve the accuracy of these tasks. This raises a chicken-and-egg problem: a disassembly is needed for static analysis, but a static analysis is needed for accurate disassembly! We overcome this problem by developing a novel static analysis approach that can operate before committing to a disassembly. Our analysis operates on the output of exhaustive disassembly that considers each possible offset in a binary as an instruction, and constructs what is known as a super-set control-flow graph (CFG). The central technical challenge in analyzing this CFG is that it mixes legitimate instructions with unintended ones, causing analysis results from invalid code paths to pollute legitimate ones. To overcome this challenge, we begin with a key new insight that if we focus on backward analyses, we can ensure accuracy of analysis results at intended instructions even though we have no idea where these intended instructions are! Moreover, our analysis operates in time that is linear in the size of the binary. Specifically, in O(n) total time, it yields analysis results for every one of the n offsets in an n-byte binary. For this task, it is orders of magnitude faster than previous techniques, as the previous techniques typically need to repeat the analysis many times. View details
Data-usage descriptors as search metadata: the case of food security data and the National Data Platform (2015-2025)
Julia Lane
Rafael Ladislau
Lauren Chenarides
Simon Porter
Manish Parashar
Scientific Data, 13 (2026)
Preview abstract Scientific data is a critical input into scientific research. Yet the research data landscape is constantly changing as new datasets emerge, others are retired, or some disappear altogether. Without a systematic way to track how datasets are used across a research field, researchers have no reliable method for identifying relevant data resources or locating communities that work with them. Data-usage descriptors can substantially advance research productivity by reducing the time that researchers spend finding new and relevant datasets in their research field, and the communities that use them. This paper describes how to generate data-usage descriptors by finding how datasets are used in publications and then linking the dataset information to the publication metadata. It also shows how usage descriptors can be used to find other related datasets and their usage. It concludes by arguing that the approach represents a critical piece of foundational infrastructure that could be deployed in repositories as part of a referenceable, navigable, and contextual data framework. This article contains a reproducible workflow for constructing data-usage descriptors, based on analyzing the full text of publications in the Dimensions database. The illustrative use case is research on food security. The illustrative repository is the National Data Platform. View details
Multi-agent cooperation through in-context co-player inference
Rajai Nasser
Alexander Meulemans
Marissa Weis
João Sacramento
Maciej Wołczyk
Rif A. Saurous
2026
Preview abstract Achieving cooperation among self-interested agents remains a fundamental challenge in multi-agent reinforcement learning. Promising recent work has shown that cooperation can be established between ``learning-aware'' agents that explicitly account for and shape the learning dynamics of their co-players. However, existing approaches typically rely on hardcoded, often inconsistent, assumptions about co-player learning rules or enforce a strict separation between ``naive learners'' updating on fast timescales and ``meta-learners'' observing these updates. Here, we demonstrate that the in-context learning capabilities of sequence models allow for co-player learning awareness without requiring hardcoded assumptions or explicit timescale separation. We show that training sequence model agents against a diverse distribution of co-players naturally induces \textit{in-context best-response} strategies, effectively functioning as learning algorithms on the fast intra-episode timescale. We find that the cooperative mechanism identified in prior work—where vulnerability to extortion drives mutual shaping—emerges naturally in this setting: in-context adaptation renders agents vulnerable to extortion, and the resulting mutual pressure to shape the opponent's in-context learning dynamics resolves into the learning of cooperative behavior. View details
Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting
David Autor
Tanya Rodchenko
Joshua Martin
Zanna Iscenko
Scott Strand
David Pearl
Melissa Ferere
NBER (2026)
Preview abstract Whether AI assistance builds or erodes professional expertise is unsettled. In a pre-registered three-month randomized controlled trial, we gave 133 practicing patent lawyers at eleven U.S. intellectual property law firms access to a custom AI drafting assistant and measured both their performance while using AI and their professional judgment afterward without it. All work was scored by blinded expert patent attorneys. Paralleling findings from other white-collar domains, AI access raised the quality of work delivered on benchmark patent drafting tasks at 10 days (0.34 SD, p = 0.03) and 90 days (0.38 SD, p = 0.01), with larger gains among junior lawyers. After three months, all subjects redlined an existing patent application without AI, a core task of patent practice requiring expert judgment. Treated lawyers outperformed controls by 0.32 SD (p = 0.04), but this advantage was concentrated entirely among senior lawyers (0.45 SD, p = 0.02). Junior lawyers showed no average gain; their scores instead bifurcated, with sharply fewer mediocre scores offset by more poor and more good ones. The largest gains from AI thus accrued to the lawyers who retained the least. Foundational expertise may be a prerequisite for extracting durable skill from AI-assisted practice. View details
POLCA: Stochastic Generative Optimization with LLM
Xuanfei Ren
Allen Nie
Tengyang Xie
Ching-An Cheng
2026
Preview abstract Optimizing complex systems, ranging from LLM prompts to multi-turn agents, traditionally requires labor-intensive manual iteration. We formalize this challenge as a stochastic generative optimization problem where a generative language model acts as the optimizer, guided by numerical rewards and text feedback to discover the best system. We introduce Prioritized Optimization with Local Contextual Aggregation (POLCA), a scalable framework designed to handle stochasticity in optimization -- such as noisy feedback, sampling minibatches, and stochastic system behaviors -- while effectively managing the unconstrained expansion of solution space. POLCA maintains a priority queue to manage the exploration-exploitation tradeoff, systematically tracking candidate solutions and their evaluation histories. To enhance efficiency, we integrate an ε-Net mechanism to maintain parameter diversity and an LLM Summarizer to perform meta-learning across historical trials. We theoretically prove that POLCA converges to near-optimal candidate solutions under stochasticity. We evaluate our framework on diverse benchmarks, including τ-bench, HotpotQA (agent optimization), VeriBench (code translation) and KernelBench (CUDA kernel generation). Experimental results demonstrate that POLCA achieves robust, sample and time-efficient performance, consistently outperforming state-of-the-art algorithms in both deterministic and stochastic problems. The codebase for this work is publicly available at this https URL. View details
OpenClaw in the Wild: Security Analysis of Autonomous Agents
Wanlun Ma
Qing-Long Han
Xiaogang Zhu
Wei Zhou
Junwu Xiong
Peter Ren
Sheng Wen
Yang Xiang
IEEE/CAA Journal of Automatica Sinica, 13 (2026), pp. 1257 - 1273
Preview abstract Autonomous self-hosted AI agent platforms are rapidly evolving from prompt-response assistants into persistent systems that can maintain long-lived state, invoke tools, ingest external content, and execute environment-changing actions. While this transition enables practical automation, it also introduces lifecycle security risks that cannot be fully explained by prompt-level analysis alone. In this paper, we present a security analysis of OpenClaw as a representative autonomous agent operating environment. We further frame OpenClaw as a concrete case study for broader security challenges in emerging agent ecosystems. We adopt a trust-boundary-first perspective and analyze how attacks propagate across five boundary classes: Channel-Access, Session-and-State, Tool-Execution, External-Content, and Extension Supply-Chain. Our results show that threats such as indirect prompt injection, memory poisoning, unsafe tool invocation, data exfiltration, and malicious skill abuse are not isolated anomalies; they are stage-specific manifestations of a common systems problem in which untrusted influence progressively crosses into higher-privilege contexts. Building on this analysis, we discuss defense-in-depth implications for OpenClaw deployments, including boundary-aware isolation, capability-scoped tool mediation, memory integrity controls, extension governance, and evidence-oriented operational oversight. The study provides a practical framework for evaluating and hardening long-running, tool-capable, autonomous AI agents in realistic deployment settings. View details
Preview abstract Quantization methods have significantly improved the compute and memory efficiency of Large Language Model (LLM) training. However, existing approaches still rely on accumulating their updates into high precision: concretely, gradient updates must be applied to a high-precision weight buffer, known as \textit{master weights}. This buffer introduces substantial memory overhead, particularly for Sparse Mixture of Experts (SMoE) models, where model parameters and optimizer states dominate memory usage. In this work, we introduce the Error-Compensating Optimizer (ECO), which \textit{for the first time} enables the complete elimination of master weights by directly accumulating updates into quantized parameters, by leveraging existing optimizer states. ECO quantizes the weights after every gradient step and injects the resulting quantization error into the optimizer's momentum buffer, creating an error-feedback loop with zero additional memory overhead for quantization. Beyond its practical efficiency, ECO comes with theoretical guarantees. Specifically, under standard assumptions, naive master weight removal can lead to unbounded drift from the ideal parameter trajectory, whereas ECO provably bounds this drift, ensuring stable convergence. We validate ECO across a range of models, including small transformers (30M--800M), Gemma-3 1B, and an SMoE 2.1B model, using FP8 quantization. In all cases, ECO achieves near-lossless accuracy compared to high-precision baselines. For large SMoE models, ECO reduces memory usage by up to 25\%, establishing a new Pareto frontier for the trade-off between static memory and training loss. View details
Preview abstract To meet aggressive time-to-market goals, modern mobile SoC architectures require the concurrent development of custom Compute-IPs and surrounding subsystem integration logic (1PIPs and 3PIPs). However, this parallel execution creates a critical verification deadlock: the subsystem cannot be validated until both the Compute-IP and volatile, in-flight 1PIPs reach physical RTL maturity. Consequently, "integration-killer" bugs—such as protocol handshaking deadlocks and clock/reset sequencing mismatches—remain hidden until late in the design cycle when RTL rework costs are prohibitive. To break this bottleneck, we present a verification-driven methodology utilizing a silicon-proven Golden Proxy, Direct-Execution Traffic Profiles, Programmable Sequencers and Automated Protocol Converters. This framework completely decouples parallel hardware dependencies, pre-pulling critical inter-IP mismatch discoveries months ahead of traditional integration milestones. View details
Preview abstract Lessons learned from building an agent to convert the R Forecast Package into a JAX api Skander Hannachi, Jasmeet Bhatia, Dennis Kashkin, Anna Novakovska - Applied AI Engineering, Google Cloud {shannachi@ jasmeetbhatia@ kashkin@ anovakovska@}google.com Accepted at ISF 2026 Abstract: The Forecast package in R, is one of the most popular and established frameworks for learning and working with statistical (local) time series models. However, over the last decade or so, the language of choice for data analysis and mathematical modeling has become Python, especially, since, unlike R, the latter provides multiple options for running models on dedicated hardware accelerators (TPUs/GPUs), and on distributed compute infrastructure. In this study, we implement an LLM agent based code conversion pipeline that automatically converts the code from the Forecast package (R, C++), into JAX/PAX, a more recent modeling framework dedicated specifically for running compute heavy modeling tasks on TPUs and GPUs. We run our specialized code conversion agent against three core models of the Forecast package: TBATs, auto.arima(), and ETS(), with an aim to provide the exact same user experience as the original R API, but with the underlying model selection, model fitting, and forecast generation running all using JAX and PAX operations, running in a Python environment. We report the results of our experiments and discuss the challenges we observed while running it. We compile these into a skills markdown file, which can be used by other agents intended to perform similar experiments. We provide the JAX based implementations, along with the skills file in an accompanying open source repo. This is not the first attempt at converting the Forecast package into Python. Those other efforts however are significantly labor intensive, especially when it comes to ensuring parity with the source modeling APIs and quality control in general. Moreover, such projects rely heavily on the long term commitment of both community members and institutional contributors to the conversion effort. The purpose of our effort is to show how this agent based process can be applied to automate any data science package upgrade or language conversion process with minimal contributors required outside of the core package maintainer team. Especially since the concept of agent skills files makes the process inherently self-improving, both within the scope of a single conversion effort, as well as across multiple long term conversion efforts. For example, the same approach can be applied to a future effort for upgrading the Forecast package to work with Julia, an even more recent and promising modeling language, while benefitting from the R-2JAX lessons learned. View details
Preview abstract As the ECMAScript specification evolves, industrial-scale JavaScript compilers face the challenge of supporting modern language syntax while maintaining compatibility for diverse execution environments. Traditionally, compilers solve this by running transpilation passes in a monolithic pipeline, where the transpilation passes are chosen to execute strictly based on a target language level. This results in significant computational waste, as compilers perform expensive Abstract Syntax Tree (AST) traversals to lower features that may not exist in the actual input source code. We present a static analysis improvement that conditionally executes transpiler passes based on accurately tracking and dynamically maintaining the exact set of language features seen in the compilation unit throughout the transpilation process. It is implemented in the production Google Closure Compiler. By populating and maintaining a FeatureSet at every JavaScript script-level, it dynamically skips running the unnecessary lowering passes. We detail the architectural safeguards - including strategic pass ordering and dynamic validation of the transpiled code for feature-correctness. Evaluation of this improvement on large-scale production applications produced a considerable reduction in compilation time and saved compute and memory usage. View details
Preview abstract Time-series forecasting has traditionally been evaluated solely on numerical accuracy, treating models as "black boxes'' that fail to capture the underlying reasoning. To address this gap, we introduce TFRBench, a novel benchmark designed to evaluate the reasoning capabilities of forecasting systems alongside their numerical accuracy. Unlike existing benchmarks, TFRBench requires models to generate verifiable natural language reasoning by analyzing cross-channel dependencies, identifying strategic trends, and justifying significant events using external context. To construct this benchmark, we propose a systematic multi-agent framework comprising Reasoning, Search, Verifier, Forecasting, and Summary agents. Our benchmark spans five diverse domains including Energy, Sales, Web/CloudOps, Transportation, and Finance, covering 10 distinct datasets. Qualitative evaluation confirms that our generated reasoning is highly faithful and effective; specifically, Large Language Models (LLMs) prompted with our generated reasoning demonstrate significantly improved forecasting accuracy compared to direct forecasting with LLMs. Conversely, benchmarking experiments reveal that off-the-shelf LLMs consistently struggle with both reasoning (shows lower LLM-as-Judge scores) and direct numerical forecasting (MAE and MASE), frequently failing to capture domain-specific dynamics. TFRBench thus establishes a new standard for interpretable, reasoning-based evaluation in time-series forecasting. View details
×