Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11612 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract We study a quantized prefix estimator for inner products that turns a randomly rotated TurboQuant-style representation into a cheap Johnson–Lindenstrauss-like search signal. The idea is simple: rotate the vectors once, keep only a short prefix of coordinates for fast scoring, and quantize the database-side prefix with an unbiased scalar quantizer. We prove that this estimator is unbiased and that its error separates cleanly into two interpretable sources: prefix truncation from using only r coordinates, and quantization error from using b bits per coordinate This separation is useful in systems because the prefix can be exposed as a lightweight filter without building a separate projection index. In ParlayANN graph search, a 64-coordinate truncated view of existing TQ4 codes can replace a separately stored JL256 filter before full-precision reranking, adding only prefix-scale and query lookup-table bookkeeping. In k-means, the same estimator accelerates the dominant point–centroid assignment kernel while preserving exact centroid norms. Empirically, the truncated-TQ filter tracks the JL recall–throughput frontier across five graph-search datasets while reusing the quantized representation already present in the index. View details
QBAT: Model-based Query Budget Autotuner for Clustering-based Approximate Nearest Neighbor Search
Jonghyun Bae
Tae Jun Ham
Alan Li
Yannis Papakonstantinou
Proceedings of the VLDB Endowment (2026), pp. 3091-3104
Preview abstract Approximate nearest neighbor search (ANNS) is a critical component in modern data-intensive applications, but its performance is often hindered by the use of a static query budget parameter. This one-size-fits-all approach, even if well-tuned, fails to account for the varying difficulty of individual queries, inevitably leading to suboptimal latency on easy queries and poor accuracy on hard ones. This paper introduces QBAT, a query-aware budget autotuner designed to resolve this dilemma. By analyzing query-specific features offline, QBAT dynamically allocates an appropriate budget for each query. We explore two predictive models: a highly accurate gradient-boosted decision tree and a simple, interpretable heuristic formula derived using the AlphaEvolve framework. These models can optimize budget allocation for both system performance or recall consistency priorities. Evaluations on large-scale datasets demonstrate that QBAT reduces total searched budget by up to 68.8% in the consistency mode on ScaNN, the state-of-the-art clustering-based ANNS method, while simultaneously enforcing a strict per-query recall target, a scenario where static budgets are notoriously inefficient and wasteful. View details
Learning from Equivalence Queries, Revisited
Mark Braverman
Roi Livni
Shay Moran
Kobbi Nissim
COLT (2026)
Preview abstract Modern machine learning systems, such as generative models and recommendation systems, often evolve through a cycle of deploying a model, observing user interactions, and updating the model intermittently based on feedback. This mode of learning contrasts with common supervised learning frameworks, which focus on loss or regret minimization over a shared sequence of prediction tasks. Motivated by this deployment-driven learning cycle, we revisit the classical model of learning from equivalence queries, introduced by Angluin, which provides a simple abstraction of such interactions: a learner repeatedly proposes hypotheses and, whenever the deployed hypothesis is inadequate, receives a counterexample tailored to that hypothesis. Under fully adversarial counterexample generation, however, this model exhibits overly pessimistic worst-case behavior. Moreover, most existing work on learning from equivalence queries considers the \emph{full-information} setting, where the learner observes not only a counterexample but also its correct label. This is an assumption that does not always align with natural interactive settings. To address these considerations, we restrict the environment to generate counterexamples in a less adversarial manner by introducing a broad class of counterexample generators, which we call \emph{symmetric}. Informally, such symmetric counterexample generators select counterexamples based only on the symmetric difference between the hypothesis and the target, and encompass natural feedback mechanisms such as random counterexamples, as well as generators that select counterexamples minimizing a prescribed complexity measure over the instance space. Within this framework, we study learning from equivalence queries under both full-information and bandit feedback. We establish tight bounds on the number of learning rounds in both settings and outline directions for future research. Our techniques rely on a game-theoretic perspective on symmetric adversaries and combine adaptive weighting algorithms with minimax arguments. View details
Preview abstract Automating AI research differs from general software engineering due to computationally expensive evaluation (e.g., model training) and opaque performance attribution. Current LLM-based agents struggle here, often generating monolithic scripts that ignore execution costs and causal factors. We introduce MARS (Modular Agent with Reflective Search), a framework optimized for autonomous AI research. MARS relies on three pillars: (1) Budget-Aware Planning via cost-constrained Monte Carlo Tree Search (MCTS) to explicitly balance performance with execution expense; (2) Modular Construction, employing a "Design-Decompose-Implement" pipeline to manage complex research repositories; and (3) Comparative Reflective Memory, which addresses credit assignment by analyzing solution differences to distill high-signal insights. MARS achieves state-of-the-art performance among open-source frameworks on MLE-Bench under comparable settings, maintaining competitiveness with the global leaderboard's top methods. Furthermore, the system exhibits qualitative "Aha!" moments, where 63% of all utilized lessons originate from cross-branch transfer, demonstrating that the agent effectively generalizes insights across search paths. View details
Open and Emergent Problems in Agentic Privacy and Security: A Contextual Angle
Sahar Abdelnabi
Borja de Balle Pigem
Sebastian Benthall
Eleanor Birrell
Kamalika Chaudhuri
Madiha Zahrah Choksi
Rachel Cummings
Adam Davies
Dj Dvijotham
Seliem El-Sayed
Ferdinando Fioretto
Matt Franchi
Adria Gascon
Roxana Geambasu
Sahra Ghalebikesabi
Hamed Haddadi
Jamie Hayes
Ashish Hooda
Amir Houmansadr
Somesh Jha
Chloé Kiddon
Tadayoshi Kohno
Abdullatif Köksal
Haoran Li
Tianshi Li
Nathan Malkin
Sarah Meiklejohn
Niloofar Mireshghallah
Sewoong Oh
Katherine Ortiz
Francesco Pinto
Franziska Roesner
Edo Roth
Jacqueline Rowe
Khawaja Shams
Ilia Shumailov
Yan Shvartzshnaider
Dawn Song
Jose Such
Octavian Suciu
Pierre Tholoniat
Trishita Tiwari
Hal Triedman
Ren Yi
Wen Zhang
Xuhui Zhou
Kassem Fawaz
Stefan Mellem
Helen Nissenbaum
Google (2026)
Preview abstract The vision of an ecosystem of general and highly capable autonomous agents challenges accepted principles of secure and privacy-preserving system engineering. Inspired by the theory of Contextual Integrity, we argue that in order for agents to act appropriately, their behaviors should comply with societal norms and expectations. This manuscript explores contextual approaches to engineering agentic systems that embody an understanding of societal norms and uphold the appropriateness of actions by design. It presents a list of key open research problems in this area to create awareness among the broader research community across academia, government, and industry. View details
Preview abstract This paper introduces XMob, a novel differentiable traffic simulation framework built in JAX to advance traditional models like SUMO’s mesoscopic simulator. By leveraging JAX’s capabilities for vectorized, hardware-accelerated computation (GPU/TPU), XMob achieves orders-of-magnitude speedups, enabling large-scale urban network simulations and extensive counterfactual analyses. A key innovation is XMob’s inherent differentiability, facilitating direct integration with gradient-based optimization for tasks such as demand calibration and network parameter estimation, significantly outperforming black-box approaches. Furthermore, XMob can be used in Physics-Informed Machine Learning (PIML) pipelines to enhance data-driven augmentation, embedding domain principles like flow conservation and shockwave theory. This ensures physically plausible and robust predictions, even for unobserved scenarios such as lane modifications. The hybrid architecture, combining a deterministic JAX core with incremental machine learning, offers a scalable and efficient solution for modern traffic simulation and optimization challenges. View details
Preview abstract We study cooperative multi-agent reinforcement learning in the setting of reward-free exploration, where multiple agents jointly explore an unknown MDP in order to learn its dynamics (without observing rewards). We focus on a tabular finite-horizon MDP and adopt a phased learning framework. In each learning phase, multiple agents independently interact with the environment. More specifically, in each learning phase, each agent is assigned a policy, executes it, and observes the resulting trajectory. Our primary goal is to characterize the tradeoff between the number of learning phases and the number of agents, especially when the number of learning phases is small. Our results identify a regime change governed by the horizon $H$. When the number of learning phases equals $H$, we present a computationally efficient algorithm that uses only $\tilde{O}(S^6 H^6 A / \epsilon^2)$ agents to obtain an $\epsilon$ approximation of the dynamics (i.e., yields an $\epsilon$-optimal policy for any reward function). We complement our algorithm with a lower bound showing that any algorithm restricted to $\rho < H$ phases requires at least $A^{H/\rho}$ agents to achieve constant accuracy. Thus, we show that having $\Theta(H)$ learning phases is both necessary and sufficient when restricting the number of agents to be polynomial. View details
Preview abstract Securing the Agentic Enterprise: Threat Modeling, Anomaly Detection, and Governing Autonomous Multi-Agent Systems addresses the critical security and governance gaps emerging as enterprises transition from human-supervised copilots to autonomous agentic workflows. As software processes gain the ability to reason, decompose natural language objectives, and execute multi-step tool calls at machine speed, traditional syntactic security boundaries (like firewalls and static analysis) become obsolete. This book provides security architects, CISOs, and platform engineers with a practical, architecture-level blueprint for securing this new paradigm. It explores novel attack vectors such as indirect prompt injections and consumption-based economic threats and provides frameworks for robust mitigation. Key topics include modernizing agentic identity, implementing semantic firewalls, transition-state anomaly detection, and applying zero-trust principles to autonomous execution contexts. Bridging the gap between high-level ethical guidelines and isolated model safety, this guide prepares practitioners to confidently deploy and govern enterprise-grade autonomous systems. View details
Preview abstract We introduce ALPS (Activation-based Length Prediction for Scheduling), a method for predicting LLM generation length from prefill activations before any tokens are generated. Unlike existing approaches that require model fine-tuning or complex entropy-weighted pooling, ALPS uses a simple linear probe on the last-token activation at intermediate layers. We discover that generation length is encoded in prefill representations: a ridge regression probe achieves R-squared > 0.85 across three model families. Validation across Llama-3.1-8B, Gemma-2-9B, and Qwen-2.5-7B demonstrates: (1) intermediate layers generally perform well, with some architectural variation; (2) simple last-token extraction outperforms complex pooling strategies; (3) activations improve substantially over surface-feature baselines (24 percentage points over input length plus lexical features). The best models achieve R-squared = 0.943 (Gemma), R-squared = 0.880 (Llama), and R-squared = 0.857 (Qwen) with MAE of 38-80 tokens. All test prompts terminated naturally (100% EOS), eliminating truncation confounds. While our evaluation uses 200 curated prompts—sufficient for demonstrating the phenomenon but requiring broader validation—cross-validation confirms generalization beyond training data. ALPS enables practical applications including budget-constrained inference, request scheduling, and resource allocation. The probe adds negligible overhead (~16KB direction vector, single dot product), making ALPS practical for production deployment. View details
Quantum Advantage in Topological Data Analysis via Mayer Homology
Anh Nghiem
Dominic Berry
Trung Phan
Guo-Wei Wei
Ryu Hayakawa
arXiv:2609.28058 (2026)
Preview abstract Prior work has explored quantum algorithms for topological data analysis (TDA), revealing the possibility of exponential quantum speedups in estimating the ratios of Betti numbers to the dimension of the combinatorial Laplacian. However, this quantity is only non-vanishing and efficient-to-quantumly-estimate when Betti numbers are exponentially large, a case for which concrete examples are rarely known. Furthermore, certain randomized classical algorithms are sometimes efficient in this regime. Thus, the prospect of achieving quantum advantage in conventional TDA appears fairly narrow. Here, we address these challenges to the quantum advantage in TDA by developing quantum algorithms for Mayer homology, which generalize simplicial homology to N-nilpotent boundary operators (∂N=0) and have recently been successfully applied to real-world TDA contexts. We introduce an efficient quantum algorithm for estimating Mayer Betti numbers and their persistent counterparts. We then prove that for high-order simplices, Mayer Betti numbers are often exponentially large in the dense regime, which ameliorates the normalization bottleneck of conventional quantum TDA. In the same regime, we argue that existing dequantization algorithms developed for conventional TDA, when applied to Mayer homology, generally lose theoretical guaranties, facing certain structural barriers that prevent their practical utilities. We also provide logical resource estimates revealing that a quantum computer with roughly a few hundred qubits and sixty million Toffoli gates could solve Mayer homology problems beyond the capabilities of known classical approaches. Finally, we discuss real-world applications of Mayer homology in genomics, supersymmetry, drug discovery, and neuroscience, revealing the potential of our quantum algorithm to deliver real-world impacts via Mayer homology. View details
Preview abstract We study the computational cost of differential privacy in terms of memory efficiency. While the trade-off between accuracy and differential privacy is well-understood, the inherent cost of privacy regarding memory use remains largely unexplored. This paper establishes for the first time an unconditional space lower bound for user-level differential privacy by introducing a novel proof technique based on a multi-player communication game. Central to our approach, this game formally links the hardness of low-memory private algorithms to the necessity of ``contribution capping''---tracking and limiting the users who disproportionately impact the dataset. We demonstrate that winning this communication game requires transmitting information proportional to the number of over-active users, which translates directly to memory lower bounds. We apply this framework, as an example, to the fundamental problem of estimating the number of distinct elements in a stream and we prove that any private algorithm requires almost $\widetilde{\Omega}(T^{1/3})$ space to achieve certain error rates in a promise variant of the problem. This resolves an open problem in the literature (by Jain et al. NeurIPS 2023 and Cummings et al. ICML 2025) and establishes the first exponential separation between the space complexity of private algorithms and their non-private $\widetilde{O}(1)$ counterparts for a natural statistical estimation task. Furthermore, we show that this communication-theoretic technique generalizes to broad classes of problems, yielding lower bounds for private medians, quantiles, and max-select. View details
Data-usage descriptors as search metadata: the case of food security data and the National Data Platform (2015-2025)
Julia Lane
Rafael Ladislau
Lauren Chenarides
Simon Porter
Manish Parashar
Scientific Data, 13 (2026)
Preview abstract Scientific data is a critical input into scientific research. Yet the research data landscape is constantly changing as new datasets emerge, others are retired, or some disappear altogether. Without a systematic way to track how datasets are used across a research field, researchers have no reliable method for identifying relevant data resources or locating communities that work with them. Data-usage descriptors can substantially advance research productivity by reducing the time that researchers spend finding new and relevant datasets in their research field, and the communities that use them. This paper describes how to generate data-usage descriptors by finding how datasets are used in publications and then linking the dataset information to the publication metadata. It also shows how usage descriptors can be used to find other related datasets and their usage. It concludes by arguing that the approach represents a critical piece of foundational infrastructure that could be deployed in repositories as part of a referenceable, navigable, and contextual data framework. This article contains a reproducible workflow for constructing data-usage descriptors, based on analyzing the full text of publications in the Dimensions database. The illustrative use case is research on food security. The illustrative repository is the National Data Platform. View details
VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion
Zador Pataki
Paul-Edouard Sarlin
Marc Pollefeys
ECCV (2026)
Preview abstract Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap. View details
×