Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11597 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Reinforcement Learning Control of Quantum Error Correction
Cameron Maxfield
Guifre Vidal
Bob Buckley
Jonathan Waltz
Christopher Wood
Reza Molavi
John Mark Kreikebaum
Rajeev Acharya
David Sobel
Abeer Vaishnav
Ali Hadjikhani
Ryuho Kudo
Wendy Leung
Brett Buchea
Ningfeng Zhu
Shirin Montazeri
Jamie Yao
Bicheng Ying
Eric Mascot
Lenny Fuste
Zhenjie Zou
Rodrigo Cortinas
Matt Lloyd
Clarke Smith
Kris Ottosson
Emma Ropes
Felix Borjans
Rebecca Potter
Sean Harrington
Jeremy Hilton
David Enriquez
Stephen Heslin
Paula Heu
Daniel Lundahl
Elliot Young
Alex Crook
Fedor Kostritsa
Roberto Rodriguez
Chia Ni
Kim Ming Lau
Priyanka Thiruraman
Martin Damyanov
Logan Oas
Dmitry Abanin
Oscar Higgott
Aaron Shorter
Steve Habegger
Aniket Maiti
Ryan Kaufman
Valerie Ehimhen
Sayra Alcaraz
Marcos Flores
Elizabeth Rossi
Aria Shahingohar
Dario Rosenstock
Travis Weidel
Steven Waltman
Kristi Wong
Murat Sarihan
Arun Kumar
Vladimir Shvarts
Matt Reagor
Alfredo Torres
Michael Qian
Anthony Megrant
Charles Neill
Christopher Hudspeth
Michael Hamilton
Bill Huggins
Laura De Lorenzo
Tan Ha
Ran Zhang
Dar Gilboa
Nicholas Bushnell
Sherman Peek
David Rhodes
Leigh Martin
Mike Shearn
Vlad Kurilovich
David Browne
Spencer Small
Brian Ballard
Will Oliver
Lior Ella
Orion Pritchard
Josh Cogan
Rachel Resnick
Dmitri Maslov
Jose Guerrero
Paul Masih Das
Theodore White
Helge Gehring
Nikita Astrakhantsev
Can Knaut
Maddy Woodson
Brooks Foxen
Frank Arute
Alejo Grajales Dau
Yaxing Zhang
Aaron Szasz
Alexander Lill
Justin Ledford
Xiaoxuan Jin
Andreas Kabel
Sid Madhuk
Orion Martin
Catherine Vollgraff Heidweiller
Gabrielle Roberts
Juan Campero
Juhwan Yoo
Robert Salazar
Michael Newman
Arpit Ranadive
James Goeders
William Giang
Gonzalo Garcia
Agnetta Cleland
Maddie Taylor
Dogan Timucin
Ross Alcaraz
Hui Kang
Johannes Bausch
William Courtney
Robert Gasca
Kevin Satzinger
Meghan Voorhees
Silas Chen
Laleh Beni
Andrew Dunsworth
Jamal Busnaina
Pavel Laptev
Kiseo Kang
Shannon Wang
Paul Donohoe
Paul Conner
Vadim Smelyanskiy
James Spencer
Benjamin Chiaro
Grayson Young
Tim Burger
ILYA Drozdov
Peter Brooks
Jordan Suchard
Austin Fowler
Jimmy Chen
Alec Eickbusch
Francisco Heras
Hung-Shen Chang
Michael Broughton
Jeanne Hartshorn
Aviv Elbag
Martin Bigdeli
Tanner Hadick
Juan Atalaya
Mahmoud Elzouka
Melvin Mathews
Alex Sztein
Markus Ansmann
Pavol Juhas
Bryan Cochrane
Murray Ich Nguyen
Ashley Maloney
Will Livingston
Roberto Collins
Ming Li
Élie Genois
Jeremiah Ford
Christopher Garrick
Sayan Das
David Peterson
Eifu Tomita
Suhas Ganjam
Reno Hiltermann
Dylan Bowers
Bryce Kobrin
Yu Chen
Dan Riley
Leon Brill
Barrett Spells
Ben Curtin
Mike Hucka
Seneca Meeks
Sebastian Molina
Tiano Lange-Dei
Georg Aigeldinger
Ashley Huff
Wing Li
ZLATKO MINEV
Monica Hansen
Sebastian Schroeder
Walt Askew
Dietrich Graumann
Elias Portoles
Stijn de Graaf
Matt Cockrell
Harold Cook
Masaya Fukami
Noah Shutty
Ed Gonzales
Robert Geiger
Amir Karamlou
Loick Le Guevel
Ebrahim Forati
Justin Vargas
Doug Thor
Joel Grebel
Lucia De Rose
LILY LI
Dave Landhuis
Emma Rosenfeld
Hsin-Yuan (Robert) Huang
Kenny Lee
Shaun Jevons
Ping Yeh
Amira Abbas
Kunal Arya
Henry Schurkus
Hector Bates
Ganesh Ramachandran
Sergey Vdovichev
Brayden Ware
Max Schaefer
Cheng Xing
Brandon Langley
Anthony Cabrera
Michel Devoret
Cody Jones
Vlad Sivak
Mert Torunbalci
Ben Kueffler
Chaitali Joshi
Raja Gosula
Joy Lee
Alexander Korotkov
Thomas Edlich
Aditya Locharla
Nathan Lacroix
George Sterling
Hao Tran
Kostyantyn Kechedzhi
Trond Andersen
Alexandre Bourassa
Aaron Lunt
Alan Fung
Alex Pizzuto
Salvatore Mandra
Alex Greene
Vitali Kutsko
Kannan Sankaragomathi
Sofia Springer
Vinicius Ferreira
Raymond Orosco
Nature (2026)
Preview abstract The promise of fault-tolerant quantum computing is challenged by environmental drift that relentlessly degrades the quality of quantum operations. The contemporary solution, halting the entire quantum computation for recalibration, is unsustainable for the long runtimes of the future algorithms \cite{reiher2017elucidating,gidney2025factor}. We address this challenge by unifying calibration with computation, granting the quantum error correction process \cite{ryan2021realization,krinner2022realizing,sivak2023real,acharya2024quantum, bluvstein2024logical,bluvstein2025architectural,lacroix2025scaling} a dual role: its error detection events are not only used to correct the logical quantum state, but are also repurposed as a learning signal, teaching a reinforcement learning (RL) agent \cite{silver2017mastering,mnih2015human,levine2016end, shalev2016safe,ouyang2022training} to continuously steer the physical control parameters and stabilize the quantum system during the computation. We experimentally demonstrate this framework on a Willow superconducting processor, improving the logical stability of the surface code 3.5-fold against injected drift. By synthesizing our full suite of technological advances, including RL fine-tuning of the entire system and near-optimal decoding \cite{senior2025scalable, beni2025tesseract}, we achieve record performance of the surface and color codes, with average logical error per cycle of $\varepsilon_L=7.7\times10^{-4}$ and $\varepsilon_L=8.2\times10^{-3}$ respectively. Simulations of surface codes up to distance-15 with tens of thousands control parameters confirm the scalability of our RL framework, revealing an optimization speed that is independent of the system size. This work thus enables a new paradigm: a quantum computer that learns to self-improve directly from its errors and never stops computing. View details
Managing and Securing Google's Fleet of Multi-Node Servers
Richard Hanley
Havard Skinnemoen
Andrés Lagar-Cavilla
Michael Wong
Jon McCune
Jeff Andersen
Kishan Prasad
Patrick Leis
Shiva Rao
Chris Koch
Jad Baydoun
Anna Sapek
Communications of the ACM, 69:3 (2026), pp. 82 - 92
Preview abstract Server hardware and software co-design for a secure, efficient cloud. View details
TDXRay: Microarchitectural Side-Channel Analysis of Intel TDX for Real-World Workloads
Tristan Hornetz
Hosein Yavarzadeh
Albert Cheu
Adria Gascon
Lukas Gerlach
Michael Schwarz
Ruiyi Zhang
IEEE Security & Privacy (S&P) (2026)
Preview abstract Confidential computing with VM-based trusted execution environments (TEEs) promises to protect code and data from a privileged cloud operator, enabling privacy-preserving workloads ranging from medical analytics to AI inference. However, most deployments exclude microarchitectural side channels from their threat model, shifting the burden to application developers who lack practical, general-purpose tools to assess (let alone mitigate) leakage. This gap is problematic: host-observable effects such as page-fault patterns, shared-cache contention, performance-counter surrogates (where available), and fine-grained timing primitives (e.g., MWAIT) can still reveal high-level secrets even when memory remains encrypted. We present TDXRay, an open-source framework that systematizes the evaluation of side-channel risk for confidential VMs in Intel TDX. TDXRay exposes unified interfaces to exercise and measure several attack primitives—including controlled-channel attacks via page tables, cache-based contention/occupancy probes, performance-counter–derived signals, and timing channels—against unmodified guest workloads. Using TDXRay, we build two end-to-end case studies: (1) a classic AES T-table attack in which a malicious hypervisor recovers the secret key from access-pattern leakage, and (2) an LLaMA inference attack in which the host infers user prompts by monitoring memory accesses during tokenization and embedding lookups. Across both, we show that a host with no direct access to guest memory can reconstruct sensitive information by observing only externalized microarchitectural signals. View details
Preview abstract Contrail cirrus represents a critical component of aviation’s non-CO2 climate impact, but its net radiative forcing, the balance between longwave warming and shortwave cooling, remains poorly constrained by direct observations. As a result, current assessments rely almost exclusively on microphysical models such as CoCiP and global climate simulations. Existing empirical estimates are largely restricted to young, linear tracks, because satellite detection masks have a poor recall of contrails once they spread and merge with natural cirrus, leaving a structural gap in our understanding of long-lived, non-linear contrail cirrus. To address this we use a causal framework that isolates the net radiative contrail effect of flight traffic over the Americas. Building on recent progress that quantified the longwave warming contrail effect using advected flight paths as a proxy for contrails, we expand this continuous treatment approach to capture the highly skewed shortwave cooling impact, delivering a 12-hour lifespan net observational radiative forcing. Our analysis reveals a statistically significant net warming energy forcing of 33.7 (95% CI: 20.8, 47.8) GJ/km flown from April 2019 to April 2020, providing a large-scale empirical quantification of long contrail lifespan impact of the same order as, though somewhat larger than, previous bottom-up simulation estimates. This observational benchmark offers an independent line of evidence on the sign and magnitude of the climate impact of contrails. View details
Preview abstract Modern user interfaces are complex composites, with elements originating from various sources, such as the operating system, apps, a web browser, or websites. Many security and privacy models implicitly depend on users correctly identifying an element's source, a concept we term ''surface attribution.'' Through two large-scale vignette-based surveys (N=4,400 and N=3,057), we present the first empirical measurement of this ability. We find that users struggle, correctly attributing UI source only 55% of the time on desktop and 53% on mobile. Familiarity and strong brand cues significantly improve accuracy, whereas UI positioning, a long-held security design concept especially for browsers, has minimal impact. Furthermore, simply adding a ''Security & Privacy'' brand cue to Android permission prompts failed to improve attribution. These findings demonstrate a fundamental gap in users' mental models, indicating that relying on them to distinguish trusted UI is a fragile security paradigm. View details
From Correctness to Collaboration: A Human-Centered Taxonomy of AI Agent Behavior in Software Engineering
Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA ’26), ACM, New York, NY, USA (2026)
Preview abstract The ongoing transition of Large Language Models in software engineering from code generators into autonomous agents requires a shift in how we define and measure success. While models are becoming more capable, the industry lacks a clear understanding of the behavioral norms that make an agent effective in collaborative software development in the enterprise. This work addresses this gap by presenting a taxonomy of desirable agent behaviors, synthesized from 91 sets of user-defined rules for coding agents. We identify four core expectations: Adhere to Standards and Processes, Ensure Code Quality and Reliability, Solve Problems Effectively, and Collaborate with the User. These findings offer a concrete vocabulary for agent behavior, enabling researchers to move beyond correctness-only benchmarks and design evaluations that reflect the realities of professional software development in large enterprises. View details
Preview abstract Probabilistic forecasting of infectious diseases is crucial for public health but relies on labor-intensive manual curation by expert modeling teams. This bespoke development bottlenecks scalability to granular geographic resolutions or emerging pathogens. Here, we present an autonomous system utilizing Large Language Model (LLM)-guided tree search \cite{aygun_ai_2025} to iteratively generate, evaluate, and optimize executable forecasting software. In a fully prospective, real-time evaluation during the 2025–2026 US respiratory season, the system autonomously discovered methodologically diverse models for influenza, COVID-19, and respiratory syncytial virus (RSV). Aggregating these machine-generated models yielded an ensemble that consistently matched or outperformed the gold-standard, human-curated Centers for Disease Control and Prevention (CDC) hub ensembles out-of-sample. The system successfully navigated data-scarce "cold start" scenarios for RSV. Moreover, controlled ablations revealed that optimizing log-scale distance metrics prevents reward hacking, while an automated judge-in-the-loop ensures structural fidelity to complex scientific theories. By autonomously translating epidemiological theory into accurate, transparent code, this framework breaks the modeling labor bottleneck, enabling the rapid deployment of expert-level disease forecasting at unprecedented scales. View details
Preview abstract We introduce a new context-enriched time series forecasting benchmark TimesX. TimesX contains a wide selection of high-quality real-world time series and diverse textual contexts from an automated generating pipeline, which helps address three main issues of existing benchmarks: (1) poor generalization due to low data volume and data being synthetic, (2) restricted forms of context, and (3) an inability to mitigate data leakage. We conduct a thorough empirical study of current multimodal solutions on TimesX. Our results suggest that most multimodal solutions that work well on existing benchmarks may fail on TimesX. In contrast, simple ensemble methods that leverage the rich textual context can outperform strong unimodal baselines and other multimodal baselines. ** Below this is what was submitted to ITP. ** We create a real world multimodal time-series forecasting benchmark that encompasses diverse domains and regions. Each time-series is annotated by various kinds of contexts like metadata, date and holiday information, dynamic events related to the time-series. This is sufficiently more advanced than other available benchmarks which rely wither on static metadata alone or synthetic examples. This forms a test bed for multimodal forecasting. We also present some baseline results showing that ensembles of publicly available LLMs and time-series foundation models can demonstrate non-trivial performance on this bechmark. View details
Preview abstract System coherency verification is vital for ensuring data consistency in complex memory hierarchies, but late integration often delays bug discovery. This paper presents a "left-shift" approach to accelerate coherency verification. We detail three key aspects: early verification via a stitched DUT for initial testing; automated stimulus generation using third-party tools like Cadence Perspec to cover complex system-level scenarios; and automated checker generation, complemented by an in-house tool (DICE). This methodology significantly reduces test/checker development time, enables faster test creation for corner cases, and results in better system-level coherency coverage, finding critical bugs earlier in the design cycle. View details
Preview abstract Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and dynamic environments, but this also introduces severe security risks. In particular, indirect prompt injection attacks can compromise agents through malicious instructions hidden in external sources such as web pages, emails, and retrieved documents. Existing defenses are largely reactive, while current automated red-teaming methods mainly optimize attack success rather than systematically uncovering hidden vulnerabilities within the agent pipeline. In this work, we propose PI-Hunter, an automated agentic red-teaming framework that shifts the focus from attack optimization to vulnerability exposure. By combining static attack-surface analysis, source-aware seeding, trajectory evaluation, and feedback-guided exploration, PI-Hunter proactively discovers vulnerable ingestion paths and localizes how malicious instructions propagate through agent reasoning. Extensive experiments across multiple benchmarks, agent architectures, attacks, and defenses show that \method~substantially improves vulnerability exposure and attack-surface coverage compared with existing automated red-teaming baselines, while remaining effective even under strong prompt injection defenses. View details
Preview abstract Zero-concentrated differential privacy (zCDP) is a variant of differential privacy (DP) that is widely used partly thanks to its nice composition property. While a tight conversion from ε-DP to zCDP exists for the worst-case mechanism, many common algorithms satisfy stronger guarantees. In this work, we derive tight zCDP characterizations for several fundamental mechanisms. We prove that the tight zCDP bound for the ε-DP Laplace mechanism is exactly (ε + e^{−ε} − 1), confirming a recent conjecture by Wang [Wan22]. We further provide tight bounds for the discrete Laplace mechanism, k-Randomized Response (for k ≤ 6), and RAPPOR. Lastly, we also provide a tight zCDP bound for the worst case bounded range mechanism. View details
Efficacy of Scalable Airline-led Contrail Avoidance
Thomas Dean
Tristan Abbott
Jill Blickstein
Alejandra Martín Frías
Mark Galyen
Rebecca Grenham
Paul Hodgson
Alan Pechman
Tyler Robarge
Dinesh Sanekommu
Aarón Sonabend
Marc E.J. Stettler
Raimund Zopp
Journal of Environmentally Compatible Air Transport System (JECATS) (2026) (to appear)
Preview abstract Contrails account for a large portion of aviation's contribution to anthropogenic climate change. Navigational contrail avoidance is a promising solution to mitigate the warming caused by contrails. Prior trials testing navigational contrail avoidance have relied on bespoke integrations of contrail forecasts into airline operations. Here, we use a randomized control trial to test the feasibility of dispatcher-led contrail avoidance integrated into standard flight planning operations using a workflow which scales to an airline's entire network. Using satellite imagery and an automated flight-contrail attribution algorithm, we observed an 11.6% reduction in contrail formation rate for the 1232 flights marked as eligible for contrail avoidance (intent-to-treat) relative to the flights in the control group (p = 0.0109). In the 112 flights which flew contrail avoidance as planned (per-protocol flights), we observed a 62.0% lower contrail formation rate relative to the flights in the control group (p < 0.001). No statistically significant difference in fuel usage was observed between the two groups. View details
Preview abstract This framework manages AI agents by establishing behavioral boundaries and a persistent identity. It uses a multi-layered stack, combining safety rules with brand guidelines, to shape an agent's reasoning. Features include authority decay to limit power if confidence drops and memory segmentation to prevent data tampering. Centralized oversight ensures these digital representatives remain aligned with company policies through continuous monitoring and testing. View details
×