Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11600 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract We study a quantized prefix estimator for inner products that turns a randomly rotated TurboQuant-style representation into a cheap Johnson–Lindenstrauss-like search signal. The idea is simple: rotate the vectors once, keep only a short prefix of coordinates for fast scoring, and quantize the database-side prefix with an unbiased scalar quantizer. We prove that this estimator is unbiased and that its error separates cleanly into two interpretable sources: prefix truncation from using only r coordinates, and quantization error from using b bits per coordinate This separation is useful in systems because the prefix can be exposed as a lightweight filter without building a separate projection index. In ParlayANN graph search, a 64-coordinate truncated view of existing TQ4 codes can replace a separately stored JL256 filter before full-precision reranking, adding only prefix-scale and query lookup-table bookkeeping. In k-means, the same estimator accelerates the dominant point–centroid assignment kernel while preserving exact centroid norms. Empirically, the truncated-TQ filter tracks the JL recall–throughput frontier across five graph-search datasets while reusing the quantized representation already present in the index. View details
Preview abstract Deep-learning methods have boosted the analytical power of Raman spectroscopy, yet they still require large, task-specific, labeled datasets and often fail to transfer across application domains. The study explores pre-trained encoders as a solution. Pre-trained encoders have significantly impacted Natural Language Processing and Computer Vision with their ability to learn transferable representations that can be applied to a variety of datasets, significantly reducing the amount of time and data required to create capable models. The following work puts forward a new approach that applies these benefits to Raman Spectroscopy. The proposed approach, RSPTE (Raman Spectroscopy Pre-Trained Encoder), is designed to learn generalizable spectral representations without labels. RSPTE employs a novel domain adaptation strategy using unsupervised Barlow Twins decorrelation objectives to learn fundamental spectral patterns from multi-domain Raman Spectroscopy datasets containing samples from medicine, biology, and mineralogy. Transferability is demonstrated through evaluation on several models created by fine-tuning RSPTE for different application domains: Medicine (detection of Melanoma and COVID), Biology (Pathogen Identification), and Agriculture. As an example, using only 20% of the dataset, models trained with RSPTE achieve accuracies ranging 50%–86% (depending on the dataset used) while without RSPTE the range is 9%–57%. Using the full dataset, accuracies with RSPTE range 81%–97%, and without pretraining 51%–97%. Current methods and state-of-the-art models in Raman Spectroscopy are compared to RSPTE for context, and RSPTE exhibits competitive results, especially with less data as well. These results provide evidence that the proposed RSPTE model can effectively learn and transfer generalizable spectral features across different domains, achieving accurate results with less data in less time (both data collection time and training time). View details
FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
Victor May
Diganta Misra
Yanqi Luo
Anjali Sridhar
Justine Gehring
Silvio Soares Ribeiro Junior
2026
Preview abstract AI coding assistants are rapidly becoming integral to modern software development. A key challenge in this space is the continual need to migrate and modernize codebases in response to evolving software ecosystems. Traditionally, such migrations have relied on rule-based systems and human intervention. With the advent of powerful large language models (LLMs), AI-driven agentic frameworks offer a promising alternative—but their effectiveness remains underexplored. In this paper, we introduce FreshBrew, a novel benchmark for evaluating AI-based agentic frameworks on project-level Java migrations. We benchmark several such frameworks, powered by state-of-the-art LLMs, and compare their performance against established rule-based tools. Our evaluation of AI agents on this benchmark of 228 repositories shows that the top-performing model, Gemini 2.5 Flash, can successfully migrate 56.5% of projects to JDK 17. Our empirical analysis reveals novel insights into the critical strengths and limitations of current agentic approaches, offering actionable insights into their real-world applicability. By releasing FreshBrew publicly upon acceptance, we aim to facilitate rigorous, reproducible evaluation and catalyze progress in AI-driven codebase modernization. View details
Preview abstract This paper provides an overview of Google's TPUs across five generations, from TPU v2 to Ironwood, highlighting their evolution as scalable, resilient, and sustainable supercomputers for AI training. It details the TPU’s stable architecture and microarchitecture, which has surprisingly easily accommodated the rapidly changing deep neural network workloads, such as the rise of Transformers. Key advancements over eight years include 10x increase in HBM capacity and bandwidth per node, a 100x increase in peak node performance, and a 3600x increase in supercomputer performance. The paper also discusses the role of optical circuit switches and built-in self test in enhancing resilience, how TPU’s carbon footprint was reduced by improving embodied carbon emissions per floating point operation and a 30x gain in performance per Watt. It concludes by identifying six features that may well characterize the successful AI accelerators of this decade. View details
Preview abstract The rapid adoption of agentic systems powered by large language models (LLMs) introduces significant security challenges distinct from plain conversational models, particularly concerning prompt injection and tool misuse due to their dynamic personas and real- world tool interactions. This paper investigates the effectiveness of hardened security prompting in a task-oriented multi-agent framework, using a coding assistant as a representative case study. We com- pare a baseline ”unhardened” agent against a ”hard- ened” version equipped with explicit security guide- lines applied across all sub-agents. Our evaluation across 150+ single-turn and 32 multi-turn attack sce- narios demonstrates that prompt hardening dramat- ically improves resilience. With a simple, approxi- mately 500-token security hardener, single-turn fail- ure rates dropped from 19.48% to 2.60%, while multi- turn failure rates decreased from 75.00% to 46.88%. Furthermore, we show that successfully bypassing the hardened agent requires significantly more adversar- ial effort and a greater number of chat turns. How- ever, the analysis also reveals a critical shift in vul- nerability taxonomy: as direct attacks fail, adver- saries exploit the agent’s core functionality via ”Func- tional Wrappers” (Intent Obfuscation), highlighting a residual risk that necessitates a shift in the defen- sive paradigm from static filters to dynamic runtime state and intent analysis. View details
Preview abstract Large-scale software systems frequently suffer from architectural rigidity caused by monolithic designs, tightly coupled integrations, and legacy technology stacks. Backend-forFrontend (BFF) architectures are increasingly adopted to address these challenges by decoupling frontend-specific requirements from backend domain services. However, designing a BFF layer requires a series of irreversible technology decisions across compute platforms, traffic routing, programming languages, frameworks, and API protocols. These decisions directly influence system latency, scalability, operational complexity, and long-term maintainability. This paper proposes a structured, metrics-driven decision framework to guide architects through foundational technology choices when designing BFF architectures. The framework decomposes the decision space into independent sub-problems, introduces weighted evaluation criteria, and applies quantitative scoring models to enable objective trade-off analysis. The approach is validated through a representative modernization scenario, demonstrating how systematic evaluation reduces architectural risk, resolves stakeholder disagreement, and improves performance and developer efficiency. The proposed framework is generic, repeatable, and applicable to a wide range of cloudnative system modernization efforts. View details
Preview abstract Prior work synthesizes tool-use LLM datasets by first generating a user query, followed by complex tool-use annotations like depth-first search (DFS). This leads to inevitable annotation failures and low efficiency in data generation. We introduce ToolGrad, an agentic framework that inverts this paradigm. ToolGrad first constructs valid tool-use chains through an iterative process guided by textual "gradients", and then synthesizes corresponding user queries. This "answer-first" approach led to ToolGrad-500, a dataset generated with more complex tool use, lower cost, and almost 100% pass rate. Experiments show that ToolGrad models outperform those trained on expensive baseline datasets and proprietary LLMs. View details
Preview abstract Image-based sexual abuse (IBSA) refers to the creation, sharing, or threats to share intimate images, whether real or synthetic. Advances in artificial intelligence (AI) now enable the generation of sexually explicit images from non-explicit source material, such as professional headshots and social media profiles. In this study, we found that across three regions (Australia, the United Kingdom, and the United States), 15% of a representative sample reported experiences of IBSA involving digitally altered images, including 6.9% who reported AI-generated IBSA (AI-IBSA) victimization specifically. Higher rates of victimization were observed among respondents under 35, LGBTQ+, and BIPOC participants, people with disabilities, and men. While victimization rates were higher among men, women were more likely to report a broader range of harms and more negative emotional and social impacts. The most common responses to victimization were platform-based actions, including blocking (69%) and reporting (58%). These findings highlight AI-IBSA as a growing and under-recognized form of harm, with important implications for prevention, platform governance, and support for victim-survivors. View details
Preview abstract While the Latin script is used informally by speakers of many languages with more complex native scripts, high quality Latin script corpora for such languages that reflect actual natural romanizations are scarce and often difficult to collect. In this work, we propose a method for mining romanized language corpora in languages for which we do not have any pre-existing samples of naturally romanized text, focusing on Tigrinya as a test case. First we examine the efficacy of learning romanizations for a language based on observed romanizations in other languages that use the same native script. We then extrinsically assess such methods by using a romanization model trained on Amharic data to bootstrap coverage of romanized Tigrinya in a language identification system. Manual evaluation by two L1 and one L2 Tigrinya speakers suggests our method extracts romanized Tigrinya text with acceptably high precision. We release code to run our mining pipeline on public web corpora, such as MADLAD-400. View details
DDRop: Generic Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes
Jesse Demeulemeester
Stefan Gloor
Patrick Jattke
David Oswald
Martin Thompson
Kaveh Razavi
Ingrid Verbauwhede
Jo Van Bulck
ACM Conference on Computer and Communications Security (CCS) (2026)
Preview abstract Trusted Execution Environments (TEEs) are increasingly deployed in the cloud to protect sensitive workloads through hardwareenforced isolation, remote attestation, and transparent memory encryption. However, to meet memory performance and size demands, modern TEEs omit cryptographic freshness guarantees, leaving them vulnerable to replay attacks by adversaries with physical memory access. Prior work demonstrated low-cost active interposition attacks on DDR4 without requiring expensive specialized equipment, but these techniques do not extend to DDR5, where existing approaches are limited to passive ciphertext side-channel analysis and rely on bus downclocking to accommodate legacy memory bus analyzers. We present the first low-cost (<200$) DDR5 RDIMM interposer capable of active fault injection at native speeds. By injecting targeted parity errors to silently discard cache line writebacks, we introduce DDRop, a new primitive that exploits the absence of cryptographic freshness to break the integrity of Intel TDX, Scalable SGX, and AMD SEV-SNP. Building on this primitive and targeting the APIs exposed by the TDX module and AMD Secure Processor, we show that adversaries can gain ciphertext access, copy arbitrary victim pages, and inject malicious secure page-table entries. We demonstrate end-to-end attacks on an up-to-date TDX platform, including forcing any TD into debug mode and forging attestation reports. While software-level mitigations, including timing-based interposer detection and API hardening, may reduce the attack surface, our results demonstrate that active DDR5 bus interposition is practical at low cost, highlighting the need for robust cryptographic memory integrity protections against physical adversaries. View details
SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao
Gerardo Vitagliano
Immanuel Trummer
H. V. Jagadish
Sebastian Schelter
Andreas Kipf
Matthew Russo
Kris Kissel
Michael Cochez
Andreas Zimmerer
Olga Ovcharenko
Thibaud Hottelier
Gautam Gupta
Tianji Cong
2026
Preview abstract We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on zero-shot abilities of state-of-the-art large language models (LLMs). They extend SQL with semantic operators, configured by natural language instructions, that are evaluated via LLMs and enable users to perform various operations on multimodal data. Our benchmark provides variety along three axis: scenarios, modalities, and operators. Included are scenarios ranging from movie review analysis to medical question-answering. Within these scenarios, we cover different data modalities, including images, audio, and text. Finally, the queries involve a diverse set of operators, including semantic filters, joins, mappings, ranking, and classification operators. We evaluate systems according to processing overheads and result quality. We present experimental results for an industrial semantic query processing engine (BigQuery), as well as academic systems (LOTUS, Palimpzest, and ThalamusDB). Our results shed light on the relative strengths and weaknesses of the evaluated systems, and hint at promising avenues for future research. View details
When “Correct” Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
Corina Pasareanu
Haizhong Zheng
Beidi Chen
Ravi Mangal
Xinyu Yang
James Song
Yibo Peng
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, San Diego, California, United States (2026), 15514–15546
Preview abstract Code agents are increasingly trusted to autonomously fix bugs on platforms such as GitHub, yet their security evaluation focuses almost exclusively on functional correctness. In this paper, we reveal a novel type of threat to real-world code agents: Functionally Correct yet Vulnerable (FCV) patches, which pass all test cases but contain vulnerable code. With our proposed FCV-Attack, which can be deliberately crafted by malicious attackers or implicitly introduced by benign developers, we show that SOTA LLMs (e.g., ChatGPT and Claude) and agent scaffolds (e.g., SWE-agent and OpenHands) are all vulnerable to this FCV threat; across 12 agent-model combinations on SWE-Bench, the attack only requires black-box access and a single query to the code agent to perform the attack. For example, for CWE-538 (information exposure vulnerability), the FCV-Attack attains an attack success rate of 40.7% on GPT-5 Mini + OpenHands. Our results reveal an important security threat overlooked by current evaluation paradigms and urge the development of security-aware defenses for code agents. View details
Beyond Tsybakov: Model Margin Noise and H-Consistency Bounds
The Nineteenth International Symposium on Artificial Intelligence and Mathematics (ISAIM 2026)
Preview abstract We introduce a new low-noise condition for classification, the *Model Margin Noise (MM noise)* assumption, and derive enhanced $H$-consistency bounds under this condition. MM noise is *weaker* than Tsybakov noise condition: it is implied by Tsybakov noise condition but can hold even when Tsybakov fails, because it depends on the discrepancy between a given hypothesis and the Bayes-classifier rather than on the intrinsic distributional minimal margin (see Figure 1 for an illustration of an explicit example). This hypothesis-dependent assumption yields enhanced $H$-consistency bounds for both binary and multi-class classification. Our results extend the enhanced $H$-consistency bounds of Mao, Mohri, and Zhong (2025a) with the same favorable exponents but under a weaker assumption than the Tsybakov noise condition; they interpolate smoothly between linear and square-root regimes for intermediate noise levels. We also instantiate these bounds for common surrogate loss families and provide illustrative tables. View details
×