Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11583 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
MAGE: Modality-Agnostic Music Generation and Editing
Muhammad Usama Saleem
Ravi Tejasvi
Rajeev Nongpiur
Mayur Jagdishbhai Patel
Pu Wang
2026
Preview abstract Multimodal music creation requires models that can both generate audio from high-level cues and edit existing mixtures in a targeted manner. Yet most multimodal music systems are built for a single task and a fixed prompting interface, making their conditioning brittle when guidance is ambiguous, temporally misaligned, or partially missing. Common additive fusion or feature concatenation further weakens cross-modal grounding, often causing prompt drift and spurious musical content during generation and editing. We propose MAGE, a modality-agnostic framework that unifies multimodal music generation and mixture-grounded editing within a single continuous latent formulation. At its core, MAGE uses a Controlled Multimodal FluxFormer, a flow-based Transformer that learns controllable latent trajectories for synthesis and editing under any available subset of conditions. To improve grounding, we introduce Audio-Visual Nexus Alignment to select temporally consistent visual evidence for the audio timeline, and a cross-gated modulation mechanism that applies multiplicative control from aligned visual and textual cues to the audio latents, suppressing unsupported components rather than injecting them. Finally, we train with a dynamic modality-masking curriculum that exposes the model to text-only, visual-only, joint multimodal, and mixture-guided settings, enabling robust inference under missing modalities without training separate models. Experiments on the MUSIC benchmark show that MAGE supports effective multimodal-guided music generation and targeted editing, achieving competitive quality while offering a lightweight and flexible interface tailored to practical music workflows. View details
Preview abstract To meet aggressive time-to-market goals, modern mobile SoC architectures require the concurrent development of custom Compute-IPs and surrounding subsystem integration logic (1PIPs and 3PIPs). However, this parallel execution creates a critical verification deadlock: the subsystem cannot be validated until both the Compute-IP and volatile, in-flight 1PIPs reach physical RTL maturity. Consequently, "integration-killer" bugs—such as protocol handshaking deadlocks and clock/reset sequencing mismatches—remain hidden until late in the design cycle when RTL rework costs are prohibitive. To break this bottleneck, we present a verification-driven methodology utilizing a silicon-proven Golden Proxy, Direct-Execution Traffic Profiles, Programmable Sequencers and Automated Protocol Converters. This framework completely decouples parallel hardware dependencies, pre-pulling critical inter-IP mismatch discoveries months ahead of traditional integration milestones. View details
An Empirical Study of Tablet Ergonomics: The Interplay of Temperature, Orientation, and Use Behaviors
Carmen Van Ommen
Mikki Phan
Arun Raghupathy
Daniel Huynh
Barbara Chaparro
Ergonomics in Design: The Quarterly of Human Factors Applications Journal (2026)
Preview abstract To balance computational performance with thermal comfort, this study explores a consolidated hotspot architecture at the top center of a tablet. We tested hotspot (39°C, 43°C, 45°C, 47°C) and ambient temperatures (25°C, 35°C) with 60 participants, measuring perception, action likelihood, and expectation. The hotspot was observed away from high contact areas, with 43°C identified as the threshold for significant discomfort. Discomfort increased with portrait mode use and higher device and ambient temperatures, while active use duration influenced acceptability. The findings underscore the importance of thermal mapping and contextual sensing, with direct applications for software throttling thresholds of coated aluminum enclosures. View details
Diffusion Controller: Framework, Algorithms and Parameterization
Tong Yang
Moonkyung Ryu
Guy Tennenholtz
Yuejie Chi
Bo Dai
Proceedings of the 43rd International Conference on Machine Learning (ICML-26), Seoul, South Korea (2026)
Preview abstract Controllable generation with diffusion models is often treated as a collection of heuristics rather than a unified optimization problem. We propose a principled control formulation by viewing the diffusion reverse process as an instance of a (generalized) linearly-solvable Markov decision process (LS-MDP). This perspective turns controllable generation into regularized optimal control around a pretrained diffusion policy, yielding tractable objectives and algorithmic updates. Under this framework, we study two practical finetuning regimes. When paired target data are available, we obtain a supervised finetuning (SFT) objective. When only a terminal reward model is available, we derive reinforcement-learning finetuning (RLFT) methods from the LS-MDP solution structure, including (i) a reward-weighted regression loss and (ii) a policy-gradient approach (with standard extensions such as PPO). Crucially, the LS-MDP optimality conditions imply an explicit relationship between the optimal and pretrained score functions. We leverage this to derive a new score-function parameterization that isolates the control signal and enables “gray-box” finetuning with substantially fewer trainable parameters. Experiments across SFT and RLFT show this parameterization improves over existing finetuning baselines while achieving stronger sample/parameter efficiency. View details
ITHICA: Intra-Thread Instruction Checking Approach for Defect-Induced Silent Data Corruptions
Ioanna Vavelidou
Eric Xue Liu
Mike Fuller
Subhasish Mitra
Caroline Trippel
59th IEEE/ACM International Symposium on Microarchitecture (MICRO) (2026)
Preview abstract Hyperscalers are reporting silent data corruptions (SDCs), presumed to be caused by silicon manufacturing defects, as a threat to datacenter reliability. To support datacenter testing efforts to detect defective CPU servers, this paper presentsITHICA, an approach and tool for automatically generating functional tests for defect-induced errors from arbitrary programs by inserting intra-thread, instruction-level error checks, primarily leveraging instruction duplication and output comparison. Our key insight is that the most pernicious defects (those most likely to escape manufacturing testing) cause inconsistent errors: two executions of the same instruction given the same inputs within the same thread can produce different architectural outputs depending on the execution context in which they run. By exploiting this insight, ITHICA uniquely enables arbitrary programs to serve as tests and localizes affected instructions concurrently with error detection. We use ITHICA to transform industrial hyperscaler tests (our baseline), datacenter programs, and common libraries into functional tests, and evaluate them on over 3,000 CPU servers. ITHICA checks detect 39% more defective servers than baseline industrial checks and yield novel findings on defect behavior that challenge conclusions drawn by prior hyperscaler fleet studies. View details
Preview abstract Despite significant strides in factual reliability, errors -- often termed hallucinations -- remain a major concern for generative AI, especially as LLMs are increasingly expected to be helpful in more complex or nuanced setups. Yet even in the simplest setting -- factoid question-answering with clear ground truth-frontier models without external tools continue to hallucinate. We argue that most factuality gains in this domain have come from expanding the model's knowledge boundary (encoding more facts) rather than improving awareness of that boundary (distinguishing known from unknown). We conjecture that the latter is inherently difficult: models may lack the discriminative power to perfectly separate truths from errors, creating an unavoidable tradeoff between eliminating hallucinations and preserving utility. This tradeoff dissolves under a different framing. If we understand hallucinations as confident errors -- incorrect information delivered without appropriate qualification -- a third path emerges beyond the answer-or-abstain dichotomy: expressing uncertainty. We propose faithful uncertainty: aligning linguistic uncertainty with intrinsic uncertainty. This is one facet of metacognition -- the ability to be aware of one's own uncertainty and to act on it. For direct interaction, acting on uncertainty means communicating it honestly; for agentic systems, it becomes the control layer governing when to search and what to trust. Metacognition is thus essential for LLMs to be both trustworthy and capable; we conclude by highlighting open problems for progress towards this objective. View details
Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation
James McClure
José Cambronero
Renyao Wei
Dorothy Chen
Livio Dalloro
SpecOps '26: Proceedings of the 1st International Workshop on Specification-Driven Development Life Cycle, ACM (Association for Computing Machinery), New York, NY, USA (2026)
Preview abstract LLM-based agents are increasingly used for coding tasks, where they have outperformed many classical approaches and scaled to repository-level tasks, such as test generation. However, when directly prompted to generate tests, these agents can fail to reason about the code and its underlying contracts, thereby missing edge cases and behavioral boundaries that affect test quality. To address this limitation, we propose Spec-Driven Test Generation, where we instruct an agent to first reason about – and explicitly document – code pre-conditions, post-conditions, and undefined behaviors. This intermediate semi-formal specification acts as a cognitive scaffold to guide subsequent test generation. Our evaluation on production bugs from Google shows that the spec-driven agent can deliver a 9.8 percentage points (p = 0.0352) improvement in bug detection rate and a 2.5 percentage point (p = 0.0034) improvement in branch coverage, compared to a traditional test generation agent baseline. Using LLM-as-a-Judge, we further show that test suites generated by the spec-driven agent are superior to the baseline and human-authored tests in 77.8% and 56.7% of the cases, respectively, and demonstrated improvements on following best practices, readability, and edge-case coverage. View details
Preview abstract Autonomous research agents can now produce competitive solutions and complete manuscripts, yet their papers routinely contain fabricated citations, method descriptions disconnected from the code, and scores on incorrect scales---failures invisible to evaluations that assess fluency rather than evidentiary grounding. The core problem is verifiability: no existing system maintains a traceable chain from each claim in the paper to its grounding evidence, and current evaluation protocols assess output fluency rather than evidentiary grounding. We address this with Chaine-of-Evidence (CoE), a verifiability standard requiring every claim to trace to its grounding evidence, and instantiate it in Scientist One, an end-to-end research system that maintains evidence chains natively, and CoE Audit, an evaluation protocol with four integrity checks targeting the most damaging chain failures. Auditing 60 papers from four systems, we find every baseline exhibits at least one failure: phantom citations at 4--25%, method-code alignment in at most 2/15 papers. Scientist One achieves zero phantom citations (0/830), the highest alignment rate (7/15), and competitive solver scores. View details
Preview abstract Zero-concentrated differential privacy (zCDP) is a variant of differential privacy (DP) that is widely used partly thanks to its nice composition property. While a tight conversion from ε-DP to zCDP exists for the worst-case mechanism, many common algorithms satisfy stronger guarantees. In this work, we derive tight zCDP characterizations for several fundamental mechanisms. We prove that the tight zCDP bound for the ε-DP Laplace mechanism is exactly (ε + e^{−ε} − 1), confirming a recent conjecture by Wang [Wan22]. We further provide tight bounds for the discrete Laplace mechanism, k-Randomized Response (for k ≤ 6), and RAPPOR. Lastly, we also provide a tight zCDP bound for the worst case bounded range mechanism. View details
Preview abstract Large language model agents increasingly act in deployment environments where failures are contextual, user-specific, and costly. In such settings, a \emph{static general-purpose guardrail is often insufficient}: whether an action should be allowed may depend on local privacy norms, organizational rules, or evolving user expectations that are difficult to enumerate fully in advance. We study \emph{lifelong deployment-time guardrail adaptation}, where a fixed base guardrail improves over time from sparse, noisy user-reported failures without repeated fine-tuning. We propose a conservative policy induction framework organized as an online--offline loop. Online, the deployed guardrail uses structured policy memory to guide runtime decisions. Offline, newly accumulated reports are converted into reusable policy items and folded back into memory through periodic refresh. The method combines three ingredients: \emph{broad policy abstraction} for sparse failure generalization, \emph{conflict-aware local policies} for mixed-label regions where broad reuse becomes too coarse, and \emph{confidence-gated reuse} based on conservative posterior lower bounds so that weakly supported memory does not influence inference too early. Across PrivacyLens+, ConFaide+, and AgentHarm, the resulting system consistently improves over a lightweight base guardrail and strong memory-based baselines in sparse-feedback regimes, remains robust to noisy feedback, traces a better cost--performance frontier than scaling the base model alone, and jointly reduces over-refusal and over-acceptance without an explicit balance knob. View details
Preview abstract While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Aggregation Layer dynamically reconciles these perspectives -- via final candidate selection, semantic synthesis, or neuro-symbolic verification -- to produce a robust global solution. We evaluate PoTRE on three frontier benchmarks: ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance. PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, surpassing the previous best official score. We demonstrate that this architectural heterogeneity achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines. View details
TDXRay: Microarchitectural Side-Channel Analysis of Intel TDX for Real-World Workloads
Tristan Hornetz
Hosein Yavarzadeh
Albert Cheu
Adria Gascon
Lukas Gerlach
Michael Schwarz
Ruiyi Zhang
IEEE Security & Privacy (S&P) (2026)
Preview abstract Confidential computing with VM-based trusted execution environments (TEEs) promises to protect code and data from a privileged cloud operator, enabling privacy-preserving workloads ranging from medical analytics to AI inference. However, most deployments exclude microarchitectural side channels from their threat model, shifting the burden to application developers who lack practical, general-purpose tools to assess (let alone mitigate) leakage. This gap is problematic: host-observable effects such as page-fault patterns, shared-cache contention, performance-counter surrogates (where available), and fine-grained timing primitives (e.g., MWAIT) can still reveal high-level secrets even when memory remains encrypted. We present TDXRay, an open-source framework that systematizes the evaluation of side-channel risk for confidential VMs in Intel TDX. TDXRay exposes unified interfaces to exercise and measure several attack primitives—including controlled-channel attacks via page tables, cache-based contention/occupancy probes, performance-counter–derived signals, and timing channels—against unmodified guest workloads. Using TDXRay, we build two end-to-end case studies: (1) a classic AES T-table attack in which a malicious hypervisor recovers the secret key from access-pattern leakage, and (2) an LLaMA inference attack in which the host infers user prompts by monitoring memory accesses during tokenization and embedding lookups. Across both, we show that a host with no direct access to guest memory can reconstruct sensitive information by observing only externalized microarchitectural signals. View details
Preview abstract Image-based sexual abuse (IBSA) refers to the creation, sharing, or threats to share intimate images, whether real or synthetic. Advances in artificial intelligence (AI) now enable the generation of sexually explicit images from non-explicit source material, such as professional headshots and social media profiles. In this study, we found that across three regions (Australia, the United Kingdom, and the United States), 15% of a representative sample reported experiences of IBSA involving digitally altered images, including 6.9% who reported AI-generated IBSA (AI-IBSA) victimization specifically. Higher rates of victimization were observed among respondents under 35, LGBTQ+, and BIPOC participants, people with disabilities, and men. While victimization rates were higher among men, women were more likely to report a broader range of harms and more negative emotional and social impacts. The most common responses to victimization were platform-based actions, including blocking (69%) and reporting (58%). These findings highlight AI-IBSA as a growing and under-recognized form of harm, with important implications for prevention, platform governance, and support for victim-survivors. View details
×