Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11612 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract We study a quantized prefix estimator for inner products that turns a randomly rotated TurboQuant-style representation into a cheap Johnson–Lindenstrauss-like search signal. The idea is simple: rotate the vectors once, keep only a short prefix of coordinates for fast scoring, and quantize the database-side prefix with an unbiased scalar quantizer. We prove that this estimator is unbiased and that its error separates cleanly into two interpretable sources: prefix truncation from using only r coordinates, and quantization error from using b bits per coordinate This separation is useful in systems because the prefix can be exposed as a lightweight filter without building a separate projection index. In ParlayANN graph search, a 64-coordinate truncated view of existing TQ4 codes can replace a separately stored JL256 filter before full-precision reranking, adding only prefix-scale and query lookup-table bookkeeping. In k-means, the same estimator accelerates the dominant point–centroid assignment kernel while preserving exact centroid norms. Empirically, the truncated-TQ filter tracks the JL recall–throughput frontier across five graph-search datasets while reusing the quantized representation already present in the index. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract Optical health sensing algorithms, such as SpO2, sleep monitoring, and metabolic health sensing, critically depend on the accurate measurement of optical emission from Light Emitting Diodes (LEDs) transmitted through user tissue and detected by a photodiode (PD). A significant challenge to the reliability of these measurements is the inherent degradation of LED optical emission intensity over time due to device aging. This degradation can confound the physiological changes being monitored. Our work quantifies the impact of LED aging on sensor signal integrity, specifically examining the Current Transfer Ratio (CTR), which is a key metric defining the ratio of received photocurrent to the LED drive current used for transmission in various health sensing algorithms. We investigate the degradation characteristics across LEDs of different wavelengths. Our findings indicate a relative CTR change due to degradation ranging from 1% to 8% within 100 hours of continuous operation which translates to approximately 3.5 to 7 years of device lifetime. Furthermore, we explore the non-linearity of this degradation and the observed initial ”overshoot” phenomenon in the CTR during aging. We discuss how understanding these dynamics could inform the development of robust specifications for different physiological sensing algorithms. Finally, we present several potential solutions to mitigate the effects of LED aging. During the product design phase, integrating a calibrating photodiode or compensating circuitry around the LED can help preemptively address degradation. In the application space, run-time calibration strategies employing two differently degraded optical paths offer a promising approach to maintain measurement accuracy. View details
SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao
Gerardo Vitagliano
Immanuel Trummer
H. V. Jagadish
Sebastian Schelter
Andreas Kipf
Matthew Russo
Kris Kissel
Michael Cochez
Andreas Zimmerer
Olga Ovcharenko
Thibaud Hottelier
Gautam Gupta
Tianji Cong
2026
Preview abstract We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on zero-shot abilities of state-of-the-art large language models (LLMs). They extend SQL with semantic operators, configured by natural language instructions, that are evaluated via LLMs and enable users to perform various operations on multimodal data. Our benchmark provides variety along three axis: scenarios, modalities, and operators. Included are scenarios ranging from movie review analysis to medical question-answering. Within these scenarios, we cover different data modalities, including images, audio, and text. Finally, the queries involve a diverse set of operators, including semantic filters, joins, mappings, ranking, and classification operators. We evaluate systems according to processing overheads and result quality. We present experimental results for an industrial semantic query processing engine (BigQuery), as well as academic systems (LOTUS, Palimpzest, and ThalamusDB). Our results shed light on the relative strengths and weaknesses of the evaluated systems, and hint at promising avenues for future research. View details
MOSAIC-GS: MOnocular Scene Reconstruction via Advanced Initialization for Complex Dynamic Environments
Svitlana Morkva
Max Wilder-Smith
Michael Oechsle
Marco Hutter
Vaishakh Patil
2026
Preview abstract We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos. Monocular reconstruction is inherently ill-posed due to the absence of sufficient multiview constraints, making accurate recovery of object geometry and temporal coherence particularly challenging. To address this, we leverage multiple geometric cues, such as depth, optical flow, dynamic object segmentation, and tracking trajectories, combined with rigidity constraints to estimate preliminary 3D scene dynamics during an advanced initialization stage. Recovering scene flow prior to the photometric optimization phase reduces reliance on motion inference from visual appearance alone, which is often ambiguous in monocular settings. To enable compact representations, fast training, and real-time rendering while supporting non-rigid deformations, the scene is decomposed into static and dynamic components, with dynamic trajectories represented as time-dependent Poly-Fourier curves for parameter-efficient motion encoding. We demonstrate that MOSAIC-GS achieves substantially faster optimization and rendering compared to existing methods, while maintaining reconstruction quality on par with state-of-the-art approaches across standard monocular dynamic scene benchmarks. View details
Preview abstract Enterprise service delivery platforms, while vital for HR operations, create significant challenges in managing the risks of Personally Identifiable Information (PII) exposure. The integration of Generative AI offers new efficiencies but also amplifies these risks. Existing solutions—ranging from manual redaction and rule-based Data Loss Prevention (DLP) to inflexible data masking—fail to provide a nuanced, integrated approach. This paper introduces the Dual-Mode Privacy Guard (DMPG), a conceptual framework that establishes a model for Augmented Compliance. The framework provides a "defense-in-depth" strategy built on three pillars: (1) a Zero-Trust AI Foundation leveraging a verifiable, non-retention API gateway to ensure data privacy; (2) a proactive "Guardrail" that uses AI to detect and flag potential PII for human-in-the-loop review; and (3) an on-demand "Tool" that allows users to create securely anonymized data assets. By differentiating between proactive monitoring and reactive utility, the DMPG shifts the compliance paradigm from a manual burden to an AI-assisted process that enhances, rather than replaces, human oversight. This paper details the framework’s platform-agnostic architecture, using Salesforce as a reference implementation, and argues for its novelty as a model for operationalizing privacy principles within modern enterprise systems. View details
Preview abstract The Private-Use Area (PUA) is an important part of the Unicode standard. It consists of several ranges of Unicode code points with no official character assignments. The PUA is primarily used as a temporary representation mechanism for characters outside the official standard to facilitate text entry and display of orthographies that cannot be adequately represented by other means. The primary downside of PUA is that characters lose their semantics if the pairing with the corresponding display font is broken. Consequently, they cannot be faithfully displayed in the general setting. Large-scale multilingual web corpora inevitably contain PUA code points of unclear provenance. We investigate the distribution of PUA characters within large-scale datasets, using filters for determining PUA tokens of linguistic interest. We analyze the resulting distributions both across scripts and writing systems, and show that PUA-bearing tokens can signal texts from under-represented languages. We explore whether an off-the-shelf large language model (LLM) can classify PUA characters as those that constitute relevant orthographic signals vs. punctuation or other noise. While the proportion of PUA-bearing paragraphs in the original corpora are small, we identify millions of paragraphs, and we argue that such data is still important for the long tail of data-scarce orthographies. Moreover, as a primary Unicode mechanism for poorly represented writing systems, the PUA is here to stay. View details
Preview abstract Regular-polygon geometry is tightly linked to cyclotomic arithmetic: Poonen and Rubinstein’s treatment of three-diagonal concurrence, for example, turns a geometric incidence condition into a short vanishing sum of roots of unity. We prove an analogous rigidity result for areas. Two congruent crossing diagonals divide a regular n-gon into four regions. For the four regions cut out by the two congruent crossing diagonals V0Vm and VkVn−m+k of a regular n-gon, we completely classify, for all parameters (n, k, m), which sums of the normalized areas a0, . . . , a3 are rational. The classification has a sharp finite–infinite contrast: a0 is rational in only five configurations, whereas the rational cases for a2 and adjacent two-region sums form infinite families. Rationality is delicately sensitive to the parameters: for the configuration (14, 3, 5), no nontrivial subset sum is rational, while the neighboring cut (14, 4, 5) gives a2 = 5/7. The proof reduces each rationality condition to trigonometric relations at rational multiples of π and combines cyclotomic norm arguments with the classification theorems of Conway–Jones and Poonen–Rubinstein. View details
Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion
Patrick Jiang
Judith Li
Moonkyung Ryu
Lily Hu
Kun Su
Liam Hebert
Hao Peng
Jiawei Han
Dima Kuzmin
Proceedings of the 43rd International Conference on Machine Learning (ICML-26), Seoul, South Korea (2026)
Preview abstract Many modern retrieval problems are set-valued: given a broad intent, the system must return a collection of results that optimizes higher-order properties (e.g., diversity, coverage, complementarity, coherence) while staying grounded to a fixed database. These objectives are inherently non-decomposable, creating a training bottleneck because property-aligned (query, content) supervision is scarce. Reinforcement learning (RL) can optimize set-level objectives via interaction, but deploying an RL-tuned LLM for fan-out retrieval is expensive at query time. Diffusion-based generative retrieval enables efficient single-pass fan-out in embedding space, but requires objective-aligned training targets. We propose R4T (Retrieve-for-Train), which uses RL once as an objective transducer: (i) train a fan-out LLM with composite set-level rewards, (ii) synthesize objective-consistent training pairs, and (iii) train a lightweight diffusion retriever to model the conditional distribution of set-valued outputs. Across Polyvore and a large-scale music playlist dataset, R4T improves retrieval quality over strong baselines while reducing query-time fan-out latency by an order of magnitude. View details
Toward a Theory of Value in AI Alignment
Shazeda Ahmed
Abeba Birhane
Jackie Kay
Kris Shrishak
2026
Preview abstract Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a commensurate rise in ways these models cause harm, spanning areas from toxic speech and hallucinations to AI agents executing unauthorized actions. Given that these models are probabilistic and general-purpose by nature, it is impossible to enumerate all possible uses and outputs of the model to reach a fully aligned end state. Within the field of AI safety, these harmful instances are often framed as “the alignment problem,” of models being “misaligned” with human values. Researchers have responded by pursuing applied and theoretical AI “value alignment” efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? The study of human values has long been part of many academic disciplines outside of computer science, yet these disciplines are seldom consulted in AI alignment. Building on the theoretical insights of Zhi-Xuan’s (2024) "preferentist paradigm" critique, we conduct a review of influential AI alignment literature. We also draw from conceptions of human values from philosophy, anthropology, and sociology, to create an analytical schema. We annotated 94 AI value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on “preferences” as a stand-in that runs the risk of reducing complex, culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and LLM-as-a-judge approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. Overall, value alignment is often reduced to an exercise in utility maximization, which we argue abstracts human values away from their lived context. In making AI value alignment’s philosophical commitments explicit, we seek to bring greater specificity and under-explored perspectives into the debate on whether and how AI can address human values View details
Preview abstract The rapid expansion of the Internet of Things (IoT) and smart home ecosystems has led to a fragmented landscape of user data management across consumer electronics (CE) such as Smart TVs, gaming consoles, and set-top boxes. Current onboarding processes on these devices are characterized by high friction due to manual data entry and opaque data-sharing practices. This paper introduces the User Data Sharing System (UDSS), a platform-agnostic framework designed to facilitate secure, privacy-first PII (Personally Identifiable Information) exchange between device platforms and third-party applications. Our system implements a Contextual Scope Enforcement (CSE) mechanism that programmatically restricts data exposure based on user intent—specifically distinguishing between Sign-In and Sign-Up workflows. Unlike cloud-anchored identity standards such as FIDO2/WebAuthn, UDSS is designed for shared, device-centric CE environments where persistent user-to-device bind-ing cannot be assumed. We further propose a tiered access model that balances developer needs with regulatory compliance (GDPR/CCPA). A proof-of-concept implementation on a reference ARMv8 Linux-based middleware demonstrates that UDSS reduces user onboarding latency by 65% and measurably reduces PII over-exposure risk through protocol-enforced data minimization. This framework provides a standardized approach to identity management in the heterogeneous CE market. View details
Preview abstract This paper proposes RD-LoRA, a rate-distortion (R-D) optimized low-rank adaptation (LoRA) framework for neural post-filtering in the upcoming AV2 coding standard. RD-LoRA adapts a pre-trained base neural model to diverse input content via online updating of LoRA parameters, whose quantity is governed by the matrix ranks. The updated parameters are then quantized, transmitted to the decoder, and merged with the pre-trained weights for post-filtering. While more parameters generally improve coding performance, they also increase transmission bitrate. To balance distortion reduction against transmission cost, we propose to dynamically allocate rank budget to each layer of the base model in a closed-loop R-D manner. Specifically, it incorporates a cost-aware importance assessment to discourage parameter-heavy updates, together with an RD-rank allocator to guide pruning based on global R-D optimization. To maintain robustness when certain layers are pruned to zero rank, we further introduce a lightweight fallback modulation mechanism. Experimental results show that, when deployed on a lightweight 25 kMACs ResNet model, RD-LoRA achieves a BD-rate reduction of 2.556% over the AV2 anchor, significantly outperforming the base model with only limited decoding overhead. View details
Preview abstract Human-Computer Interaction research and design pedagogy rely on idealized process models, such as the Double Diamond, to describe how user experiences are designed. These models assume an orderly, linear design process that, while easy to understand, fails to capture the iterative and collaborative reality of professional practice. A few qualitative studies have successfully captured this complexity -- still, they often suffer from retrospective narrative smoothing and lack systemic scale. To understand how design unfolds in real products, we analyzed historical snapshots of 102 Figma files from a multi-national technology company and investigated the true trajectories of the design process at scale. Our analysis reveals that while the established process models might be applicable, the operational details are highly non-linear. Rather than a straight line from ideation toward completion, design advances are repeatedly reset to the ideation stage as feedback is received. We argue that by treating design files as operational telemetry, the industry can move beyond abstract frameworks to build practices and collaborative tools that support the non-linear realities of modern product development. View details
Preview abstract in many large language model (LLM) applications, a serialized prompt contains both a user request and external records such as webpages, email, memory, or tool outputs. The same embedded directive may need to be applied for one task and treated as data for another. IBBench-Light tests this contrast with 12 semantic tasks rendered through four source wrappers and three embedding forms. The construction yields 144 matched records and 288 prompts per model. Four 4-bit open-weight instruction models produce 1,152 archived single-run greedy responses. Under the exact output contract, paired exact-contract accuracy (PECA) ranges from 1.4% to 67.4%; Qwen3-4B and Mistral-7B have the two highest point estimates on this four-model panel. Task-cluster resampling describes variation across the 12 semantic bases, while standalone-target scoring separates some target-selection failures from response-form errors. A post-hoc leading-target rule changes Phi-4-mini’s paired score substantially; the archived run lacks the stop-reason metadata needed to resolve whether generation termination caused these continuations. The findings describe synthetic, single-turn tasks under one user-role serialization. Adaptive attacks, role-level hierarchy, and tool-mediated effects require separate experiments View details
×