Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11597 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract The accelerated integration of generative AI technologies and agentic AI tools, particularly those like ChatGPT, into workplace settings has introduced complex challenges concerning data governance, regulatory compliance, and organizational privacy (GDPR 2016; CCPA/CPRA). This study introduces the Digital Shadow AI Risk Theoretical Framework (DART)—a novel theoretical framework designed to systematically identify, classify, and address the latent risks arising from the widespread, and often unregulated, use of AI systems in professional environments (NIST, 2023; OECD AI Policy Observatory, 2023). DART introduces six original, interrelated constructs developed in this study: Unintentional Disclosure Risk, Trust-Dependence Paradox, Data Sovereignty Conflict, Knowledge Dilution Phenomenon, Ethical Black Box Problem, and Organizational Feedback Loops. Each construct reflects a unique dimension of risk that emerges as organizations increasingly rely on AI-driven tools for knowledge work and decision-making. The framework is empirically tested through a mixed-methods research design involving hypothesis testing and statistical analysis of behavioral data gathered from cross-sectional surveys of industry professionals. Two cross-industry surveys (Survey-1: 416 responses, 374 analyzed; Survey-2: 203 responses, 179 analyzed) and CB-SEM tests supported seven of eight hypotheses; H4 (sovereignty) was not significant; H7 (knowledge dilution) was confirmed in replication. The findings highlight critical gaps in employee training, policy awareness, and risk mitigation strategies—underscoring the urgent need for updated governance frameworks, comprehensive AI-use policies, and targeted educational interventions. This paper contributes to emerging scholarship by offering a robust model for understanding and mitigating digital risks in AI-enabled workplaces, providing practical implications for compliance officers, risk managers, and organizational leaders aiming to harness the benefits of generative AI responsibly and securely. The novelty of DART lies in its explicit theorization of workplace-level behavioral risks—especially Shadow AI, which unlike Shadow IT externalizes organizational knowledge into adaptive systems—thereby offering a unified framework that bridges fragmented literatures and grounds them in empirical evidence. View details
From Unfinished to Done: Bridging the MDE Implementation Gap with Constrained LLMs
Christian Kirchhof
Lukas Netz
Dennis Mertens
Bernhard Rumpe
Proceedings of the ACM/IEEE 29th International Conference on Model Driven Engineering Languages and Systems, ACM (2026), pp. 152-163
Preview abstract Model-driven engineering (MDE) excels at ensuring architectural consistency and managing complexity, yet extending generated code skeletons with project-specific business logic often remains a time-consuming manual task. Conversely, large language models (LLMs) offer immense flexibility in coding but may produce vibe-coded results—software that looks correct but fails to adhere to intended architectures and implicit requirements. An open challenge lies in harnessing the adaptability of LLMs to handle implementation details without sacrificing the deterministic correctness of the MDE scaffolding. Previous attempts have largely relied on unconstrained coding assistants or required a human in the loop for interacting with the LLM. To overcome this, we introduce Agentic Gap Filling, a hybrid pipeline where LLM prompts are embedded directly within the model to drive the generation of code skeletons. The generated code contains protected TODO regions, which LLM agents subsequently resolve and a jury of LLM review agents and separate testing agents immediately validates. We integrate this approach into the MontiThings framework, demonstrating a workflow where models define the architecture of a project while artificial intelligence (AI) agents act as developers and reviewers, filling in the logic. This synergy ensures the structural integrity of the generated system while providing the flexibility of LLMs to adapt to various domain requirements. Our evaluation across 32 test configurations demonstrates that this multiagent workflow achieves up to 93.21% functional accuracy after four iterations, while preserving the structural integrity of the generated system. View details
Learning Conditional Averages
Marco Bressan
Nataly Brukhim
Nicolo Cesa-Bianchi
Emmanuel Esposito
Shay Moran
Maximilian Thiessen
COLT (2026)
Preview abstract We introduce the problem of learning \emph{conditional averages} in the PAC framework. The learner receives a sample labeled by an unknown target concept from a known concept class, as in standard PAC learning. However, instead of learning the target concept itself, the goal is to predict, for each instance, the average label over its \emph{neighborhood}---an arbitrary subset of points that contains the instance. In the degenerate case where all neighborhoods are singletons, the problem reduces exactly to classic PAC learning. More generally, it extends PAC learning to a setting that captures learning tasks arising in several domains, including explainability, fairness, and recommendation systems. %including explainability, fairness, and recommendation systems. Our main contribution is a complete characterization of when conditional averages are learnable, together with sample complexity bounds that are tight up to logarithmic factors. The characterization hinges on the joint finiteness of two novel combinatorial parameters, which depend on both the concept class and the neighborhood system, and are closely related to the independence number of the associated neighborhood graph. View details
Preview abstract This article introduces OpenClaw, an AI-powered workflow harness designed to reduce operational friction for developers. Unlike traditional chatbots, OpenClaw manages and persists context across complex engineering tasks, enabling asynchronous operations and mobile-first interactions. The post explores practical use cases, including incident triage from mobile, asynchronous pull request reviews, quick infrastructure scripting, and automating routine operational tasks. It also delves into the key architectural layers of an OpenClaw-like system—Connectors, Gateway/Session Manager, Agent Runtime, Memory/Configuration, and Skills/Tools. The article emphasizes the importance of security, observability, and proper integration with existing developer ecosystems, positioning OpenClaw as a shift towards reducing context switching and enhancing developer productivity by automating the workflows around coding. View details
Toward a Theory of Value in AI Alignment
Shazeda Ahmed
Abeba Birhane
Jackie Kay
Kris Shrishak
2026
Preview abstract Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a commensurate rise in ways these models cause harm, spanning areas from toxic speech and hallucinations to AI agents executing unauthorized actions. Given that these models are probabilistic and general-purpose by nature, it is impossible to enumerate all possible uses and outputs of the model to reach a fully aligned end state. Within the field of AI safety, these harmful instances are often framed as “the alignment problem,” of models being “misaligned” with human values. Researchers have responded by pursuing applied and theoretical AI “value alignment” efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? The study of human values has long been part of many academic disciplines outside of computer science, yet these disciplines are seldom consulted in AI alignment. Building on the theoretical insights of Zhi-Xuan’s (2024) "preferentist paradigm" critique, we conduct a review of influential AI alignment literature. We also draw from conceptions of human values from philosophy, anthropology, and sociology, to create an analytical schema. We annotated 94 AI value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on “preferences” as a stand-in that runs the risk of reducing complex, culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and LLM-as-a-judge approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. Overall, value alignment is often reduced to an exercise in utility maximization, which we argue abstracts human values away from their lived context. In making AI value alignment’s philosophical commitments explicit, we seek to bring greater specificity and under-explored perspectives into the debate on whether and how AI can address human values View details
Preview abstract The Abkhaz-Adyghe and Nakh-Daghestanian language families encompass 35 living languages that possess arguably the most complex modern Cyrillic orthographies due to their very sophisticated phonology. The relevant online data displays idiosyncratic patterns among which the use of confusable characters in input methods is the most prevalent. This work studies one such character---letter \emph{palochka}---that is shared by most writing systems in question. We investigate whether the patterns including variants of this character alone can act as language data markers when mining these languages in a large-scale web-crawled data. Using a wide-coverage off-the-shelf LID model (GlotLID) we further investigate the data mined using such patterns and estimate the effects of confusable character normalization on quality of paragraph-level LID predictions in 14 supported languages. According to GlotLID, the normalization significantly increases the recall (discovery of new language data) for some languages while degrading it for others. However, manual evaluation reveals that only 41\% of wins and 46\% of losses are accurate due to GlotLID prediction errors. We argue that despite finding useful signal higher precision LID approaches tailored to these long-tail languages are needed to improve the quality of mined data. View details
Assessing Global Flood Risk from Atmospheric Rivers through Physically Guided Machine Learning
Assaf Shmuel
Oleg Zlydenko
Martin Gauch
Colin Price
Scientific Reports (2026)
Preview abstract Atmospheric rivers (ARs) are narrow corridors of concentrated moisture transport that play a crucial role in the global water cycle, delivering both beneficial rainfall and severe floods. Here, we develop a physically guided, explainable machine learning framework to predict flood occurrence during AR conditions worldwide by integrating AR characteristics with meteorological and topographic variables. The best model achieves a receiver operating characteristic area under the curve (ROC AUC) of 0.94, outperforming a logistic regression baseline at 0.81. Despite their limited footprint, we find that one third of large midlatitude floods occur under AR conditions, reflecting their disproportionate role in global flood risk. SHAP analysis highlights integrated vapor transport, precipitation, and elevation as dominant predictors. We show the non-linear amplification of flood risk under combined conditions of high soil moisture and persistent AR activity, underscoring the importance of antecedent wetness in modulating flood risk. Using consistent reanalysis inputs, we find that model-estimated high-flood-risk AR conditions increased globally by over 10% from 1980 to 2020 and shifted poleward. Validation on an independent satellite-based flood database shows comparable skill. We estimate that roughly 90% of the global population lives in regions that experience at least one AR annually, underscoring the broad societal relevance of AR dynamics. These findings highlight the value of physically guided machine learning for mapping and monitoring AR-related flood risk globally, offering actionable insights for preparedness and climate adaptation. View details
Preview abstract Mid-air gestures in Extended Reality (XR) often lead to fatigue, discomfort and imprecision, limiting their suitability for extended use. Surface-based interactions offer a compelling alternative, providing improved accuracy, speed, and comfort. However, current egocentric vision-based methods struggle with reliable surface inputs due to challenges in hand tracking and surface-plane estimation from oblique and occluded viewing angles. To this extent, we introduce SurfaceXR, a novel sensor fusion approach that combines headset based hand tracking with micro-vibration data sampled from commodity smartwatch IMUs to enable precise and robust inputs on arbitrary surfaces. Our system is designed with flexibility in mind - it can function using only hand tracking, only IMU sensing, or optimally with both modalities combined. Our user study across 12 participants validates SurfaceXR's effectiveness in augmenting surface touch tracking and 8 class hand-surface gesture recognition, demonstrating significant improvements over single-modality approaches. Enabled by SurfaceXR, we demonstrate a series of interactive apps for both AR and VR, ranging from on-surface sketching, text entry and gesture based navigation. View details
Preview abstract Contrail cirrus represents a critical component of aviation’s non-CO2 climate impact, but its net radiative forcing, the balance between longwave warming and shortwave cooling, remains poorly constrained by direct observations. As a result, current assessments rely almost exclusively on microphysical models such as CoCiP and global climate simulations. Existing empirical estimates are largely restricted to young, linear tracks, because satellite detection masks have a poor recall of contrails once they spread and merge with natural cirrus, leaving a structural gap in our understanding of long-lived, non-linear contrail cirrus. To address this we use a causal framework that isolates the net radiative contrail effect of flight traffic over the Americas. Building on recent progress that quantified the longwave warming contrail effect using advected flight paths as a proxy for contrails, we expand this continuous treatment approach to capture the highly skewed shortwave cooling impact, delivering a 12-hour lifespan net observational radiative forcing. Our analysis reveals a statistically significant net warming energy forcing of 33.7 (95% CI: 20.8, 47.8) GJ/km flown from April 2019 to April 2020, providing a large-scale empirical quantification of long contrail lifespan impact of the same order as, though somewhat larger than, previous bottom-up simulation estimates. This observational benchmark offers an independent line of evidence on the sign and magnitude of the climate impact of contrails. View details
TCO-driven Storage Provisioning for Exascale Data Centers
Timothy Kim
Prashant Nema
Jai Menon
Rashmi Vinayak
Gregory R. Ganger
2026
Preview abstract Recent changes in data temperatures and storage device characteristics, both mechanical disk drives (HDDs) and solidstate drives (SSDs), expand the set of deployment options for exascale storage. Until recently, exascale storage systems followed a pattern of placing most data on HDDs with smaller amounts of SSD storage used for caching and performance-critical workloads. Exascale storage provisioning and dataset placement trade-offs have now changed. This paper describes a total cost of ownership (TCO) model that captures primary aspects of modern deployments and uses it to explore the new trade-off space. Using capacity and performance telemetry information for 43 production datasets+workloads at two large hyperscalers, we show significant changes from prior analyses of workloads and storage placement decisions across a multitude of storage device types. We also introduce a storage cluster TCO optimizer that identifies the lowest-TCO grouping and assignment of datasets to device types, exposing a number of insights that can help guide future deployments. For example, our analysis shows that the highest-density SSDs are particularly favorable for clusters with heavy AI/ML workloads but are only cost-effective at exascale when combined with high-density HDDs. Finally, we use our framework to evaluate how storage provisioning and overall TCO change as a function of key parameters like device write amplification, cluster power bounds, and the maximum number of device types allowed. View details
Rolling Shutter Relative Pose Estimation Made Practical
Daniel Barath
European Conference on Computer Vision (ECCV) (2026)
Preview abstract Rolling shutter (RS) cameras equip virtually all consumer devices, yet RS-aware relative pose estimation has remained impractical: the state-of-the-art solver requires a minimum of 20 point correspondences, making RANSAC-based robust estimation prohibitively expensive due to the exponential dependence of the iteration count on the sample size. We make RS relative pose estimation practical by introducing affine correspondences (ACs) into the RS two-view geometry. We derive novel \emph{RS-corrected affine constraints} that account for the coupling between point perturbations and the row-dependent essential matrix, providing two equations per correspondence beyond the standard epipolar constraint. Building on these constraints, we develop a linearized algebraic solver that estimates pose and RS motion from only 7 ACs. The solver exploits the physical smallness of RS parameters to linearize the constraints, eliminates the 12 RS unknowns via null-space projection, and solves the remaining degree-20 system via action matrices in 1.2\,ms. On the TUM RS benchmark, our method achieves the best pose and RS parameter accuracy among all tested methods and, uniquely among RS solvers, provides accurate translational velocity estimates -- which are poorly conditioned from point correspondences alone due to a $\vec{v}$-$\vec{t}$ coupling. On the global-shutter EuRoC MAV dataset, the solver achieves comparable accuracy to the standard 5-point algorithm, demonstrating that it generalizes well to the GS setting. Code will be made public. View details
Preview abstract This talk addresses the challenges of operating Google's monitoring systems at scale, handling terabytes of telemetry data and preventing overload from diverse workloads. We'll explore how Google's internal client library and Monarch, its planet-scale time-series database, work together for cost-effective data collection. Key principles include a distributed push model, dynamic client-side data reduction, centralized retention, and periodic metric analysis. The session will then bridge these concepts to the open-source world, discussing our work with OpenTelemetry's OpAMP protocol to achieve similar scalable and efficient telemetry collection. Attendees will gain insights into adapting these principles for cost savings and learn about our collaboration with the OpAMP SIG to benefit the broader community. View details
Preview abstract The field of Human-Computer Interaction is approaching a critical inflection point, moving beyond the era of static, deterministic systems into a new age of self-evolving systems. We introduce the concept of Adaptive generative interfaces that move beyond static artifacts to autonomously expand their own feature sets at runtime. Rather than relying on fixed layouts, these systems utilize generative methods to morph and grow in real-time based on a user’s immediate intent. The system operates through three core mechanisms: Directed synthesis (generating new features from direct commands), Inferred synthesis (generating new features for unmet needs via inferred commands), and Real-time adaptation (dynamically restructuring the interface's visual and functional properties at runtime). To empirically validate this paradigm, we executed a within-subject (repeated measures) comparative study (N=72) utilizing 'Penny,' a digital banking prototype. The experimental design employed a counterbalanced Latin Square approach to mitigate order effects, such as learning bias and fatigue, while comparing Deterministic interfaces baseline against an Adaptive generative interfaces. Participant performance was verified through objective screen-capture evidence, with perceived usability quantified using the industry-standard System Usability Scale (SUS). The results demonstrated a profound shift in user experience: the Adaptive generative version achieved a System Usability Scale (SUS) score of 84.38 ('Excellent'), significantly outperforming the Deterministic version’s score of 53.96 ('Poor'). With a statistically significant mean difference of 30.42 points (p < 0.0001) and a large effect size (d=1.04), these findings confirm that reducing 'navigation tax' through adaptive generative interfaces directly correlates with a substantial increase in perceived usability. We conclude that deterministic interfaces are no longer sufficient to manage the complexity of modern workflows. The future of software lies not in a fixed set of pre-shipped features, but in dynamic capability sets that grow, adapt, and restructure themselves in real-time to meet the specific intent of the user. This paradigm shift necessitates a fundamental transformation in product development, requiring designers to transcend traditional, linear workflows and evolve into 'System Builders'—architects of the design principles and rules that facilitate this new age of self-evolving software. View details
×