Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11557 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract We prove the following asymptotically tight lower bound for k-color discrepancy: For any k ≥ 2, there exists a hypergraph with n vertices such that its k-color discrepancy is at least Ω(√n). This improves on the previously known lower bound of Ω(√n/ log k) due to Caragiannis et al. [CLS25]. As an application, we show that our result implies improved lower bounds for group fair division. View details
Learning Conditional Averages
Marco Bressan
Nataly Brukhim
Nicolo Cesa-Bianchi
Emmanuel Esposito
Shay Moran
Maximilian Thiessen
COLT (2026)
Preview abstract We introduce the problem of learning \emph{conditional averages} in the PAC framework. The learner receives a sample labeled by an unknown target concept from a known concept class, as in standard PAC learning. However, instead of learning the target concept itself, the goal is to predict, for each instance, the average label over its \emph{neighborhood}---an arbitrary subset of points that contains the instance. In the degenerate case where all neighborhoods are singletons, the problem reduces exactly to classic PAC learning. More generally, it extends PAC learning to a setting that captures learning tasks arising in several domains, including explainability, fairness, and recommendation systems. %including explainability, fairness, and recommendation systems. Our main contribution is a complete characterization of when conditional averages are learnable, together with sample complexity bounds that are tight up to logarithmic factors. The characterization hinges on the joint finiteness of two novel combinatorial parameters, which depend on both the concept class and the neighborhood system, and are closely related to the independence number of the associated neighborhood graph. View details
Preview abstract This article presents a novel approach to automating operations tasks, particularly incident triage, by using AI agents defined entirely in Markdown. These agents orchestrate actions across various observability tools (e.g., Datadog, Splunk) and use the file system for state and communication, mimicking the Unix philosophy. The system enables parallel investigations, cross-tool validation, and structured reporting without traditional coding frameworks. View details
Compact Conformal Subgraphs
Kamesh Munagala
Aravindan Vijayaraghavan
ICML (2026)
Preview abstract Conformal prediction provides rigorous uncertainty guarantees for model outputs but can produce prohibitively large prediction sets in structured domains such as routing, planning, or sequential recommendation. We introduce \emph{graph-based conformal compression}, a framework for constructing compact subgraphs that preserve the statistical validity of conformal prediction while reducing structural complexity. We study a formulation that selects a smallest subgraph capturing a prescribed fraction of conditional probability mass, and reduce to a weighted version of densest $k$-subgraphs in hypergraphs, in the regime where the subgraph has a large fraction of edges. We design efficient approximation algorithms that achieve constant factor coverage and size trade-offs. Our results highlight an algorithmic regime, distinct from classical densest-$k$-subgraph hardness settings, where the problem can be approximated efficiently, bridging conformal prediction with combinatorial graph compression. We finally validate our algorithmic approach on synthetic and real-world instances of trip planning and navigation, showing in each case that our approach handily beats natural baselines. View details
Holistic Latent Diffusion Acceleration: Unifying Spatial, Temporal, and Architectural Efficiency
Ruyi An
Xin Yuan
Xixi Hu
Hongliang Fei
Mingyuan Zhou
Keyang Xu
ICML 2026 Workshop on Structured Probabilistic Inference & Generative Modeling
Preview abstract Latent Diffusion Models (LDM) face three compounding efficiency challenges in practical deployment: i) the temporal latency of iterative sampling; ii) the architectural overhead of heavy backbone parameter counts; and iii) the spatial cost of high-dimensional latent grids. While recent acceleration methods have made substantial progress on temporal distillation and architectural compression, the spatial axis is often inherited from the teacher tokenizer and treated as fixed. In this work, we recast latent resolution as an optimizable efficiency axis and introduce a unified framework that optimizes all three dimensions simultaneously. We introduce a novel strategy of Score-Compatible Tokenizer Distillation (SCTD), which leverages score-matching principles to align a spatially compact latent space with the induced distribution of a frozen, powerful teacher model, distilling the teacher's generative prior into a compressed, lower-dimensional compatible manifold. With flexibility provided by SCTD, we can surrogate a computationally heavy teacher backbone with a lightweight student architecture operating strictly within this new compressed space. Finally, we apply temporal distillation to collapse the sampling trajectory, producing a one-step generator that operates at peak efficiency. Our method yields a student generator outperforming existing single-axis acceleration methods in efficiency and throughput, while maintaining competitive generation quality. With reduced peak memory usage and latency, our method enables resource-constrained deployment and high-volume serving of high-fidelity LDM. View details
Preview abstract Social scientists rely on hypothesis testing to support their research conclusions, but our standard procedures are designed for testing one hypothesis rather than adjudicating between rival possibilities. We develop a new framework, “classification testing”, as an alternative. Instead of selecting one hypothesis to test, a researcher conducting a classification test decides what qualitative distinctions (“classes”) are most substantively relevant; the test either assigns the estimand to a class with error control similar to that of a conventional hypothesis test, or declares the result inconclusive. We argue that classification testing is superior to current practice not just when the objective is to adjudicate between rival possibilities but also when there is one research hypothesis to be tested, because classification testing exposes that hypothesis to refutation. We illustrate the framework by applying it to a well-known media experiment and offer an R package to aid in implementation. View details
Incentivizing Data Collaboration: A Mechanism Design Approach
Ali Makhdoumi
Azarakhsh Malekian
Ali Daei Naby
2026
Preview abstract We study the problem of incentivizing strategic agents to truthfully contribute high-quality data in collaborative learning settings, where each agent benefits from improved estimation based on others’ data. Each agent privately observes the quality of their data, and agents may misreport it if not incentivized properly. We cast this problem with a Bayesian mechanism design framework in which the platform aims to find the optimal data-sharing mechanism that jointly determines allocations and payments to maximize both estimation accuracy and platform revenue. We prove that the optimal mechanism that incentivizes truthful reporting takes the form of a \emph{personalized threshold and pricing} mechanism, in which each agent is allocated the learned estimator if their reported quality exceeds a (personalized) threshold and is charged a price based on the relevance of other agents' data in the learning task. We analyze this mechanism in a canonical Gaussian mean estimation task, derive a closed-form solution to the optimal mechanism, and highlight how data correlation affects the mechanism. We further extend the model to allow agents to exert costly efforts to improve their data quality before collaboration. We show that ''free-riding'' is mitigated as the optimal data-sharing mechanism induces a supermodular game: each agent is incentivized to exert more effort when others exert more. Finally, we show that equilibrium efforts form a complete lattice, and in the highest-effort equilibrium, each agent increases effort as others' data becomes more relevant in the learning task. View details
Type-Aware Ranking of Urban Similarity from Aerial Imagery
Idan Kligvasser
Yotam Intrator
Yuval Desheh
Aviad Barzilai
Niv Efron
Ehud Rivlin
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops (2026), pp. 821-829
Preview abstract Estimating and ranking cross-city similarity from aerial imagery is a fundamental challenge in remote sensing and geospatial representation learning. Urban environments differ widely in road layout, marking conventions, and infrastructure design, yet standard visual representations often struggle to disentangle these meaningful structural variations from superficial appearances. In this work, we propose a type-aware contrastive learning framework that measures urban similarity by explicitly modeling distinct infrastructure elements. Leveraging open-vocabulary retrieval, we construct a globally diverse dataset of road-related features, such as intersections, crosswalks, and bus lanes, and train a type-conditioned Vision Transformer that fuses visual features with CLIP-derived semantic embeddings. Crucially, we introduce an adaptive per-type contrastive loss that dynamically emphasizes infrastructure categories with high discriminative power while down-weighting less informative types. To quantify city-level similarity, we aggregate per-type cosine similarities via a lightweight classifier to generate a global city-to-city similarity matrix. Experiments demonstrate that this type-aware approach significantly improves clustering quality and successfully generalizes to unseen cities, establishing a scalable, interpretable foundation for comparative urban analysis. View details
A simple and efficient implementation of strong call by need by an abstract machine
Małgorzata Biernacka
Witold Charatonik
Journal of Functional Programming, Volume 36 (2026)
Preview abstract We present an abstract machine for a strong call-by-need strategy in the lambda calculus. The machine has been derived automatically from a higher-order evaluator that uses the technique of memothunks to implement laziness. The derivation has been done with the use of an off-the-shelf transformation tool implementing the "functional correspondence" between higher-order interpreters and abstract machines, and it yields a simple and concise description of the machine. We prove that the resulting machine conservatively extends the lazy version of Krivine machine for the weak call-by-need strategy, and that it simulates the normal-order strategy in bilinear number of steps. View details
Usability Hasn’t Peaked: Exploring How Expressive Design Overcomes the Usability Plateau
Alyssa Sheehan
Bianca Gallardo
Ying Wang
Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26), April 13–17, 2026, Barcelona, Spain (2026)
Preview abstract Critics have argued that mobile usability has largely been optimized, and that only incremental gains are possible. We set out to explore if the newest generation of design systems, which promote greater flexibility and a return to design basics, could produce substantially more usable designs while maintaining or increasing aesthetic judgments. Through a study with 48 diverse participants completing tasks in 10 different applications, we found that in designs created following Material 3 Expressive guidelines, users fixated on the correct screen element for a task 33% faster, completed tasks 20% faster, and rated experiences more positively compared to versions designed using the previous Material design system. These improvements in performance and aesthetic ratings challenge the premise of a usability plateau and show that mobile usability has not peaked. We illustrate specific opportunities to make mobile experiences more usable by returning to design fundamentals while highlighting risks of added flexibility. View details
Preview abstract As organizations pursue AI transformation in large environments, ambitions often collide with the practical urgency of fixed-deadline data center exits. This article argues that a non-negotiable migration date should not be viewed merely as an infrastructure constraint, but as a critical deadline for establishing AI readiness. AI systems amplify the strengths and weaknesses of the underlying architecture; fragmented data or inconsistent infrastructure will lead to unreliable AI outcomes. To build a foundation for future intelligence, architects must prioritize resilience, standardization, and governance early in the migration process. Ultimately, successful AI transformation depends less on the speed of deployment and more on foundational architectural decisions such as Infrastructure-as-Code and unified telemetry—made before the migration concludes. View details
Multi-agent cooperation through in-context co-player inference
Rajai Nasser
Alexander Meulemans
Marissa Weis
João Sacramento
Maciej Wołczyk
Rif A. Saurous
2026
Preview abstract Achieving cooperation among self-interested agents remains a fundamental challenge in multi-agent reinforcement learning. Promising recent work has shown that cooperation can be established between ``learning-aware'' agents that explicitly account for and shape the learning dynamics of their co-players. However, existing approaches typically rely on hardcoded, often inconsistent, assumptions about co-player learning rules or enforce a strict separation between ``naive learners'' updating on fast timescales and ``meta-learners'' observing these updates. Here, we demonstrate that the in-context learning capabilities of sequence models allow for co-player learning awareness without requiring hardcoded assumptions or explicit timescale separation. We show that training sequence model agents against a diverse distribution of co-players naturally induces \textit{in-context best-response} strategies, effectively functioning as learning algorithms on the fast intra-episode timescale. We find that the cooperative mechanism identified in prior work—where vulnerability to extortion drives mutual shaping—emerges naturally in this setting: in-context adaptation renders agents vulnerable to extortion, and the resulting mutual pressure to shape the opponent's in-context learning dynamics resolves into the learning of cooperative behavior. View details
Preview abstract Source-to-source compilers may perform inefficiently by executing transpilation passes on scripts that do not contain the specific language features a pass is designed to transform, potentially leading to redundant processing. A compiler can analyze a script to generate a per-script feature map, for example, by identifying language features in its abstract syntax tree (AST). Before executing a transpilation pass, the compiler can check this map and may bypass the pass for that script if the specific feature targeted by the pass is not present. This feature map can also be dynamically updated throughout the compilation process as other passes transform the code. This method of conditional pass execution based on content-aware analysis may reduce redundant AST traversals, which could decrease overall compilation time and computational resource consumption. View details
×