Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11583 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Unveiling the Global Landscape of Android Security Updates
Haiyun Deng
Abbas Acar
Esteban Luques
Harun Oz
Ahmet Aris
Selcuk Uluagac
IEEE Transactions on Dependable and Secure Computing (2026)
Preview abstract Android is the world’s leading mobile operating system, with over three billion active devices. Detecting vulnerabilities and ensuring timely patch deployment are critical to maintaining security. The Android Open Source Project (AOSP) has enhanced the transparency of security updates through Security Patch Levels. However, challenges related to update speed and availability persist. In 2022, Google reported that half of the zero-day vulnerabilities discovered in the wild were variations of vulnerabilities that had already been patched. Recent research mainly highlights delays in update distribution, often attributing them to fragmentation and focusing primarily on flagship devices or limited time-frames. Our approach takes a device-centric perspective to investigate Android update patterns, analyzing 567K security update records from 2014 to 2024, covering 904 distinct devices from six key Original Equipment Manufacturers (OEMs) across 98 countries. Our extensive analysis revealed notable differences in update release timing across OEMs, device types, and regions. Our study also examines documented vulnerabilities and weaknesses, while assessing OEM compliance with Android security guidelines. Our study shows that ∼89.7% of vulnerabilities on unpatched Android devices are exploitable without user interaction and with low attack complexity. We also identified delays linked to fragmentation and OEM-specific challenges, and provide actionable insights for improvement. View details
Large-scale, interpretable gene regulatory network inference through biologically informed matrix factorization
Soel Micheletti
Viola Fanfani
Julia Vogt
John Quackenbush
Jonas Fischer
Alexander Marx
Panagiotis Mandros
bioRxiv (2026)
Preview abstract Gene regulatory networks (GRNs) provide a mechanistic framework for understanding how transcription factors coordinate gene expression to establish cellular identity and phenotype. Methods that integrate gene expression with motif-derived regulatory priors and other sources of biological information have substantially advanced gene regulatory network inference by reconstructing condition-specific regulatory architecture. These approaches estimate the evidence supporting regulatory interactions and have proven remarkably successful in a wide range of biological applications. A complementary view of regulatory networks, however, seeks to estimate the effect of those interactions on gene expression itself, providing a framework in which regulatory edges can be interpreted as activating or inhibitory influences on transcription. We developed Giraffe, a biologically informed matrix factorization framework that jointly estimates transcription factor activities and gene regulatory networks by integrating gene expression, motif-based regulatory priors, and transcription factor protein-protein interactions. Giraffe estimates signed partial regulatory effects whose magnitude and sign can be interpreted as the strength and direction of transcriptional regulation. Building directly on the biological framework established by methods such as PANDA, Giraffe provides a complementary representation of gene regulatory networks that emphasizes mechanistic interpretation while remaining scalable, flexible, and computationally efficient. Across synthetic benchmarks, six human tissues, yeast transcription factor perturbation experiments, and liver hepatocellular carcinoma, Giraffe accurately reconstructs regulatory interactions while distinguishing activating from inhibitory regulation with high accuracy. The inferred networks recover known features of tissue-specific regulation, correctly classify regulatory effects in transcription factor perturbation experiments, and identify biologically coherent changes in regulatory programs associated with liver cancer. Together, these results demonstrate that estimating the direction of transcriptional regulation provides a complementary perspective on gene regulatory networks that facilitates biological interpretation and hypothesis generation. View details
Preview abstract We study advertising in conversational LLM platforms, where the platform gradually learns a user's preferences before deciding when and what ad to offer. We model this as a stochastic control problem in which the platform balances the value of information acquisition against the risk of user departure. Although the optimal policy that maximizes welfare admits a threshold structure, computing it exactly is infeasible without strong assumptions. We develop an approximate threshold policy based on upper and lower bounds on the continuation value and prove explicit performance guarantees that scale with uncertainty, learning dynamics, and heterogeneity of the advertisers. Furthermore, we extend this framework to revenue maximization, where advertisers hold private valuations. We establish a "virtual welfare equivalence" in this dynamic setting, demonstrating that the revenue-optimal incentive-compatible mechanism is implemented by applying the welfare-maximizing policy to advertisers' properly defined virtual valuations. This mechanism generalizes classical optimal auction theory to environments where the information structure evolves endogenously through user interaction, enabling us to derive approximately revenue-optimal mechanisms. View details
AI, Identity, and Ethical Governance: Building Trust in High-Stakes Systems
Ibrahim Waziri Jr.
Abhilasha Bhargav - Spantzel
RSAC (2026)
Preview abstract As AI redefines identity verification in high stakes systems, it introduces novel risks like deepfake fraud and algorithmic bias, creating a critical trust deficit. This session will provide a practical framework for ethical governance, equipping leaders to build and manage secure, fair, and fundamentally trustworthy AI systems by design. View details
Preview abstract A growing body of qualitative research has identified contextual risk factors that elevate people’s chances of experiencing digital-safety attacks. However, the lack of quantitative data on the population level distribution of these risk factors prevents policymakers and tech companies from developing targeted, evidence-based interventions to improve digital safety. To address this gap, we surveyed 5,001 adults in the United States to analyze: (1) the frequency of and relationship between digital-safety attacks (e.g., scams, harassment, account hacking), and (2) how these attacks align with 10 contextual risk factors. Nearly half of our respondents identify as resource constrained, which significantly correlates with higher likelihood of experiencing four common attacks. We also present qualitative insights to expand our understanding of the factors beyond the existing literature (e.g., “prominence” included high-visibility roles in local communities). This study provides the first large-scale quantitative analysis correlating digital-safety attacks with contextual risk factors and demographics. View details
Preview abstract We introduce a new context-enriched time series forecasting benchmark TimesX. TimesX contains a wide selection of high-quality real-world time series and diverse textual contexts from an automated generating pipeline, which helps address three main issues of existing benchmarks: (1) poor generalization due to low data volume and data being synthetic, (2) restricted forms of context, and (3) an inability to mitigate data leakage. We conduct a thorough empirical study of current multimodal solutions on TimesX. Our results suggest that most multimodal solutions that work well on existing benchmarks may fail on TimesX. In contrast, simple ensemble methods that leverage the rich textual context can outperform strong unimodal baselines and other multimodal baselines. ** Below this is what was submitted to ITP. ** We create a real world multimodal time-series forecasting benchmark that encompasses diverse domains and regions. Each time-series is annotated by various kinds of contexts like metadata, date and holiday information, dynamic events related to the time-series. This is sufficiently more advanced than other available benchmarks which rely wither on static metadata alone or synthetic examples. This forms a test bed for multimodal forecasting. We also present some baseline results showing that ensembles of publicly available LLMs and time-series foundation models can demonstrate non-trivial performance on this bechmark. View details
Preview abstract The rapid evolution of autonomous artificial intelligence systems has catalyzed the emergence of the “Agentic Economy,” a paradigm where decentralized intelligent agents execute complex workflows through heterogeneous tool environments. However, the scalability of this economy is fundamentally constrained by the N × M integration bottleneck, wherein N models require bespoke, brittle connectors for M data sources. This paper presents a comprehensive theoretical and quantitative evaluation of Universal Agentic Interoperability Protocols (UAIP), such as the Model Context Protocol (MCP). We formalize a Multi-Attribute Utility Analysis (MAUA) framework to assess the transition from point-to-point (P2P) RESTful architectures to standardized, stateful RPC-based hubs. Our analysis incorporates high-fidelity metrics for Architectural Entropy (Sa), Technical Debt Decay (δTD), and Protocol Efficiency (η). Through a large-scale enterprise simulation involving 100 agents and 500 tools, we demonstrate that UAIP adoption reduces total system configuration entropy by 84% and collapses maintenance overhead by an order of magnitude. The results provide a rigorous basis for the standardization of context exchange in future multi-agent ecosystems, ensuring that the next generation of AI infrastructure remains scalable, secure, and vendor-neutral. View details
Preview abstract Using generative artificial intelligence with sensitive data may present challenges, as transmitting personally identifiable information or protected health information to third-party providers can introduce security risks, and some data masking techniques can reduce reasoning capabilities. A described system uses a proxy, masking layer that can intercept data within an enterprise's secure perimeter. This layer can substitute sensitive strings with persistent, structured semantic tokens that may be enriched with non-sensitive metadata hints to help preserve context. An external artificial intelligence can perform reasoning on this abstracted data, and its tokenized response can be re-hydrated into readable text on a client device (e.g., a smartphone, computer, or wearable device). This approach may allow third-party models to reason on proprietary information without direct access to the underlying plaintext data, which can assist organizations in managing data sovereignty while maintaining functional utility. View details
Inference Perf: A Benchmarking Tool for GenAI Inference
Sachin Varghese
Jason Kramberger
Brendan Slabe
Chen Wang
Yuan Tang
Journal of Open Source Software (2026)
Preview abstract Inference Perf is a generative AI (GenAI) inference performance benchmarking tool aimed at benchmarking and analyzing the performance of inference deployments. It is designed to be model-server agnostic, allowing for apples-to-apples comparisons across different model servers and serving stacks. As part of the inference benchmarking and metrics standardization effort in the Kubernetes wg-serving, it seeks to standardize tooling and metrics for measuring inference performance across the Kubernetes and model server communities. View details
Preview abstract Large-scale cloud-native systems generate continuous streams of operational alerts across distributed microservice architectures. On-call engineers must manually triage these alerts by correlating signals from heterogeneous observability tools, a process that is time-consuming, cognitively demanding, and prone to error. Despite advances in monitoring and anomaly detection, incident triage remains largely manual. This paper presents a declarative, large language model (LLM)–driven multi-agent approach to automating incident triage and Service Level Objective (SLO) monitoring. The proposed design constrains agent behavior using domain-expertauthored investigation workflows, enabling deterministic execution and reproducibility while preserving operational safety. The framework integrates a unified tool execution layer for interacting with diverse observability systems and an enhanced retrieval-augmented generation (RAG) pipeline optimized for operational knowledge retrieval. The approach has been evaluated in a production cloud environment spanning multiple microservices and geographic regions. Results show reductions in high-severity incident triage time from approximately 30 minutes to under 5 minutes, alert acknowledgement latency from minutes to seconds, and service onboarding effort from weeks to days. These findings suggest that constrained multi-agent systems can substantially reduce on-call cognitive load while maintaining reliability and human oversight. View details
Preview abstract In large-scale distributed enterprises, traditional Knowledge Management (KM) systems face a critical failure mode: static documentation cannot keep pace with evolving operational realities and regional nuances. This "knowledge latency" forces employees out of self-service workflows and into costly support ticketing queues. This paper introduces SENTINEL, a geo-contextual AI framework designed to shift enterprise support from reactive retrieval to proactive interception. The architecture employs a novel dual-engine system integrated into an omni-present interface. The first engine utilizes Large Language Models (LLMs) to conduct pre-emptive, historical case-grounded audits of documentation, generating a "Contextual Density" score that identifies friction zones. The second engine is an autonomous Retrieval-Augmented Generation (RAG) agent that surfaces in-situ via a location-intelligent assistant window, resolving queries in real-time. By functioning as a strategic "defensive barrier" at the point of origin, SENTINEL demonstrates how a proactive AI assistant can drive high-fidelity, in-situ case deflection. View details
Holistic Latent Diffusion Acceleration: Unifying Spatial, Temporal, and Architectural Efficiency
Ruyi An
Xin Yuan
Xixi Hu
Hongliang Fei
Mingyuan Zhou
Keyang Xu
ICML 2026 Workshop on Structured Probabilistic Inference & Generative Modeling
Preview abstract Latent Diffusion Models (LDM) face three compounding efficiency challenges in practical deployment: i) the temporal latency of iterative sampling; ii) the architectural overhead of heavy backbone parameter counts; and iii) the spatial cost of high-dimensional latent grids. While recent acceleration methods have made substantial progress on temporal distillation and architectural compression, the spatial axis is often inherited from the teacher tokenizer and treated as fixed. In this work, we recast latent resolution as an optimizable efficiency axis and introduce a unified framework that optimizes all three dimensions simultaneously. We introduce a novel strategy of Score-Compatible Tokenizer Distillation (SCTD), which leverages score-matching principles to align a spatially compact latent space with the induced distribution of a frozen, powerful teacher model, distilling the teacher's generative prior into a compressed, lower-dimensional compatible manifold. With flexibility provided by SCTD, we can surrogate a computationally heavy teacher backbone with a lightweight student architecture operating strictly within this new compressed space. Finally, we apply temporal distillation to collapse the sampling trajectory, producing a one-step generator that operates at peak efficiency. Our method yields a student generator outperforming existing single-axis acceleration methods in efficiency and throughput, while maintaining competitive generation quality. With reduced peak memory usage and latency, our method enables resource-constrained deployment and high-volume serving of high-fidelity LDM. View details
Preview abstract While AI scientists increasingly automate research tasks through advanced language models, generating publication-ready illustrations remains a labor-intensive bottleneck in the scientific workflow. To lift this burden, we introduce PaperBanana, an agentic framework for automated generation of publication-ready academic diagrams. Powered by Nano-Banana-Pro and Gemini-3-Pro, PaperBanana orchestrates a team of specialized agents to retrieve reference examples, devise detailed plans for content and style, render the image, and perform iterative refinement based on self-critique. To rigorously evaluate our framework and address the absence of dedicated benchmarks for automated academic illustration, we introduce PaperBananaBench, comprising 292 test cases for methodology diagrams curated from NeurIPS 2025 publications. Comprehensive experiments demonstrate that PaperBanana consistently outperforms vanilla Nano-Banana-Pro across all four dimensions—faithfulness, conciseness, readability, and aesthetics—achieving human-level performance. We further show that PaperBanana seamlessly extends to statistical plots through targeted adaptations. Collectively, PaperBanana enables AI scientists to fully automate the generation of publication-ready academic illustrations. View details
×