Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11583 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Every year an estimated 200,000 people go missing in the UK alone. Missing persons investigations involve challenging time-critical sensemaking tasks based on fragmented data sources. This paper describes a mixed-methods participatory study evaluating data science and AI-driven techniques (summarisation, fact extraction, and data visualisation) for supporting these investigations as part of a human-centered workflow. A series of human-AI interfaces were iteratively designed and tested with search officers and domain experts at Police Scotland. Based on findings, we describe: (1) user and information needs for missing persons investigations; (2) Insights on the benefits and challenges of applying LLM-based techniques in high-risk contexts; and (3) lessons for integrating AI for information and sensemaking tasks in policing more broadly. We highlight that in high-stakes contexts, where accuracy and context-sensitivity are paramount, AI techniques must be balanced with other approaches and designed in close partnership with end-users. View details
Preview abstract Quantization methods have significantly improved the compute and memory efficiency of Large Language Model (LLM) training. However, existing approaches still rely on accumulating their updates into high precision: concretely, gradient updates must be applied to a high-precision weight buffer, known as \textit{master weights}. This buffer introduces substantial memory overhead, particularly for Sparse Mixture of Experts (SMoE) models, where model parameters and optimizer states dominate memory usage. In this work, we introduce the Error-Compensating Optimizer (ECO), which \textit{for the first time} enables the complete elimination of master weights by directly accumulating updates into quantized parameters, by leveraging existing optimizer states. ECO quantizes the weights after every gradient step and injects the resulting quantization error into the optimizer's momentum buffer, creating an error-feedback loop with zero additional memory overhead for quantization. Beyond its practical efficiency, ECO comes with theoretical guarantees. Specifically, under standard assumptions, naive master weight removal can lead to unbounded drift from the ideal parameter trajectory, whereas ECO provably bounds this drift, ensuring stable convergence. We validate ECO across a range of models, including small transformers (30M--800M), Gemma-3 1B, and an SMoE 2.1B model, using FP8 quantization. In all cases, ECO achieves near-lossless accuracy compared to high-precision baselines. For large SMoE models, ECO reduces memory usage by up to 25\%, establishing a new Pareto frontier for the trade-off between static memory and training loss. View details
Preview abstract In large-scale distributed enterprises, traditional Knowledge Management (KM) systems face a critical failure mode: static documentation cannot keep pace with evolving operational realities and regional nuances. This "knowledge latency" forces employees out of self-service workflows and into costly support ticketing queues. This paper introduces SENTINEL, a geo-contextual AI framework designed to shift enterprise support from reactive retrieval to proactive interception. The architecture employs a novel dual-engine system integrated into an omni-present interface. The first engine utilizes Large Language Models (LLMs) to conduct pre-emptive, historical case-grounded audits of documentation, generating a "Contextual Density" score that identifies friction zones. The second engine is an autonomous Retrieval-Augmented Generation (RAG) agent that surfaces in-situ via a location-intelligent assistant window, resolving queries in real-time. By functioning as a strategic "defensive barrier" at the point of origin, SENTINEL demonstrates how a proactive AI assistant can drive high-fidelity, in-situ case deflection. View details
Preview abstract High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections. While large-scale generative priors have shown promise in static image inpainting, they often lack the structural constraints necessary to maintain identity, expression, and pose, leading to visible “identity drift” in both static images and dynamic sequences. In this paper, we propose a novel distillation framework that addresses the stochastic nature of generative priors. Our pipeline first extracts high-fidelity outputs from Gemini, regularizes them via dense geometric landmark constraints to preserve identity, expression, and pose, and finally applies physically-based simulation of lens optics to generate realistic paired training data. This process transfers Gemini’s photo-realistic, multi-view knowledge into a specialized restoration architecture, JFSnet (Joint Feature-Spatial network). JFSnet integrates DINOv2-based semantic features with a convolutional decoder for spatial reconstruction, leveraging equivariance constraints to improve high-frequency detail preservation and stability. Evaluations show that our approach significantly outperforms existing diffusion and GAN-based baselines, achieving the lowest FID scores and ranking highest in user preference studies across visual fidelity, identity preservation, and temporal consistency. View details
Unveiling the Global Landscape of Android Security Updates
Haiyun Deng
Abbas Acar
Esteban Luques
Harun Oz
Ahmet Aris
Selcuk Uluagac
IEEE Transactions on Dependable and Secure Computing (2026)
Preview abstract Android is the world’s leading mobile operating system, with over three billion active devices. Detecting vulnerabilities and ensuring timely patch deployment are critical to maintaining security. The Android Open Source Project (AOSP) has enhanced the transparency of security updates through Security Patch Levels. However, challenges related to update speed and availability persist. In 2022, Google reported that half of the zero-day vulnerabilities discovered in the wild were variations of vulnerabilities that had already been patched. Recent research mainly highlights delays in update distribution, often attributing them to fragmentation and focusing primarily on flagship devices or limited time-frames. Our approach takes a device-centric perspective to investigate Android update patterns, analyzing 567K security update records from 2014 to 2024, covering 904 distinct devices from six key Original Equipment Manufacturers (OEMs) across 98 countries. Our extensive analysis revealed notable differences in update release timing across OEMs, device types, and regions. Our study also examines documented vulnerabilities and weaknesses, while assessing OEM compliance with Android security guidelines. Our study shows that ∼89.7% of vulnerabilities on unpatched Android devices are exploitable without user interaction and with low attack complexity. We also identified delays linked to fragmentation and OEM-specific challenges, and provide actionable insights for improvement. View details
Preview abstract Advanced reasoning typically requires Chain-of-Thought prompting, which is accurate but incurs prohibitive latency and substantial test-time inference costs. The standard alternative, fine-tuning smaller models, often sacrifices interpretability while introducing significant resource and operational overhead. To address these limitations, we introduce Prompt-Level Distillation (PLD). We extract explicit reasoning patterns from a Teacher model and organize them into a structured list of expressive instructions for the Student model's System Prompt. Evaluated on the StereoSet and Contract-NLI datasets using Gemma-3 4B, PLD improved Macro F1 scores from 57\% to 90.0\% and 67\% to 83\% respectively, enabling this compact model to match frontier performance with negligible latency overhead. These expressive instructions render the decision-making process transparent, allowing for full human verification of logic, making this approach ideal for regulated industries such as law, finance, and content moderation, as well as high-volume use cases and edge devices. View details
CoDaS: AI Co-Data-Scientist for Biomarker Discovery via Wearable Sensors
Juro Gottweis
CJ Park
Salman Rahman
Ahmed Metwally
Hong Yu
Ivor Rendulic
Yuzhe Yang
Petar Sirkovic
Daniel McDuff
Shwetak Patel
Nicolas Stroppa
Yubin Kim
Mark Malhotra
Orson Xu
Sam Schmidgall
Tim Althoff
Elahe Vedadi
Cynthia Breazeal
Hae Won Park
(2026)
Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion
Pengcheng Jiang
Judith Yue Li
Moonkyung Ryu
R. Lily Hu
Kun Su
Liam Hebert
Hao Peng
Jiawei Han
Dima Kuzmin
2026
Preview abstract Many modern retrieval problems are set-valued: given a broad intent, the system must return a collection of results that optimizes higher-order properties (e.g., diversity, coverage, complementarity, coherence) while staying grounded to a fixed database. Set-valued objectives are inherently non-decomposable and are not captured by existing supervised (query, content) datasets which only prioritize top-1 retrieval. While reinforcement learning (RL) can optimize set-level objectives via interaction, deploying an RL-tuned LLM for fan-out retrieval is prohibitively expensive at query time. Conversely, diffusion-based generative retrieval enables efficient single-pass fan-out in embedding space, but requires objective-aligned training targets. To address these issues, we propose R4T (Retrieve-for-Train), which uses RL once as an objective transducer in a three step process: (i) train a fan-out LLM with composite set-level rewards, (ii) synthesize objective-consistent training pairs, and (iii) train a lightweight diffusion retriever to model the conditional distribution of set-valued outputs. Across Polyvore and a music playlist dataset, R4T improves retrieval quality over strong baselines while reducing query-time fan-out latency by an order of magnitude. View details
Preview abstract The rapid evolution of autonomous artificial intelligence systems has catalyzed the emergence of the “Agentic Economy,” a paradigm where decentralized intelligent agents execute complex workflows through heterogeneous tool environments. However, the scalability of this economy is fundamentally constrained by the N × M integration bottleneck, wherein N models require bespoke, brittle connectors for M data sources. This paper presents a comprehensive theoretical and quantitative evaluation of Universal Agentic Interoperability Protocols (UAIP), such as the Model Context Protocol (MCP). We formalize a Multi-Attribute Utility Analysis (MAUA) framework to assess the transition from point-to-point (P2P) RESTful architectures to standardized, stateful RPC-based hubs. Our analysis incorporates high-fidelity metrics for Architectural Entropy (Sa), Technical Debt Decay (δTD), and Protocol Efficiency (η). Through a large-scale enterprise simulation involving 100 agents and 500 tools, we demonstrate that UAIP adoption reduces total system configuration entropy by 84% and collapses maintenance overhead by an order of magnitude. The results provide a rigorous basis for the standardization of context exchange in future multi-agent ecosystems, ensuring that the next generation of AI infrastructure remains scalable, secure, and vendor-neutral. View details
Preview abstract Superconducting qubits are a leading platform for realizing fault-tolerant quantum computers. Current generations demonstrate fast, high fidelity quantum gates and readout on the order of hundreds of nanoseconds, while maintaining coherence times exceeding one hundred microseconds. Achieving this state-of-the-art performance requires a tight co-design, balancing fundamental physics, microwave engineering, and semiconductor fabrication. Readout designs, in particular, benefit from this multidisciplinary approach. In this talk, we discuss the current challenges for readout in superconducting quantum processors from the perspective of Google Quantum AI. We examine the intersection of device physics and microwave engineering constraints, illustrating how optimizing both is essential for scaling next-generation quantum systems. View details
Incentivizing Data Collaboration: A Mechanism Design Approach
Ali Makhdoumi
Azarakhsh Malekian
Ali Daei Naby
2026
Preview abstract We study the problem of incentivizing strategic agents to truthfully contribute high-quality data in collaborative learning settings, where each agent benefits from improved estimation based on others’ data. Each agent privately observes the quality of their data, and agents may misreport it if not incentivized properly. We cast this problem with a Bayesian mechanism design framework in which the platform aims to find the optimal data-sharing mechanism that jointly determines allocations and payments to maximize both estimation accuracy and platform revenue. We prove that the optimal mechanism that incentivizes truthful reporting takes the form of a \emph{personalized threshold and pricing} mechanism, in which each agent is allocated the learned estimator if their reported quality exceeds a (personalized) threshold and is charged a price based on the relevance of other agents' data in the learning task. We analyze this mechanism in a canonical Gaussian mean estimation task, derive a closed-form solution to the optimal mechanism, and highlight how data correlation affects the mechanism. We further extend the model to allow agents to exert costly efforts to improve their data quality before collaboration. We show that ''free-riding'' is mitigated as the optimal data-sharing mechanism induces a supermodular game: each agent is incentivized to exert more effort when others exert more. Finally, we show that equilibrium efforts form a complete lattice, and in the highest-effort equilibrium, each agent increases effort as others' data becomes more relevant in the learning task. View details
Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration
Suddhasatwa Bhaumik
Nilesh Jaiswal
Arjit Shukla
Divya Malhotra
Aniket Agrawal
Saurabh Garg
Suchit Puri
Google Cloud India, Google, S. No, AP81, 83, N Main Rd, near Hard Rock Cafe, Koregaon Park Annexe, Mundhwa, Pune, Maharashtra 411036 (2026)
Preview abstract As enterprises modernize legacy systems (e.g., monolithic Java architectures to Python microservices), Large Language Models (LLMs) have become instrumental in automated code translation. However, traditional vector-based Retrieval-Augmented Generation (Standard RAG) struggles with topological relationships, fetching isolated text chunks that frequently sever inheritance chains and lead to high compilation failure rates. This paper presents a comparative analysis between Standard RAG and a novel Hierarchical Context-Resident Graph (HCRG) methodology. Our pipeline utilizes tree-sitter for polyglot Abstract Syntax Tree (AST) extraction, mapping architectural edges into a Google Cloud Spanner Property Graph, and serializing this structure into a Gemini (on Vertex AI) Context Cache to enable topological, parent-first code translation. By shifting evaluation from naive text-overlap to a custom 7-metric framework measuring Software Engineering (SE) utility, empirical evaluations on the spring-petclinic-genai repository demonstrate significant structural improvements. Graph RAG decisively mitigates dependency loss, dropping the API hallucination rate from 56.4% to 16.2%. Furthermore, it improves Dependency Resolution Quality (DRQ) from 34.8% to 65.9% and enhances Parent-Child Consistency (PCC) from 26.7% to 45.5%. Interestingly, traditional lexical metrics fail to capture this divergence; both methodologies achieved an identical 91% average CodeBLEU score, effectively masking Standard RAG’s structural failures behind syntactically plausible but broken code. However, the results indicate that Graph RAG is not strictly superior across all dimensions. Providing the LLM with dense, global structural context introduces new vulnerabilities: Graph RAG suffers a severe degradation in Cyclomatic Complexity Consistency (dropping from Standard RAG’s 71.6% to 46.7%) due to defensive over-engineering by the LLM, alongside a slight drop in Docstring Preservation (67.0% down to 61.0%) caused by prompt attention dilution. Ultimately, this research validates that while Graph RAG trades an increase in code complexity for critical reductions in API hallucinations, it offers a substantially more viable and architecturally sound path for automated enterprise codebase modernisation. View details
Preview abstract Using generative artificial intelligence with sensitive data may present challenges, as transmitting personally identifiable information or protected health information to third-party providers can introduce security risks, and some data masking techniques can reduce reasoning capabilities. A described system uses a proxy, masking layer that can intercept data within an enterprise's secure perimeter. This layer can substitute sensitive strings with persistent, structured semantic tokens that may be enriched with non-sensitive metadata hints to help preserve context. An external artificial intelligence can perform reasoning on this abstracted data, and its tokenized response can be re-hydrated into readable text on a client device (e.g., a smartphone, computer, or wearable device). This approach may allow third-party models to reason on proprietary information without direct access to the underlying plaintext data, which can assist organizations in managing data sovereignty while maintaining functional utility. View details
×