Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11487 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract AI agents equipped with tool-calling capabilities are susceptible to \emph{Indirect Prompt Injection} (IPI) attacks. In this attack scenario, malicious commands hidden within \emph{untrusted} content trick the agent into performing unauthorized actions. Existing defenses can reduce attack success but often suffer from the \emph{over-defense dilemma}: they deploy expensive, \emph{always-on} sanitization that degrades utility and latency even in benign scenarios. We revisit IPI through an operational causal lens: a successful injection manifests as a \emph{grounding collapse} where the user request no longer provides decisive support for the agent's privileged action, while a particular untrusted segment provides disproportionate marginal support. Based on this signature, we propose \texttt{CausalArmor}, a selective defense framework that (i) computes lightweight, normalized leave-one-out attributions at privileged decision points, and (ii) triggers targeted sanitization only when an untrusted segment dominates the user intent. Additionally, CausalArmor employs \emph{retroactive Chain-of-Thought masking} to prevent the agent from acting on ``poisoned" reasoning traces. Experiments on AgentDojo and DoomArena demonstrate that CausalArmor matches the security of aggressive defenses with explainability while preserving utility and latency of AI agents. View details
Preview abstract In this work, we prove computational lower bounds against differentially private (DP) coreset. Specifically, assuming the existence of one-way functions, we show that no polynomial-time $(\epsilon, 1/n^{\omega(1)})$-DP algorithm can compute $(\alpha, \beta)$-coreset for $k$-means in the $\ell_\infty$ metric for some constant $\alpha > 1$. For the Euclidean metric, we show a similar result but only for $\alpha = 1 + \Theta\left(\frac{1}{d^2}\right)$ where $d$ is the dimension. View details
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
Ofer Meshi
Guy Tennenholtz
Jihwan Jeong
Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL-26), Rabat, Morocco (2026), pp. 5270-5304
Preview abstract LLM-based user simulators are a scalable solution for improving conversational AI, but a critical realism gap undermines their effectiveness. To close this gap, we introduce a framework for building and validating high-fidelity simulators. We present a novel dataset of human-AI shopping conversations designed to capture a wide spectrum of user experiences. To measure fidelity, we propose a hybrid evaluation protocol that combines statistical alignment with a learned, discriminator-based Human-Likeness Score. Our most sophisticated simulator, trained via reinforcement learning with iterative critique, achieves a significant leap in realism. Critically, we demonstrate through counterfactual validation that our simulator—trained exclusively on optimal interactions—realistically adapts its behavior to suboptimal system responses, mirroring real user reactions and marking a key advance in creating reliable simulators for robust AI development. View details
Preview abstract The management of a hybrid workforce comprising human and autonomous computational agents may be challenged by the use of separate systems for human capital and software assets, which can create a governance gap. A system can provide a unified framework for managing a hybrid workforce. For example, the system may utilize a labor service mesh to analyze and route tasks to either a human intent tier or an agentic execution tier. A potential principle of the system is structural symmetry, where computational agents can be assigned digital identities and managed through a lifecycle process that may parallel human resource functions, such as onboarding, performance evaluation, and structured offboarding. This integrated approach can facilitate a unified system of record and governance model for an organization's intelligence capacity. View details
Towards A Human-in-the-Loop Framework for Reliable Patch Evaluation using an LLM-as-a-Judge
Renyao Wei
José Cambronero
AI-SQE '26: Proceedings of the 1st International Workshop on AI for Software Quality Evaluation - Judgment, Metrics, Benchmarks, and Beyond, ACM (Association for Computing Machinery), New York, NY, USA (2026), pp. 19 - 28
Preview abstract Reliable evaluation is crucial for advancing Automated Program Repair (APR), but prevailing benchmarks that rely on execution-based evaluation methods (pass@k) often fail to capture the patch quality required for real-world adoption. This creates a significant gap between automated metrics and true patch validity (valid@k), a discrepancy observed across several state-of-the-art techniques. To develop a scalable solution for measuring valid@k, we first study the human evaluation process itself. While manual assessment can determine validity, we find it suffers from poor inter-rater reliability (Fleiss' Kappa k=0.307). Our foundational insight is that this inconsistency is largely resolved when evaluators use a shared, high-quality rubric, which significantly improves agreement. Building on this finding, we propose an LLM-as-a-Judge framework that operationalizes rubric-guided evaluation at scale. Our method employs a human-in-the-loop workflow where an LLM first generates a candidate rubric for a given bug, which a human expert then reviews and refines into a "golden" evaluation standard. This golden rubric is then used by an LLM judge to assess the validity of candidate patches. In an evaluation on 48 bugs and 115 patches, our LLM judge demonstrates substantial agreement with the consensus of human developers. This work contributes a scalable and reliable methodology for approximating valid@k, providing a much-needed high-fidelity signal for measuring true progress in the field of automated program repair. View details
Preview abstract Managing compiler build errors that can arise during infrastructure upgrades in large, polyglot codebases may be challenging, as manual remediation can be slow and some automated tools may not support modern language syntax. A system can provide automated error remediation by ingesting compiler diagnostics and analyzing source code using an Abstract Syntax Tree (AST). A recursive scope resolution algorithm, for example, can traverse the AST to identify a specific and narrowly-scoped code block at which to apply an error suppression. Conversely, this algorithmic complexity can be bypassed when lexical scope resolution is not required, and the system can identify the specific location of error suppressions directly from the error's exact coordinates. The system may then generate and apply language-specific patches, such as structured comments for JavaScript source files or line-scoped comments for TypeScript source files, for example, by using a transactional rewrite engine. This approach can provide a scalable method for managing automated code remediation, which may facilitate infrastructure upgrades by reducing the need for manual intervention. View details
Model Checking for Flexible Networking Protocols
Andrew Johnson
Sean Song
Proceedings of the 38th International Conference on Computer Aided Verification (CAV 2026)
Preview abstract Operating a network is a daunting task. Operating one at a global scale, with stringent service objectives and requirements to be available during maintenance and failures, is even more so. At Google, we operate such a network. This paper details our experience applying formal methods to some of the networking protocols that are developed and maintained by in-house engineers. These protocols centrally route network traffic to respond to changes in demand, react to network failures, and allow for maintenance and upgrades. We used formal methods to target a class of bugs stemming from unclear specifications, unintended system interactions, and logical errors at the specification level. We show how we modeled our protocols using an off-the-shelf model checker and a custom harness to scale the model horizontally. We were able to recreate several recent bugs and verify that the fixes implemented were correct. Finally, we present a method called state projection that we used to increase confidence in the coverage of our models, which we added to the TLC model checker for TLA+. We created 7 different TLA+ models and showed that they were effective at recreating bugs and verifying our fixes to those bugs. View details
Large-scale, interpretable gene regulatory network inference through biologically informed matrix factorization
Soel Micheletti
Viola Fanfani
Julia Vogt
John Quackenbush
Jonas Fischer
Alexander Marx
Panagiotis Mandros
bioRxiv (2026)
Preview abstract Gene regulatory networks (GRNs) provide a mechanistic framework for understand- ing how transcription factors coordinate gene expression to establish cellular identity and phenotype. Methods that integrate gene expression with motif-derived regulatory priors and other sources of biological information have substantially advanced gene regulatory network inference by reconstructing condition-specific regulatory architec- ture. These approaches estimate the evidence supporting regulatory interactions and have proven remarkably successful in a wide range of biological applications. A com- plementary view of regulatory networks, however, seeks to estimate the effect of those interactions on gene expression itself, providing a framework in which regulatory edges can be interpreted as activating or inhibitory influences on transcription. We developed Giraffe, a biologically informed matrix factorization framework that jointly estimates transcription factor activities and gene regulatory networks by integrating gene expression, motif-based regulatory priors, and transcription factor protein-protein interactions. Giraffe estimates signed partial regulatory effects whose magnitude and sign can be interpreted as the strength and direction of transcriptional regulation. Building directly on the biological framework established by methods such as PANDA, Giraffe provides a complementary representation of gene regulatory networks that emphasizes mechanistic interpretation while remaining scalable, flexible, and computationally efficient. Across synthetic benchmarks, six human tissues, yeast transcription factor pertur- bation experiments, and liver hepatocellular carcinoma, Giraffe accurately recon- structs regulatory interactions while distinguishing activating from inhibitory regula- tion with high accuracy. The inferred networks recover known features of tissue-specific regulation, correctly classify regulatory effects in transcription factor perturbation ex- periments, and identify biologically coherent changes in regulatory programs associated with liver cancer. Together, these results demonstrate that estimating the direction of transcriptional regulation provides a complementary perspective on gene regulatory networks that facilitates biological interpretation and hypothesis generation. View details
Preview abstract In large-scale distributed enterprises, traditional Knowledge Management (KM) systems face a critical failure mode: static documentation cannot keep pace with evolving operational realities and regional nuances. This "knowledge latency" forces employees out of self-service workflows and into costly support ticketing queues. This paper introduces SENTINEL, a geo-contextual AI framework designed to shift enterprise support from reactive retrieval to proactive interception. The architecture employs a novel dual-engine system integrated into an omni-present interface. The first engine utilizes Large Language Models (LLMs) to conduct pre-emptive, historical case-grounded audits of documentation, generating a "Contextual Density" score that identifies friction zones. The second engine is an autonomous Retrieval-Augmented Generation (RAG) agent that surfaces in-situ via a location-intelligent assistant window, resolving queries in real-time. By functioning as a strategic "defensive barrier" at the point of origin, SENTINEL demonstrates how a proactive AI assistant can drive high-fidelity, in-situ case deflection. View details
Preview abstract Context: The cost of frontier large language model inference has fallen by two orders of magnitude since 2023, yet the techno-economic forces governing AI value capture remain poorly understood. No existing work provides a unified, multi-layer framework connecting hardware physics to commercial pricing to actuarial constraints. Objectives: This survey aims to establish that Generative AI (GenAI) monetization is structurally bound by five interdependent techno-economic layers: (1) the physical constraints of memory bandwidth and compute, (2) deflationary architectural innovations, (3) the algorithmic economics of inference-time compute, (4) the verification economics governing outcome-based pricing, and (5) the macro-legal realities of enterprise liability. Methods: We conduct a Multivocal Literature Review (MLR) adapting the PRISMA protocol, synthesizing peer-reviewed and grey literature sources—vendor documentation, SLAs, and API pricing data (2022–2026). Two reviewers independently screened all records (Cohen’s κ ≥ 0.81 across all decision stages). Results: We contribute four primary artifacts. First, the Viability Inequality, an analytical model formalizing the conditions under which outcome-based AI pricing is economically sustainable. Second, the Billing Fallacy: aggregate cost growth is driven by Agentic Recursion, not quadratic attention complexity. Third, the Verifiability Bifurcation: objective task domains enable outcome pricing, while subjective domains depend on proxy-based models. Fourth, the Multi-Layer Techno-Economic Taxonomy (M-TET), a unified five-layer framework mapping the full monetization stack from silicon-anchored token pricing through actuarial risk ceilings. Conclusion: GenAI monetization is not a commercial pricing exercise but a dynamic negotiation across hardware, algorithmic, economic, and actuarial layers. In subjective and hybrid task domains, the binding constraint on outcome-based pricing is the cost of verification, not generation. AI value capture depends on engineering low-cost, high-fidelity Verification Engines. View details
Preview abstract This talk addresses the challenges of operating Google's monitoring systems at scale, handling terabytes of telemetry data and preventing overload from diverse workloads. We'll explore how Google's internal client library and Monarch, its planet-scale time-series database, work together for cost-effective data collection. Key principles include a distributed push model, dynamic client-side data reduction, centralized retention, and periodic metric analysis. The session will then bridge these concepts to the open-source world, discussing our work with OpenTelemetry's OpAMP protocol to achieve similar scalable and efficient telemetry collection. Attendees will gain insights into adapting these principles for cost savings and learn about our collaboration with the OpAMP SIG to benefit the broader community. View details
Preview abstract Quantization methods have significantly improved the compute and memory efficiency of Large Language Model (LLM) training. However, existing approaches still rely on accumulating their updates into high precision: concretely, gradient updates must be applied to a high-precision weight buffer, known as \textit{master weights}. This buffer introduces substantial memory overhead, particularly for Sparse Mixture of Experts (SMoE) models, where model parameters and optimizer states dominate memory usage. In this work, we introduce the Error-Compensating Optimizer (ECO), which \textit{for the first time} enables the complete elimination of master weights by directly accumulating updates into quantized parameters, by leveraging existing optimizer states. ECO quantizes the weights after every gradient step and injects the resulting quantization error into the optimizer's momentum buffer, creating an error-feedback loop with zero additional memory overhead for quantization. Beyond its practical efficiency, ECO comes with theoretical guarantees. Specifically, under standard assumptions, naive master weight removal can lead to unbounded drift from the ideal parameter trajectory, whereas ECO provably bounds this drift, ensuring stable convergence. We validate ECO across a range of models, including small transformers (30M--800M), Gemma-3 1B, and an SMoE 2.1B model, using FP8 quantization. In all cases, ECO achieves near-lossless accuracy compared to high-precision baselines. For large SMoE models, ECO reduces memory usage by up to 25\%, establishing a new Pareto frontier for the trade-off between static memory and training loss. View details
Mobility-Embedded POIs: Learning What a Place Is and How It’s Used from Human Movement
Shushman Choudhury
Shang-Ling Hsu
Cyrus Shahabi
Forty-third International Conference on Machine Learning (2026)
Preview abstract Recent progress in geospatial foundation models (GeoFMs) has highlighted the importance of learning general-purpose representations for real-world locations, particularly Points of Interest (POIs) where human activity concentrates. Yet, ex- isting POI representations remain largely static, drawing from textual metadata (e.g., category labels, descriptions) and spatial attributes (e.g., coordinates, neigh- borhood context), all of which describe what a place is, but not how it is actu- ally used. We argue that human mobility provides a complementary and dynamic signal, capturing real-world visitation patterns that reveal how places function in practice. To this end, we introduce Mobility Embedded POIs (ME-POIs), a pretraining framework that learns POI representations directly from sequences of human visits. Each visit is encoded as a contextualized embedding that captures the POI’s static attributes as well as its temporal and sequential context, includ- ing when the visit occurs and which visits surround it. These visit embeddings are aligned with learnable POI embeddings via a contrastive objective, grounding POI representations in their real-world usage patterns. To address the long tail of sparsely visited POIs, we transfer visitation distributions from data-rich anchors to sparse locations, leveraging multi-scale spatial proximity to capture local and regional patterns, and functional similarity to enable transfer across semantically related POIs. We demonstrate the utility of ME-POIs for a set of automated map enrichment tasks, critical in geospatial intelligence. We show empirically that by embedding visitation dynamics, ME-POIs outperform text- and location-only baselines, proving that mobility-informed embeddings provide a stronger founda- tion for modeling place function and change. View details
×