Yossi Matias

Yossi Matias

Yossi Matias is Vice President, Google, and the Head of Google Research.

For the past two decades, Yossi Matias has been a key figure in Google’s executive technology leadership. Today, as GM of Google Research, Yossi leads pioneering global research teams at the forefront of science and technology, driving breakthrough research through the magic cycle to real-world impact on diverse areas. From foundational machine learning, algorithms, and computing systems to AI for societal impact in health, education and planetary intelligence; and from foundational advancements in generative AI to quantum computing and advancing science.

Recent research breakthroughs include speculative decoding, generative UI, genAI factuality, AI Co-Scientist, Earth AI, Quantum Echo algorithm, DeepSomatic, MedGemma, Flood Forecasting, FireSat, LearnLM, Empirical Research Assistance, Gemini for Science.

Yossi was previously on Google Search leadership for over a decade, driving strategic features and technologies (including Google Trends and Autocomplete) and pioneered Conversational AI innovations to help transform the phone experience and help remove barriers of modality and languages. He was also the founding lead of Google center in Israel and supported other global sites. During his tenure at Google Yossi founded and spearheaded initiatives such as Google's AI for Social Good, Crisis Resilience, Google for Startups Accelerator, social and cultural initiatives seeding Google Arts & Culture, and programs fostering startups, sustainability, and AI literacy for youth.

Prior to Google Yossi was on the Computer Science faculty at Tel Aviv University, a visiting professor at Stanford, and a Research Scientist at Bell Labs. He’s published over 200 papers and is the inventor of over 80 patents. He pioneered some of the early technologies for internet privacy, contextual search, and the effective analysis of Big Data. He is a recipient of the Gödel Prize, an ACM Fellow, and a recipient of the ACM Kanellakis Theory and Practice Award for seminal work on streaming algorithms, data sketches, and large-scale data analytics

Yossi has a track record of impact-driven breakthrough research and innovation, and extensive product leadership, transforming products and advancing AI to help address global challenges.

-----------------------------------------------

More about work Yossi has been leading in recent years:

Search: Google Search leadership for over a decade included Autocomplete, Google Trends, Search Console, and Search experiences in weather, sports, dictionary and more.

Generative UI: pioneering work on generative UI, enables AI models to create immersive experiences and interactive tools and simulations, all generated completely on the fly for any prompt, launched in Gemini app and Google Search AI Mode.

Speculative Decoding: extensive research on efficiency for generative AI, including Speculative Decoding which has impact across the industry.

Generative AI Factuality: extensive research work on consistency and multi-modal factuality, published benchmarks (TRUE, FACTS) and leading to Double Check.

Conversational AI: Pioneering innovations in conversational AI as the "ultimate user interface" towards ambient intelligence and new experiences. Google Duplex - from a defining moment in AI to helping people and businesses well over 1 trillion times to getting things done faster directly from Search. Helping transform the phone experience (Call Screen, Hold for Me) and helping remove barriers of modality and languages making content and communication more universally available (Live Caption, Live Relay, Euphonia, Read Aloud).

Google Earth AI bringing together geospatial models Gemini advanced reasoning to Google Earth AI, to help tackle the planet's most critical needs.

Health AI: Work on Google’s Health AI is driving AI research to help transform healthcare from innovation to impact and help make healthcare more accessible for everyone, with multiple breakthroughs including Med-PaLM(1, 2), MedGemini(3, 4), AMIE(5,6), MedGemma.

Accelerating Scientific discovery: Using AI to drive scientific research, powering science breakthroughs with greater real-world benefit. Accelerating scientific discovery with AI Co-Scientist (Nature paper) and Empirical Research Assistant (Nature paper) powering Gemini for Science.

Climate Resilience: Leadership of AI for Climate and Sustainability, work on climate crisis mitigation (Greenlight,Contrails) as well as climate crisis nowcasting and forecasting - with leadership work on Google’s Crisis Response initiative (SOS alerts, flood forecasting, wildfire detection, FireSat). From research to climate resilience, Google Earth AI.

Education: Developing LearnLM. - enabling the best LLM for education tasks. Reimagening the textbook with Learn Your Way.

Special initiatives: Founding lead of Google’s AI for Social Good, Google for Startup Accelerator (from supporting early stage entrepreneurs to particular focus on AI & ML and to focus on Sustainability, expanding to regional programs and globally). Founding lead of Mind the Gap and Hello Tech. Pioneered an initiative of bringing online hundreds of heritage collections (including the Dead Sea Scrolls and the Nelson Mandela archive), and helped establish Google’s Arts and Culture.

Global sites: Founded and has lead Google’s center in Israel, through growth to over 2500 on staff, and founding lead of Campus TLV. Also supported Google’s growth (4X) in Bangalore, India, and oversaw Google’s Expanding Research Center in Africa, innovations for Africa and the world, initiating AI Community Center in Accra and supporting the future of AI Research in Africa and globally..

Additional work: Sketches, streaming algorithms, approximate query answering (seminal work - AMS, Synopses, AQUA Project - see awards below), Privacy and Security - early work on privacy and personalization (see also NYTimes article) based on the novel Janus function, early lightweight security primitives, and foundations for BLE ephemeral IDs); Parallel computation (highly parallel randomized algorithms, parallel models, parallel scheduling..); Compression (LZ improvements, compression in networks,.. ) and more (see publications).

Awards: Yossi is an ACM Fellow for contributions to the analysis of large data sets and data streams. His foundational work on data streams, data synopses and sketches, motivated by computational challenges in what was then the world’s largest data warehouses, was recognized with the Gödel Prize in Theoretical Computer Science, and with the ACM Kanellakis Theory and Practice Award for the instrumental role it played "in the development of the field of streaming algorithms, which is one of the most prolific and highly regarded areas of data management research" and its broad applicability to large-scale data analytics.

Authored Publications
Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
Towards expert-level medical AI for real-time video consultations
Mahvish Nagda
Jihyeon Lee
Matthew Thompson
CJ Park
Tim Strother
Roma Ruparel
Teya Bergamaschi
Suhana Bedi
Meet Shah
Pavel Dubov
Toshiyuki Fukuzawa
Sam Schmidgall
Craig Schiff
Joseph Xu
Aliya Rysbek
Yana Lunts
Jan Freyberg
Rebecca Hemenway
David Racz
Carey Radebaugh
Joelle Barral
Kavi Goel
Kat Chou
James Manyika
Gregory Wayne
Yun Liu
Ethan Goh
Christina Chen
Ryutaro Tanno
arXiv (2026)
Preview abstract Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility, but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice. View details
An AI system to help scientists write expert-level empirical software
Eser Aygün
Anastasiya Belyaeva
Gheorghe Comanici
Hao Cui
Renee Johnston
Zahra Shamsi
David Smalling
James Thompson
Sarah Martinson
Lai Wei
Yuchen Zhou
Qian-Ze Zhu
Matthew Abraham
Erica Brand
Anna Bulanova
Jeffrey Cardille
Chris Co
Scott Ellsworth
Grace Joseph
Malcolm Kane
Ryan Krueger
Johan Kartiwa
Jackson Cui
Paul Raccuglia
Julie Wang
Kat Chou
James Manyika
Lizzie Dorfman
Shibl Mourad
Nature (2026)
Preview abstract The cycle of scientific discovery is frequently bottlenecked by the slow, manual creation of software to support computational experiments. To address this, we present Empirical Research Assistance (ERA), an AI system that creates expert-level scientific software whose goal is to maximize a quality metric. The system uses a Large Language Model (LLM) and Tree Search (TS) to systematically improve the quality metric and intelligently navigate the large space of possible solutions. ERA achieves expert-level results when it explores and integrates complex research ideas from external sources. The effectiveness of tree search is demonstrated across a diverse range of tasks. In bioinformatics, ERA discovered 40 novel methods for single-cell data analysis that outperformed the top human-developed methods on a public leaderboard. In epidemiology, ERA generated 14 models that outperformed the CDC ensemble and all other individual models for forecasting COVID-19 hospitalizations. ERA also produced expert-level software for geospatial analysis, neural activity prediction in zebrafish, and numerical solution of integrals, and a novel rule-based construction for time series forecasting. By devising and implementing novel solutions to diverse tasks, ERA represents a significant step towards accelerating scientific progress. Keywords: Tree Search, Generative AI, Scorable Scientific Tasks, Empirical Software View details
Preview abstract Despite significant strides in factual reliability, errors -- often termed hallucinations -- remain a major concern for generative AI, especially as LLMs are increasingly expected to be helpful in more complex or nuanced setups. Yet even in the simplest setting -- factoid question-answering with clear ground truth-frontier models without external tools continue to hallucinate. We argue that most factuality gains in this domain have come from expanding the model's knowledge boundary (encoding more facts) rather than improving awareness of that boundary (distinguishing known from unknown). We conjecture that the latter is inherently difficult: models may lack the discriminative power to perfectly separate truths from errors, creating an unavoidable tradeoff between eliminating hallucinations and preserving utility. This tradeoff dissolves under a different framing. If we understand hallucinations as confident errors -- incorrect information delivered without appropriate qualification -- a third path emerges beyond the answer-or-abstain dichotomy: expressing uncertainty. We propose faithful uncertainty: aligning linguistic uncertainty with intrinsic uncertainty. This is one facet of metacognition -- the ability to be aware of one's own uncertainty and to act on it. For direct interaction, acting on uncertainty means communicating it honestly; for agentic systems, it becomes the control layer governing when to search and what to trust. Metacognition is thus essential for LLMs to be both trustworthy and capable; we conclude by highlighting open problems for progress towards this objective. View details
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
Guy Mor-Lan
Omer Goldman
Matan Eyal
Adi Mayrav Gilady
Sivan Eiger
Reut Tsarfaty
ACL (2026) (to appear)
Preview abstract Multilingual large language models (LLMs) have minimized the fluency gap between languages. This advancement, however, exposes models to the risk of biased behavior, as knowledge and norms may propagate across languages. In this work, we aim to quantify models' inter- and intra-lingual biases, via their ability to answer locale-ambiguous questions. To this end, we present LocQA, a test set containing 2,156 questions in 12 languages, referring to various locale-dependent facts such as laws, dates, and measurements. The questions do not contain indications of the locales they relate to, other than the querying language itself. LLMs' responses to LocQA locale-ambiguous questions thus reveal models' implicit priors. We used LocQA to evaluate 32 models, and detected two types of structural biases. Inter-lingually, we show a global bias towards answers relevant to the US-locale, even when models are asked in languages other than English. Moreover, we discovered that this global bias is exacerbated in models that underwent instruction tuning, compared to their base counterparts. Intra-lingually, we show that when multiple locales are relevant for the same language, models act as demographic probability engines, prioritizing locales with larger populations. Taken together, insights from LocQA may help in shaping LLMs' desired local behavior, and in quantifying the impact of various training phases on different kinds of biases. View details
A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
Peter Brodeur
Jacob M. Koshy
Khaled Saab
Ava Homiar
Roma Ruparel
Charles Wu
Ryutaro Tanno
Joseph Xu
Amy Wang
David Stutz
Hannah M. Ferrera
David Barrett
Lindsey Crowley
Jihyeon Lee
Spencer E. Rittner
Selena K. Zhang
Elahe Vedadi
Christine G. Kohn
Kavita Kulkarni
Vinay Kadiyala
Sara Mahdavi
Wendy Du
David Feinbloom
Renee Wong
Petar Sirkovic
Alessio Orlandi
Juro Gottweis
Joelle Barral
Kat Chou
James Manyika
Rob Fields
Jonathan X. Li
Marc L. Cohen
Adam Rodman
arXiv (2026)
Preview abstract Large language model (LLM)-based AI systems have shown promise for patient-facing diagnostic and management conversations in simulated settings. Translating these systems into clinical practice requires assessment in real-world workflows with rigorous safety oversight. We report a prospective, single-arm feasibility study of an LLM-based conversational AI, the Articulate Medical Intelligence Explorer (AMIE), conducting clinical history taking and presentation of potential diagnoses for patients to discuss with their provider at urgent care appointments at a leading academic medical center. 100 adult patients completed an AMIE text-chat interaction up to 5 days before their appointment. We sought to assess the conversational safety and quality, patient and clinician experience, and clinical reasoning capabilities compared to primary care providers (PCPs). Human safety supervisors monitored all patient-AMIE interactions in real time and did not need to intervene to stop any consultations based on pre-defined criteria. Patients reported high satisfaction and their attitudes towards AI improved after interacting with AMIE (p < 0.001). PCPs found AMIE's output useful with a positive impact on preparedness. AMIE's differential diagnosis (DDx) included the final diagnosis, per chart review 8 weeks post-encounter, in 90% of cases, with 75% top-3 accuracy. Blinded assessment of AMIE and PCP DDx and management (Mx) plans suggested similar overall DDx and Mx plan quality, without significant differences for DDx (p = 0.6) and appropriateness and safety of Mx (p = 0.1 and 1.0, respectively). PCPs outperformed AMIE in the practicality (p = 0.003) and cost effectiveness (p = 0.004) of Mx. While further research is needed, this study demonstrates the initial feasibility, safety, and user acceptance of conversational AI in a real-world setting, representing crucial steps towards clinical translation. View details
Preview abstract Although large language models have shown promise in diagnostic dialogue, their capabilities for effective management reasoning, including disease progression, therapeutic response and safe medication prescription, have remained underexplored. We have advanced the previously demonstrated diagnostic capabilities of the Articulate Medical Intelligence Explorer (AMIE) using a new large-language-model-based agentic system optimized for multivisit clinical management and dialogue. To ground the reasoning of AMIE in authoritative clinical knowledge, we leveraged the long-context capabilities of Gemini, combining in-context retrieval with structured reasoning to align its output with up-to-date clinical practice guidelines and drug formularies. In a randomized, blinded virtual Objective Structured Clinical Examination study, AMIE was compared to 21 primary care physicians (PCPs) across 100 multivisit case scenarios designed to reflect the guidance of the UK National Institute for Health and Care Excellence and BMJ Best Practice guidelines. AMIE was non-inferior to PCPs in management reasoning, as assessed by specialists, and scored better both with respect to preciseness of treatment and investigation, and in terms of its alignment with and grounding in clinical guidelines. To benchmark medication reasoning, we developed RxQA, a multiple-choice question benchmark that was derived from two national drug formularies (from the USA and UK) and validated by board-certified pharmacists. Although AMIE and PCPs both benefited from the ability to access external drug information, AMIE outperformed PCPs on higher-difficulty questions. Although further research will be needed before real-world translation of AMIE, its strong performance across evaluations marks a significant step towards use of conversational artificial intelligence as a tool in disease management. View details
Preview abstract Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself. View details
A unified acoustic-to-speech-to-language embedding space captures the neural basis of natural language processing in everyday conversations
Uri Hasson
Samuel A. Nastase
Harshvardhan Gazula
Aditi Rao
Tom Sheffer
Werner Doyle
Orrin Devinsky
aditi singh
Adeen Flinker
Patricia Dugan
Bobbi Aubrey
Sasha Devore
Daniel Friedman
Leonard Niekerken
Catherine Kim
Haocheng Wang
Zaid Zada
Gina Choe
Nature Human Behaviour (2025)
Preview abstract This study introduces a unified computational framework connecting acoustic, speech and word-level linguistic structures to study the neural basis of everyday conversations in the human brain. We used electrocorticography to record neural signals across 100 h of speech production and comprehension as participants engaged in open-ended real-life conversations. We extracted low-level acoustic, mid-level speech and contextual word embeddings from a multimodal speech-to-text model (Whisper). We developed encoding models that linearly map these embeddings onto brain activity during speech production and comprehension. Remarkably, this model accurately predicts neural activity at each level of the language processing hierarchy across hours of new conversations not used in training the model. The internal processing hierarchy in the model is aligned with the cortical hierarchy for speech and language processing, where sensory and motor regions better align with the model’s speech embeddings, and higher-level language areas better align with the model’s language embeddings. The Whisper model captures the temporal sequence of language-to-speech encoding before word articulation (speech production) and speech-to-language encoding post articulation (speech comprehension). The embeddings learned by this model outperform symbolic models in capturing neural activity supporting natural speech and language. These findings support a paradigm shift towards unified computational models that capture the entire processing hierarchy for speech comprehension and production in real-world conversations. View details
A personal health large language model for sleep and fitness coaching
Anastasiya Belyaeva
Zhun Yang
Nick Furlotte
Chace Lee
Erik Schenck
Yojan Patel
Jian Cui
Robby Bryant
Ryan Gomes
Allen Jiang
Roy Lee
Javier Perez
Jamie Rogers
Cathy Speed
Shyam Tailor
Megan Walker
Jeffrey Yu
Tim Althoff
Conor Heneghan
Mark Malhotra
Leor Stern
Shwetak Patel
Shravya Shetty
Jiening Zhan
Daniel McDuff
Nature Medicine (2025)
Preview abstract Although large language models (LLMs) show promise for clinical healthcare applications, their utility for personalized health monitoring using wearable device data remains underexplored. Here we introduce the Personal Health Large Language Model (PH-LLM), designed for applications in sleep and fitness. PH-LLM is a version of the Gemini LLM that was finetuned for text understanding and reasoning when applied to aggregated daily-resolution numerical sensor data. We created three benchmark datasets to assess multiple complementary aspects of sleep and fitness: expert domain knowledge, generation of personalized insights and recommendations and prediction of self-reported sleep quality from longitudinal data. PH-LLM achieved scores that exceeded a sample of human experts on multiple-choice examinations in sleep medicine (79% versus 76%) and fitness (88% versus 71%). In a comprehensive evaluation involving 857 real-world case studies, PH-LLM performed similarly to human experts for fitness-related tasks and improved over the base Gemini model in providing personalized sleep insights. Finally, PH-LLM effectively predicted self-reported sleep quality using a multimodal encoding of wearable sensor data, further demonstrating its ability to effectively contextualize wearable modalities. This work highlights the potential of LLMs to revolutionize personal health monitoring via tailored insights and predictions from wearable data and provides datasets, rubrics and benchmark performance to further accelerate personal health-related LLM research. View details
Towards accurate differential diagnosis with large language models
Daniel McDuff
Amy Wang
Karan Singhal
Yash Sharma
Kavita Kulkarni
Le Hou
Yong Cheng
Sara Mahdavi
Sushant Prakash
Anupam Pathak
Shwetak Patel
Ewa Dominowska
Juro Gottweis
Joelle Barral
Kat Chou
Jake Sunshine
Nature (2025)
Preview abstract A comprehensive differential diagnosis is a cornerstone of medical care that is often reached through an iterative process of interpretation that combines clinical history, physical examination, investigations and procedures. Interactive interfaces powered by large language models present new opportunities to assist and automate aspects of this process. Here we introduce the Articulate Medical Intelligence Explorer (AMIE), a large language model that is optimized for diagnostic reasoning, and evaluate its ability to generate a differential diagnosis alone or as an aid to clinicians. Twenty clinicians evaluated 302 challenging, real-world medical cases sourced from published case reports. Each case report was read by two clinicians, who were randomized to one of two assistive conditions: assistance from search engines and standard medical resources; or assistance from AMIE in addition to these tools. All clinicians provided a baseline, unassisted differential diagnosis prior to using the respective assistive tools. AMIE exhibited standalone performance that exceeded that of unassisted clinicians (top-10 accuracy 59.1% versus 33.6%, P = 0.04). Comparing the two assisted study arms, the differential diagnosis quality score was higher for clinicians assisted by AMIE (top-10 accuracy 51.7%) compared with clinicians without its assistance (36.1%; McNemar’s test: 45.7, P < 0.01) and clinicians with search (44.4%; McNemar’s test: 4.75, P = 0.03). Further, clinicians assisted by AMIE arrived at more comprehensive differential lists than those without assistance from AMIE. Our study suggests that AMIE has potential to improve clinicians’ diagnostic reasoning and accuracy in challenging cases, meriting further real-world evaluation for its ability to empower physicians and widen patients’ access to specialist-level expertise. View details
×