Andrew W. Senior

Andrew W. Senior

Andrew Senior received his PhD from the University of Cambridge, for his thesis "Offline cursive handwriting recognition with recurrent neural networks", having previously worked on speech recognition at LIMSI at the University of Paris XI. He joined IBM Research in 1994 where he worked in the areas of handwriting, audio-visual speech, face and fingerprint recognition as well as video privacy and visual tracking. In 2008 he taught at Columbia University before joining Google Research where he worked on speech recognition with deep neural networks and recurrent neural networks. He has coauthored a "Guide to Biometrics", and over seventy scientific papers; holds forty-six patents. His research interests range across deep learning, speech, computer vision and visual art.
Authored Publications
Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
Reinforcement Learning Control of Quantum Error Correction
Cameron Maxfield
Guifre Vidal
Bob Buckley
Jonathan Waltz
Christopher Wood
Reza Molavi
John Mark Kreikebaum
Rajeev Acharya
David Sobel
Abeer Vaishnav
Ali Hadjikhani
Ryuho Kudo
Wendy Leung
Brett Buchea
Ningfeng Zhu
Shirin Montazeri
Jamie Yao
Bicheng Ying
Eric Mascot
Lenny Fuste
Zhenjie Zou
Rodrigo Cortinas
Matt Lloyd
Clarke Smith
Kris Ottosson
Emma Ropes
Felix Borjans
Rebecca Potter
Sean Harrington
Jeremy Hilton
David Enriquez
Stephen Heslin
Paula Heu
Daniel Lundahl
Elliot Young
Alex Crook
Fedor Kostritsa
Roberto Rodriguez
Chia Ni
Kim Ming Lau
Priyanka Thiruraman
Martin Damyanov
Logan Oas
Dmitry Abanin
Oscar Higgott
Aaron Shorter
Steve Habegger
Aniket Maiti
Ryan Kaufman
Valerie Ehimhen
Sayra Alcaraz
Marcos Flores
Elizabeth Rossi
Aria Shahingohar
Dario Rosenstock
Travis Weidel
Steven Waltman
Kristi Wong
Murat Sarihan
Arun Kumar
Vladimir Shvarts
Matt Reagor
Alfredo Torres
Michael Qian
Anthony Megrant
Charles Neill
Christopher Hudspeth
Michael Hamilton
Bill Huggins
Laura De Lorenzo
Tan Ha
Ran Zhang
Dar Gilboa
Nicholas Bushnell
Sherman Peek
David Rhodes
Leigh Martin
Mike Shearn
Vlad Kurilovich
David Browne
Spencer Small
Brian Ballard
Will Oliver
Lior Ella
Orion Pritchard
Josh Cogan
Rachel Resnick
Dmitri Maslov
Jose Guerrero
Paul Masih Das
Theodore White
Helge Gehring
Nikita Astrakhantsev
Can Knaut
Maddy Woodson
Brooks Foxen
Frank Arute
Alejo Grajales Dau
Yaxing Zhang
Aaron Szasz
Alexander Lill
Justin Ledford
Xiaoxuan Jin
Andreas Kabel
Sid Madhuk
Orion Martin
Catherine Vollgraff Heidweiller
Gabrielle Roberts
Juan Campero
Juhwan Yoo
Robert Salazar
Michael Newman
Arpit Ranadive
James Goeders
William Giang
Gonzalo Garcia
Agnetta Cleland
Maddie Taylor
Dogan Timucin
Ross Alcaraz
Hui Kang
Johannes Bausch
William Courtney
Robert Gasca
Kevin Satzinger
Meghan Voorhees
Silas Chen
Laleh Beni
Andrew Dunsworth
Jamal Busnaina
Pavel Laptev
Kiseo Kang
Shannon Wang
Paul Donohoe
Paul Conner
Vadim Smelyanskiy
James Spencer
Benjamin Chiaro
Grayson Young
Tim Burger
ILYA Drozdov
Peter Brooks
Jordan Suchard
Austin Fowler
Jimmy Chen
Alec Eickbusch
Francisco Heras
Hung-Shen Chang
Michael Broughton
Jeanne Hartshorn
Aviv Elbag
Martin Bigdeli
Tanner Hadick
Juan Atalaya
Mahmoud Elzouka
Melvin Mathews
Alex Sztein
Markus Ansmann
Pavol Juhas
Bryan Cochrane
Murray Ich Nguyen
Ashley Maloney
Will Livingston
Roberto Collins
Ming Li
Élie Genois
Jeremiah Ford
Christopher Garrick
Sayan Das
David Peterson
Eifu Tomita
Suhas Ganjam
Reno Hiltermann
Dylan Bowers
Bryce Kobrin
Yu Chen
Dan Riley
Leon Brill
Barrett Spells
Ben Curtin
Mike Hucka
Seneca Meeks
Sebastian Molina
Tiano Lange-Dei
Georg Aigeldinger
Ashley Huff
Wing Li
ZLATKO MINEV
Monica Hansen
Sebastian Schroeder
Walt Askew
Dietrich Graumann
Elias Portoles
Stijn de Graaf
Matt Cockrell
Harold Cook
Masaya Fukami
Ed Gonzales
Robert Geiger
Amir Karamlou
Loick Le Guevel
Ebrahim Forati
Justin Vargas
Doug Thor
Joel Grebel
Lucia De Rose
LILY LI
Dave Landhuis
Emma Rosenfeld
Hsin-Yuan (Robert) Huang
Kenny Lee
Shaun Jevons
Ping Yeh
Amira Abbas
Kunal Arya
Henry Schurkus
Hector Bates
Ganesh Ramachandran
Sergey Vdovichev
Brayden Ware
Max Schaefer
Cheng Xing
Brandon Langley
Anthony Cabrera
Michel Devoret
Cody Jones
Vlad Sivak
Mert Torunbalci
Ben Kueffler
Chaitali Joshi
Raja Gosula
Joy Lee
Alexander Korotkov
Thomas Edlich
Aditya Locharla
Nathan Lacroix
George Sterling
Hao Tran
Kostyantyn Kechedzhi
Trond Andersen
Alexandre Bourassa
Aaron Lunt
Alan Fung
Alex Pizzuto
Salvatore Mandra
Alex Greene
Vitali Kutsko
Kannan Sankaragomathi
Sofia Springer
Vinicius Ferreira
Raymond Orosco
Nature (2026)
Preview abstract The promise of fault-tolerant quantum computing is challenged by environmental drift that relentlessly degrades the quality of quantum operations. The contemporary solution, halting the entire quantum computation for recalibration, is unsustainable for the long runtimes of the future algorithms \cite{reiher2017elucidating,gidney2025factor}. We address this challenge by unifying calibration with computation, granting the quantum error correction process \cite{ryan2021realization,krinner2022realizing,sivak2023real,acharya2024quantum, bluvstein2024logical,bluvstein2025architectural,lacroix2025scaling} a dual role: its error detection events are not only used to correct the logical quantum state, but are also repurposed as a learning signal, teaching a reinforcement learning (RL) agent \cite{silver2017mastering,mnih2015human,levine2016end, shalev2016safe,ouyang2022training} to continuously steer the physical control parameters and stabilize the quantum system during the computation. We experimentally demonstrate this framework on a Willow superconducting processor, improving the logical stability of the surface code 3.5-fold against injected drift. By synthesizing our full suite of technological advances, including RL fine-tuning of the entire system and near-optimal decoding \cite{senior2025scalable, beni2025tesseract}, we achieve record performance of the surface and color codes, with average logical error per cycle of $\varepsilon_L=7.7\times10^{-4}$ and $\varepsilon_L=8.2\times10^{-3}$ respectively. Simulations of surface codes up to distance-15 with tens of thousands control parameters confirm the scalability of our RL framework, revealing an optimization speed that is independent of the system size. This work thus enables a new paradigm: a quantum computer that learns to self-improve directly from its errors and never stops computing. View details
Preview abstract Multichannel ASR systems commonly separate speech enhancement, including localization, beamforming and postfiltering, from acoustic modeling. In this paper, we perform multichannel enhancement jointly with acoustic modeling in a deep neural network framework. Inspired by beamforming, which leverages differences in the fine time structure of the signal at different microphones to filter energy arriving from different directions, we explore modeling the raw time-domain waveform directly. We introduce a neural network architecture which performs multichannel filtering in the first layer of the network and show that this network learns to be robust to varying target speaker direction of arrival, performing as well as a model that is given oracle knowledge of the true target speaker direction. % Next, we show how performance can be improved by \emph{factoring} the first layer to separate the multichannel spatial filtering operation from a single channel filterbank which computes a frequency decomposition. % We also introduce an adaptive variant, which updates the spatial filter coefficients at each time frame based on the previous inputs. % Finally we demonstrate that these approaches can be implemented more efficiently in the frequency domain. Overall, we find that such multichannel neural networks give a relative word error rate improvement of more than 5\% compared to a traditional beamforming-based multichannel ASR system and more than 10\% compared to a single channel waveform model. View details
Preview abstract Multichannel ASR systems commonly separate speech enhancement, including localization, beamforming and postfiltering, from acoustic modeling. In this chapter, we perform multi-channel enhancement jointly with acoustic modeling in a deep neural network framework. Inspired by beamforming, which leverages differences in the fine time structure of the signal at different microphones to filter energy arriving from different directions, we explore modeling the raw time-domain waveform directly. We introduce a neural network architecture which performs multichannel filtering in the first layer of the network and show that this network learns to be robust to varying target speaker direction of arrival, performing as well as a model that is given oracle knowledge of the true target speaker direction. Next, we show how performance can be improved by factoring the first layer to separate the multichannel spatial filtering operation from a single channel filterbank which computes a frequency decomposition. We also introduce an adaptive variant, which updates the spatial filter coefficients at each time frame based on the previous inputs. Finally we demonstrate that these approaches can be implemented more efficiently in the frequency domain. Overall, we find that such multichannel neural networks give a relative word error rate improvement of more than 5% compared to a traditional beamforming-based multichannel ASR system and more than 10% compared to a single channel waveform model. View details
WaveNet: A Generative Model for Raw Audio
Aäron van den Oord
Sander Dieleman
Karen Simonyan
Oriol Vinyals
Alexander Graves
Nal Kalchbrenner
Koray Kavukcuoglu
Arxiv (2016)
Preview abstract This paper introduces WaveNet, a deep generative neural network trained end-to-end to model raw audio waveforms, which can be applied to text-to-speech and music generation. Current approaches to text-to-speech are focused on non-parametric, example-based generation (which stitches together short audio signal segments from a large training set), and parametric, model-based generation (in which a model generates acoustic features synthesized into a waveform with a vocoder). In contrast, we show that directly generating wideband audio signals at tens of thousands of samples per second is not only feasible, but also achieves results that significantly outperform the prior art. A single trained WaveNet can be used to generate different voices by conditioning on the speaker identity. We also show that the same approach can be used for music audio generation and speech recognition. View details
Preview abstract We present a new procedure to train acoustic models from scratch for large vocabulary speech recognition requiring no previous model for alignments or boot-strapping. We augment the Connectionist Temporal Classification (CTC) objective function to allow training of acoustic models directly from a parallel corpus of audio data and transcribed data. With this augmented CTC function we train a phoneme recognition acoustic model directly from the written-domain transcript. Further, we outline a mechanism to generate a context-dependent phonemes from a CTC model trained to predict phonemes and ultimately train a second CTC model to predict these context-dependent phonemes. Since this approach does not require training of any previous non-CTC model it drastically reduces the overall data-to-model training time from 30 days to 10 days. Additionally, models obtain from this flatstart-CTC procedure outperform the state-of-the-art by XX-XX\%. View details
Preview abstract Both Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) have shown improvements over Deep Neural Networks (DNNs) across a wide variety of speech recognition tasks. CNNs, LSTMs and DNNs are complementary in their modeling capabilities, as CNNs are good at reducing frequency variations, LSTMs are good at temporal modeling, and DNNs are appropriate for mapping features to a more separable space. In this paper, we take advantage of the complementarity of CNNs, LSTMs and DNNs by combining them into one unified architecture. We explore the proposed architecture, which we call CLDNN, on a variety of large vocabulary tasks, varying from 200 to 2,000 hours. We find that the CLDNN provides a 4-6% relative improvement in WER over an LSTM, the strongest of the three individual models. View details
Large Vocabulary Automatic Speech Recognition for Children
Olivier Siohan
Melissa Carroll
Noah Coccaro
Qi-Ming Jiang
Françoise Beaufays
Interspeech (2015)
Preview abstract Recently, Google launched YouTube Kids, a mobile application for children, that uses a speech recognizer built specifically for recognizing children’s speech. In this paper we present techniques we explored to build such a system. We describe the use of a neural network classifier to identify matched acoustic training data, filtering data for language modeling to reduce the chance of producing offensive results. We also compare long short-term memory (LSTM) recurrent networks to convolutional, LSTM, deep neural networks (CLDNN). We found that a CLDNN acoustic model outperforms an LSTM across a variety of different conditions, but does not specifically model child speech relatively better than adult. Overall, these findings allow us to build a successful, state-of-the-art large vocabulary speech recognizer for both children and adults. View details
Learning acoustic frame labeling for speech recognition with recurrent neural networks
Hasim Sak
Ozan Irsoy
Alex Graves
Françoise Beaufays
Johan Schalkwyk
ICASSP (2015), pp. 4280-4284
Preview abstract This paper describes a series of experiments to extend the application of Context-Dependent (CD) long short-term memory (LSTM) recurrent neural networks (RNNs) trained with Connectionist Temporal Classification (CTC) and sMBR loss. Our experiments, on a noisy, reverberant voice search task, include training with alternative pronunciations and the application to child speech recognition; combination of multiple models, and convolutional input layers. We also investigate the latency of CTC models and show that constraining forward-backward alignment in training can reduce the delay for a real-time streaming speech recognition system. Finally we investigate transferring knowledge from one network to another through alignments View details
×