Improved Hierarchical Patient Classification with Language Model Pretraining over Clinical Notes

Jonas Beachey Kemp; Alvin Rishi Rajkomar; Andrew Mingbo Dai

Improved Hierarchical Patient Classification with Language Model Pretraining over Clinical Notes

Jonas Beachey Kemp

Alvin Rishi Rajkomar

Andrew Mingbo Dai

NeurIPS ML4H – extended abstract (2019)

Download Google Scholar

Abstract

Clinical notes in electronic health records contain highly heterogeneous writing styles, including non-standard terminology or abbreviations. Using these notes in predictive modeling has traditionally required preprocessing (e.g. taking frequent terms or topic modeling) that removes much of the richness of the source data. We propose a pretrained hierarchical recurrent neural network model that parses minimally processed clinical notes in an intuitive fashion, and show that it improves performance for discharge diagnosis classification tasks on the Medical Information Mart for Intensive Care III (MIMIC-III) dataset, compared to models that treat the notes as an unordered collection of terms or that conduct no pretraining. We also apply an attribution technique to examples to identify the words that the model uses to make its prediction, and show the importance of the words' nearby context.

Research Areas

Machine intelligence

Explore our many areas of focus

Building a collaborative ecosystem

Shaping the future together

Translating discovery into real-world impact

Improved Hierarchical Patient Classification with Language Model Pretraining over Clinical Notes

Abstract

Research Areas

Meet the teams driving innovation

Google AI

Google Cloud

Google DeepMind

Google Labs