What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation

Vitaly Feldman

Chiyuan Zhang

2020

Download Google Scholar

Abstract

Deep learning algorithms are well-known to have a propensity for fitting the training data very well including outliers and mislabeled data. Such memorization of training data has attracted significant research interest but has not been given a compelling explanation so far. A recent work proposes a theoretical explanation for this phenomenon based on a combination of two insights (Feldman, 2019). First, natural image and data distributions are (informally) known to be long-tailed, that is have a significant fraction of rare and atypical examples. Second, in a simple theoretical model such memorization is necessary for achieving close-to-optimal generalization error when the data distribution is long-tailed. However, no direct empirical evidence for this explanation or even an approach for obtaining such evidence were.
In this work we design experiments to test the key ideas in this theory. The experiments require estimation of the influence of each training example on the accuracy at each test example as well as memorization values of training examples. Estimating these quantities directly is computationally prohibitive but we show that closely-related subsampled influence and memorization values can be estimated much more efficiently. Our experiments demonstrate the significant benefits of memorization for generalization on several standard benchmarks. They also provide quantitative and visually compelling evidence for the theory put forth in (Feldman, 2019).

Research Areas

Machine Intelligence

Defining the technology of today and tomorrow.

Philosophy

People

Research areas

Foundational ML & Algorithms

Computing Systems & Quantum AI

Science, AI & Society

Projects

Publications

Resources

Shaping the future, together.

Student programs

Faculty programs

Conferences & events

What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation

Abstract

Research Areas

Meet the teams driving innovation