Tradeoffs in Data Augmentation: An Empirical Study

Ekin Dogus Cubuk; Ethan S Dyer; Rapha Gontijo Lopes; Sylvia Smullin

Tradeoffs in Data Augmentation: An Empirical Study

Ekin Dogus Cubuk

Ethan S Dyer

Rapha Gontijo Lopes

Sylvia Smullin

ICLR (2021)

Download Google Scholar

Abstract

Though data augmentation has become a standard component of deep neural network training, the underlying mechanism behind the effectiveness of these techniques remains poorly understood. In practice, augmentation policies are often chosen using heuristics of distribution shift or augmentation diversity. Inspired by these, we conduct an empirical study to quantify how data augmentation improves model generalization. We introduce two interpretable and easy-to-compute measures: Affinity and Diversity. We find that augmentation performance is predicted not by either of these alone but by jointly optimizing the two.

Explore our many areas of focus

Building a collaborative ecosystem

Shaping the future together

Translating discovery into real-world impact

Tradeoffs in Data Augmentation: An Empirical Study

Abstract

Research Areas

Meet the teams driving innovation

Google AI

Google Cloud

Google DeepMind

Google Labs