Abstract
Large Language Models (LLMs), such as OpenAI’s GPT-4 or Google’s Bard, have created unprecedented opportunities for analyzing and generating language data on a massive scale. Because language is core to all areas of psychology, this new technology holds the potential to transform the field. In this Review, we first present emerging applications of LLMs for psychological measurement, experimentation, and practice across areas of psychology. We show how LLMs can make certain tasks can be made profoundly more efficient (content analysis, questionnaire item generation, systematic reviews), while also unlocking entirely new research questions and methods across areas. Second, we review the foundations of LLMs. We explain how the way that LLMs were constructed (i.e. to predict the next word or utterance, not to reason like a human) is both the source of their strengths and their limitations. Third, we examine three major concerns with the application of LLMs to psychology, and how each might be overcome. Finally, we recommend several necessary investments that can help address these concerns. These include: (a) field-initiated “keystone” datasets; (b) increased standardization; and (c) investments in shared computing and analysis infrastructure to ensure that the future of LLM-powered R&D is equitable.