Can large language models help predict results from a complex behavioural science study?

Open

Steffen Lippert, Anna Dreber, Magnus Johannesson, Warren Tierney, Wilson Cyrus-Lai, Eric Luis Uhlmann, Thomas Pfeiffer

2024 Royal Society Open Science Vol. 11 Issue 9 Article Cited by 11 Quartile

Abstract

We tested whether large language models (LLMs) can help predict results from a complex behavioural science experiment. In study 1, we investigated the performance of the widely used LLMs GPT-3.5 and GPT-4 in forecasting the empirical findings of a large-scale experimental study of emotions, gender, and social perceptions. We found that GPT-4, but not GPT-3.5, matched the performance of a cohort of 119 human experts, with correlations of 0.89 (GPT-4), 0.07 (GPT-3.5) and 0.87 (human experts) between aggregated forecasts and realized effect sizes. In study 2, providing participants from a university subject pool the opportunity to query a GPT-4 powered chatbot significantly increased the accuracy of their forecasts. Results indicate promise for artificial intelligence (AI) to help anticipate - at scale and minimal cost - which claims about human behaviour will find empirical support and which ones will not. Our discussion focuses on avenues for human-AI collaboration in science. © 2024 The Author(s).

Affiliations

Department of Economics, University of Auckland, Auckland, New Zealand; Department of Economics, Stockholm School of Economics, Stockholm, Sweden; Department of Economics, University of Innsbruck, Innsbruck, Austria; Organisational Behaviour Area/Marketing Area, INSEAD, Singapore, Singapore; Graduate School of Business, Stanford University, CA, United States; New Zealand Institute for Advanced Study, Massey University, Auckland, New Zealand