FabScore: Fine-Grained Evaluation of Fabrications in Automated AI Research

James Xu Zhao
See-Kiong Ng
Dongfu Jiang
Bryan Hooi
Qianyun Guo
Hui Chen
Pang Wei Koh
Muhao Chen
Yiwei Wang
2026
Google Scholar

Abstract

In automated AI research, scientific rigor is not merely a matter of producing coherent papers and executable code; rather, it fundamentally requires that the experimental results claimed in a paper are faithfully supported by the accompanying implementation and verifiable through actual execution logs. When this alignment breaks down, AI-generated research may contain fabrications: discrepancies between the methods described in the paper and what the code actually implements, or between the reported results and those obtained by running the code. This motivates a systematic evaluation of fabrications that systematically examines the consistency between a paper's claimed experimental outcomes and its accompanying code files. In automated AI research, scientific rigor is not merely a matter of producing coherent papers and executable code; rather, it fundamentally requires that the experimental results claimed in a paper are faithfully supported by the accompanying implementation and verifiable through actual execution logs. When this alignment breaks down, AI-generated research may contain fabrications: discrepancies between the methods described in the paper and what the code actually implements, or between the reported results and those obtained by running the code. This motivates a systematic evaluation of fabrications that systematically examines the consistency between a paper's claimed experimental outcomes and its accompanying code files. This is a placeholder. This is a placeholder. This is a placeholder.
×