Abstract
Verification is rapidly emerging as a key primitive for both training and real-world deployment of large language models (LLMs). In practice, this often involves using imperfect LLM judges and reward models, as ground truth acquisition can be time-consuming and expensive. This naturally motivates asking if verification can be improved by leveraging multiple verifiers, and whether such ensembling can be done without itself requiring ground truth labels. We introduce Fully Unsupervised Score Ensembling (FUSE ), a method for ensembling verifiers that is (i) completely unsupervised and (ii) can operate on a query-conditional basis. Our main insight is that when conditional dependencies between verifiers are carefully controlled, certain spectral algorithms in the ensembling literature enjoy strong unsupervised guarantees. Despite requiring zero ground truth labels, FUSE matches or improves upon semi-supervised alternatives in test-time scaling experiments with diverse sets of generator models, verifiers, and benchmarks.