Offline Retrieval Evaluation Without Evaluation Metrics

Andres Ferraro

Fernando Diaz

Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (2022)

Download Google Scholar

Abstract

Offline evaluation of information retrieval and recommendation has traditionally focused on distilling the quality of a ranking into a scalar metric such as average precision or normalized discounted cumulative gain. We can use this metric to compare multiple systems’ performance on the same query or user. Although evaluation metrics provide a convenient summary of system performance, they can also obscure subtle behavior in the original ranking and can carry assumptions about user behavior and utility not supported across retrieval scenarios. We propose recall-paired preference (RPP), a metric-free evaluation method based on directly comparing ranked lists. RPP simulates multiple user subpopulations per query and compares systems across these pseudo-populations. Our results across multiple search and recommendation tasks demonstrate that RPP substantially improves discriminative power while being robust to missing data and correlating well with existing metrics.

Research Areas

Information Retrieval and the Web

Defining the technology of today and tomorrow.

Philosophy

People

Teams

AI/ML Foundations  & Capabilities

Algorithms & Optimization

Computing Paradigms

Responsible Human-Centric Technology

Science & Societal Impact

Projects

Publications

Resources

Shaping the future, together.

Student programs

Faculty programs

Conferences & events

Offline Retrieval Evaluation Without Evaluation Metrics

Abstract

Research Areas

Learn more about how we conduct our research

Defining the technology of today and tomorrow.

Philosophy

People

Teams

AI/ML Foundations & Capabilities

Algorithms & Optimization

Computing Paradigms

Responsible Human-Centric Technology

Science & Societal Impact

Projects

Publications

Resources

Shaping the future, together.

Student programs

Faculty programs

Conferences & events

Offline Retrieval Evaluation Without Evaluation Metrics

Abstract

Research Areas

Learn more about how we conduct our research

AI/ML Foundations  & Capabilities