ResearchRCTs for Human-AI Evaluation

First page of RCTs for Human-AI Evaluation: Methodological Challenges and Practical Solutions

AIES · 2026

RCTs for Human-AI Evaluation: Methodological Challenges and Practical Solutions

Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest

This report examines human uplift studies — randomized controlled trial (RCT)-style evaluations of artificial intelligence (AI) systems — that inform decisionmaking for AI. Drawing on interviews with 16 practitioners, the research identifies methodological challenges across the study life cycle and documents emerging solutions, including standardized task libraries and versioned evaluation infrastructure.