When models disagree a lot—measured by high ensemble variance or low margin—it...
https://page-wiki.win/index.php/Week_1_Disagreement_Instrumentation_Checklist
When models disagree a lot—measured by high ensemble variance or low margin—it often signals tricky or risky inputs. By flagging the top 1-2% of these disputed cases for human review, teams can catch errors early