Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: ai-infrastructure.net
Model judges can be unreliable in several distinct ways: their scores may change with presentation details, contradict one another, or favor style over substance. More fundamentally, a judge can be highly repeatable yet measure the wrong construct, so apparent agreement and stable rankings do not by themselves establish valid evaluation. The practical risk is that benchmark or product improvements become optimization for judge preferences rather than genuine quality, especially when aggregation hides uncertainty or rubrics fail to explain verdicts.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: Validity, construct, and benchmark design
A deeper measurement-focused account argues that the central problem is not only whether judges are consistent, but whether they evaluate the intended construct at all. On this view, benchmark design, rubric coherence, aggregation, and hidden representations can create systematic errors that survive ordinary reliability checks. A stable score may therefore reflect a robust bias rather than dependable quality measurement.
Deeper threads worth pulling on next.