Weighing mainstream and alternative accounts…
Two lenses on the same evidence, given equal space. Source weight and the primary source ratio show what each rests on.
What every lens accepts.
Specific positions people hold on this question. Say whether you agree, add evidence, or submit a view of your own.
Deeper threads worth pulling on next.
Investigated
The sources broadly support human evaluation as important for assessing generated text and LLMs, but identify substantial limits in reliability, reproducibility, evaluator performance, and generalizability. Studies report that many human-evaluation protocols omit enough design information to prevent reliable repetition, while non-experts can struggle to distinguish machine-generated from human text without training. One perspective emphasizes better protocols, calibration, statistical analysis, and hybrid human–LLM evaluation; another stresses that evaluation embeds values and may not resolve governance or social-impact questions. The main disagreement is whether the central problem is chiefly methodological and repairable through improved evaluation design, or whether some limits arise from normative choices that evaluation alone cannot settle.
Two lenses on the same evidence, given equal space. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: How better study design can address known limits
The mainstream methodological view treats human evaluation as necessary but fallible measurement. Its evidence points to weaknesses in sampling, evaluator selection, instructions, metrics, reporting, reproducibility, and statistical analysis. It therefore favors explicit protocols, calibration, adjudication, uncertainty estimates, and carefully designed combinations of human and automated judgments rather than abandoning human evaluation.
0 agree · 0 disagree (50% agree)
Lens adapted to this topic: What methodological fixes may leave unresolved
The critical or outsider view argues that human-evaluation problems are not only technical defects. Evaluation choices can encode assumptions about which outcomes matter, whether context can be ignored, whether impacts can be quantified, and whether failures are comparable. On this account, better protocols are useful but cannot by themselves settle value conflicts, distributional effects, or broader governance questions.
0 agree · 0 disagree (50% agree)
What every lens accepts.
Specific positions people hold on this question. Say whether you agree, add evidence, or submit a view of your own.
How it works: Agree/disagree is about the view. Evidence is scored on helpfulness, verified primary sources, and flags. New submissions are reviewed.
No perspectives on record yet.
Every investigation starts with one voice. Be the first to put a viewpoint — and the evidence behind it — on the record.
Deeper threads worth pulling on next.