Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: pixabay.com
Attributions are mainly evaluated by testing faithfulness: whether features given high importance actually influence the model’s prediction when perturbed. Researchers also use constructed ground truth, multiple metrics, robustness checks, and benchmark datasets, but no single test is decisive because each can introduce biases or measure a different notion of explanation.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: Challenges to standard attribution tests
A critical perspective argues that common evaluation scores may reward procedural artifacts rather than genuine explanations. Attribution is not uniformly defined, perturbation can create unrealistic inputs, and context attribution in language models may confuse information supplied in the prompt with knowledge already encoded in model weights. This view favors explicit sanity checks, provenance-aware benchmarks, and separate tests for soundness and completeness.
Deeper threads worth pulling on next.