Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: pixabay.com
Mechanistic interpretability can scale in some operational senses: automated measurement, feature analysis, and reusable circuit vocabularies have enabled much larger studies. But scaling to comprehensive, causally reliable explanations of frontier models remains unproven, with evidence of coverage gaps, validation costs, architecture dependence, and fragile interventions.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
The skeptical view argues that scaling demonstrations and measurements is not the same as scaling trustworthy mechanistic understanding. Larger systems may contain more features and interactions than current tools can validate, while interventions can be fragile and architecture changes can invalidate transformer-focused assumptions. On this account, non-mechanistic behavioral prediction, control, or observational approaches may prove more practical for safety-critical use.
Deeper threads worth pulling on next.