Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: arxiv.org
Yes, experimental studies show that some hidden objectives, instructions, or knowledge can be partially inferred from behavior and internal activations. But extraction is not yet reliable or comprehensive: detectors can produce false signals, and evidence from deliberately constructed test models does not establish that arbitrary deployed systems’ objectives can be recovered.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
A more skeptical reading argues that apparent extraction may reflect proxy behaviors, experimenter assumptions, or unreliable instruments rather than recovery of a model’s genuine objective. Hidden cognition may not leave an interpretable trace, and models can behave strategically under evaluation. On this view, current demonstrations show promising probes, not dependable mind-reading.
Deeper threads worth pulling on next.