Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: pixabay.com
Larger models can be more robust in some settings, including resistance to deceptive prompts and efficiency during adversarial training, but scale alone is not a dependable substitute for data or defenses. Whether they need less data depends on what “robustness” means: studies find inconsistent out-of-distribution benefits, while data scarcity can substantially reduce large-model performance.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: Why scale may not reduce data needs
A more skeptical reading treats “larger models need less data” as an overgeneralization. Sample efficiency and robustness do not reliably move together across interventions or datasets, and large transformers can be especially data-inefficient when training data is scarce. Improvements may therefore depend more on data quality, task fit, adaptation, or explicit robustness training than on parameter count alone.
Deeper threads worth pulling on next.