Weighing mainstream and alternative accounts…
Two lenses on the same evidence, given equal space. Source weight and the primary source ratio show what each rests on.
What every lens accepts.
Specific positions people hold on this question. Say whether you agree, add evidence, or submit a view of your own.
Deeper threads worth pulling on next.
Investigated
Adversarial training generally seeks to reduce error under malicious or worst-case input changes, but several theoretical and empirical studies find that this can increase error on unperturbed inputs. The trade-off is not universal or fixed: it varies with data quality and amount, model size and architecture, perturbation type and strength, overparameterization, and the training method. Some methods and settings report improvements in both robustness and standard performance, including robust self-training and selected neural-network experiments, while application studies such as robot learning report net performance costs. The main disagreement is whether robustness–accuracy tension is a fundamental constraint in practical systems or mainly a consequence of particular data, models, attacks, and evaluation choices.
Two lenses on the same evidence, given equal space. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: What the evidence supports about the trade-off
The dominant research account is that adversarial robustness and ordinary accuracy can be competing objectives, especially when the threat model makes predictive features vulnerable. Theory identifies conditions producing a trade-off, and experiments observe it across models and applications. However, mainstream work also treats the trade-off as conditional rather than an unavoidable law: architecture, data, perturbation strength, and methods such as robust self-training can change the outcome.
0 agree · 0 disagree (50% agree)
Lens adapted to this topic: Why the trade-off may be avoidable or overstated
A serious dissenting view argues that broad claims of an inherent trade-off overgeneralize from particular threat models, datasets, hypothesis classes, and training procedures. Results showing robust self-training or other methods improving both metrics suggest that standard accuracy need not always be sacrificed. A further critical position questions whether adversarial-ML evaluations—particularly for large language models—define meaningful threats and measure progress rigorously enough to support strong conclusions.
0 agree · 0 disagree (50% agree)
What every lens accepts.
Specific positions people hold on this question. Say whether you agree, add evidence, or submit a view of your own.
How it works: Agree/disagree is about the view. Evidence is scored on helpfulness, verified primary sources, and flags. New submissions are reviewed.
No perspectives on record yet.
Every investigation starts with one voice. Be the first to put a viewpoint — and the evidence behind it — on the record.
Deeper threads worth pulling on next.