Rabbit hole · 3 connected questions
How do choices and adaptivity in empirical attacks versus formal certification create a persistent evaluation gap in claimed model robustness?
How these converge
All three topics are about how we judge whether a model is robust to adversarial perturbations. PGD is the common, practical baseline for generating attacks but is highly sensitive to choices (loss, norm, step size, initialization). Adaptive PGD variants deliberately change those choices to find much stronger failures, showing that a non-adaptive PGD can give a misleading picture of robustness. Certified defenses sit on the other side: they provide attack-independent guarantees but make modeling, scalability, or tightness tradeoffs. Together they point to a single concrete problem: empirical robustness claims depend on which attacks and hyperparameters were tried, while certificates avoid that fragility at the cost of restricted threat models or loose bounds—so evaluating defenses requires understanding both adaptive attack capability and certification limits.
Where these converge
PGD is a fragile baseline whose outcomes depend on concrete attack choices
PGD's success as an evaluation tool hinges on specific implementation choices (loss function, norm, projection, initialization, iteration budget). That means a model that appears robust under one PGD setup may be vulnerable under another, making PGD-based claims contingent on those concrete parameters rather than on a single underlying property of the model.
Adaptive attack variants concretely expose weaknesses missed by standard PGD
Variants that change step sizes, objectives, projections, or norms systematically find stronger adversarial examples in many cases. This is not a philosophical point but a reproducible mechanism: adaptivity in the attack search algorithm can turn an apparently robust model into a broken one, demonstrating that non-adaptive PGD can overestimate robustness.
Certificates remove dependence on attack search but introduce their own concrete limits
Certified defenses provide formal guarantees that a prediction cannot change within a specified perturbation set, eliminating the need to enumerate attacks. However, those guarantees depend on the chosen threat model, the tightness of the certificate, and scalability to realistic inputs; in practice these limits mean certificates may not cover the same perturbations or datasets where adaptive empirical attacks are effective.
The core, actionable tension: empirical attack adaptivity versus certification scope
Putting the pieces together reveals a specific evaluation gap: adaptive empirical attacks demonstrate how evaluation protocol choices lead to false security claims, while certificates offer attack-independent assurance but only for particular, sometimes narrow, perturbation models and with potential looseness or scalability issues. Resolving real robustness requires methods that address both adaptive attack strength and certification applicability.
The chain
Keep going: open any topic above to find its own related questions.