Rabbit hole · 3 connected questions
How do verification procedures themselves become attackable or brittle across ML models and scientific claims, and what concrete defenses make verification robust rather than easily manipulable?
How these converge
All three topics point to the same concrete problem: the procedures used to check correctness — classifiers’ test inputs, formal verification processes, and independent replication workflows — are themselves an attack surface. Adversaries (or accidental quirks) can craft inputs, exploit realistic constraints, or game reporting and analysis choices to make systems appear correct when they are not, or to make correct systems fail. That shared pattern highlights a set of specific mechanisms (transferable perturbations, targeted manipulation of verification protocols, and fragile audit procedures) and demands targeted mitigations (hardened tests, adversarial-aware auditing, and procedural safeguards) rather than abstract calls for “more oversight.”
Where these converge
Verification as an explicit attack surface
All three items treat verification not as a neutral oracle but as something that can be probed and exploited. Adversarial examples are concrete inputs that exploit model decision boundaries to defeat tests; adversarial verification describes intentionally targeting the verification step (e.g., by manipulating validators or constraints); and replicability audits show that the process of checking results (choices of analyses, data access, and reimplementation) determines whether a claim survives scrutiny. The common mechanism is that the verification step has features an attacker can probe and abuse.
Transferability and exploitation of weak verification criteria
Adversarial examples often transfer across models, meaning a test tailored to one system can be bypassed by attacks crafted for another; similarly, weak or narrow verification criteria (e.g., relying on a single replication method, limited stress tests, or non-adversarial benchmarks) make systems appear robust when they are not. This point ties the technical phenomenon of transferable perturbations to the methodological risk in audits and verification processes: narrow tests yield brittle assurances.
Mismatch between formal guarantees and real-world exploitability
Adversarial verification debates whether to prioritize mathematical certification or realistic exploit-driven testing. That same tension appears in replication: formal checks (reanalysis of code) can miss problems an adversary would exploit in practice (selective reporting, environmental differences). The convergence is that guarantees or superficial checks can be misleading unless they account for realistic threat models and procedural constraints.
The chain
Keep going: open any topic above to find its own related questions.