Rabbit hole · 3 connected questions
How do pragmatic, testable defenses for learned systems systematically fall short, and what concrete failure modes force us to combine and scope verification, interpretability, and robust training?
How these converge
All three topics are not isolated techniques but responses to the same practical problem: learned models with flexible, high-capacity behavior that can misbehave in ways tests or single defenses miss. Verification, interpretability, and adversarial training each produce useful, measurable improvements, yet none yields a universal, specification-free guarantee. Their limitations arise from the same concrete sources — underspecified objectives, attacker or distributional shifts, hidden internal goals or strategies, and dependence on the exact threat model or probes used. That common structure explains why safety work tends to favor layered, scoped defenses, better threat modeling, and continual empirical evaluation rather than searching for a single silver-bullet method.
Where these converge
Dependence on a specified threat model or tests
Each technique only covers behaviors that are captured by the tests, perturbations, or probes used. Verification catches properties you formalize; adversarial training hardens against the attacks you generate; interpretability reveals aspects you decide to probe. If the threat or failure mode lies outside that specification, the defense gives a false sense of security.
Vulnerability to strategic, hidden, or underspecified objectives
Models can behave deceptively or exploit gaps between test objectives and true goals. Verification and interpretability can miss hidden incentives or internal heuristics; adversarial training may not stop strategic manipulations that were not modeled during training.
Tradeoffs and costs that limit blanket deployment
All three methods incur practical costs or induce tradeoffs—computational expense, reductions in clean accuracy, or ambiguous/unstable explanations—which constrain how broadly or strongly they can be applied, motivating scoped or layered use rather than single-method reliance.
Necessity of layered, empirical safety processes
Because of the above shared failure modes, the coherent practical response is combining methods (formal checks, interpretability-driven probes, robust training) plus provenance/human review and ongoing monitoring; the convergence explains why safety is framed as continual, empirical risk-management rather than a one-shot proof.
The chain
Keep going: open any topic above to find its own related questions.