An attacker's analysis of Anthropic's next-generation Constitutional Classifiers — the two-stage cascade, linear probe screening, 0.05% flag rate, 40x cost reduction, and what the absence of a universal jailbreak actually means.
A technical survey of every major agentic defense system published in 2025-2026 — SafeAgent, AgentSentry, CoopGuard, Constitutional Classifiers, RvB, FlowGuard, SHIELD, DynaTrust, TrajGuard, AegisLLM, Stable Agentic Control, and PocketAgents — with an attacker's perspective on each.
Deep technical analysis of FlowSteer — a prompt-only planning-time attack against multi-agent LLM systems that achieves +55% malicious success rate by manipulating workflow formation itself, completely bypassing post-hoc defenses like G-Safeguard and ARGUS.
I ran 12 agentic defense systems against a two-turn attack. Every one of them failed. Here's the vulnerability they all share and the architectural fix that closes it.
A position paper arguing that predictions used by learning-augmented algorithms should be treated as untrusted algorithmic inputs, not neutral estimates. Maps adversarial machine learning threat models into algorithm design and proposes benchmark dimensions for robust, security-aware online decision-making.