Why single-prompt testing isn't enough
Most AI red-teaming practice was built for single-turn, single-model interactions: send an adversarial prompt, check the response. Agentic workflows break that model because the risk isn't only in what the model says—it's in what it does with the tools, memory and state it has access to across multiple steps.
A model can pass every isolated safety check and still produce an unsafe outcome once it's given the ability to call APIs, write files, or trigger downstream actions in sequence.