Module 02
Adversarial agents
Not every attack fits in one message. The techniques that work best today are multi-turn: they escalate gradually, lean on what the model itself has already said, and reach the objective without any single message asking for anything forbidden. This module is the adaptive half of the suite.
Defences read messages, attacks use conversations
Almost all protection deployed today evaluates the current turn: an injection classifier, a pattern list, a judgement on the incoming message. An attack that spreads its intent across eight turns gives that classifier nothing to flag, because in no single turn is there anything to flag.
In isolation, every turn is legitimate
Asking about a compound's history, then its industrial synthesis, then the detail of one step: no single question is the one a filter is watching for.
The model leans on itself
Crescendo and echo chamber exploit the same thing: once the previous answer is in context, the model treats it as an accepted premise and contradicting itself becomes much harder.
Adaptation cannot be precomputed
Turn four depends on what turn three answered. No library can hold, in advance, the conversation needed against your configuration.
The published figures are high
Crescendo reports 98% success against GPT-4. Foot-in-the-door, 94% on average across seven models. These are not marginal lab cases.
How we test it
A 7ONE agent takes an objective from the library and works out how to reach it through conversation, choosing its technique from how your model responds.
crescendo · illustrative run
Filter verdict, turn by turnThe filter is right on all four turns: none of them, on its own, asks for anything forbidden. The intent is distributed, and only appears if you read the whole conversation.
- 01
Objective and turn budget
We fix which behaviour to provoke and the maximum number of turns. The objective comes from the catalogue; the route does not.
- 02
Technique selection
The agent picks among the available families depending on the surface: gradual escalation, incremental commitment, context poisoning, evaluator framing, narrative distraction or long-context saturation.
- 03
Adaptive conversation
Each turn is written after reading the previous answer. If the model refuses, the agent reframes rather than repeats; if it partially concedes, the agent advances from there.
- 04
Locating the break point
We record at which turn the boundary gave way and with which rewording, because that is what has to be defended afterwards.
- 05
Isolated execution
Anything that can mutate real state runs exclusively in sandbox, like the rest of the suite.
What it delivers
- 01
The full transcript
The entire conversation, turn by turn, with the exact point where the model gave way. That is what makes the finding reproducible.
- 02
Depth to break point
How many turns your model holds, per technique and per category. A model that survives ten turns is in a different position from one that folds at three.
- 03
Remediation built for multi-turn
Controls that evaluate the conversation rather than the message: intent-drift tracking, per-turn revalidation and limits on accumulated context.
Rationale
The six techniques this module runs are public, documented, and each reports its own success rate. The suite runs them against your model instead of leaving you to guess how it would do.
Next step
Test your model before an attacker does
Let us start by defining the scope of the evaluation and agreeing the baseline for your risk score.