Module 07
Poeticised attacks
This is the cleanest case of an uncomfortable principle: a model's defences are sensitive to form, not only to content. The same request, with the same intent, written in verse instead of prose, passes where the prose runs into a refusal.
The filter recognises form, not intent
Safety classifiers learn from examples. The examples they saw are written in the plain prose people use to ask for things. An identical intent dressed in metaphor falls outside that distribution and stops resembling what the classifier learned to reject.
The effect is measured
With the same intent in verse rather than prose, the average success rate rises from 8.08% to 43.07%. With hand-crafted poems it reaches 62%, and with some providers it exceeds 90%.
It cuts across providers
The measurement covers 25 frontier models from nine providers. This is not one model's quirk: it is a property of how alignment is trained.
One turn, no adaptation
Every attack in the study is single-turn. No conversation and no iteration required: rewriting is enough.
The effect is not uniform
Depending on the case, the verse version does better or worse than the original. That is why the suite keeps both and evaluates them in parallel, rather than replacing one with the other.
How it works
Poeticisation is a transformation applied to the corpus, not a separate attack. Every library entry has its pair.
Average attack success rate across 25 frontier models
- 01
Intent preservation
The variant has to ask for exactly the same thing. If the verse dilutes the objective, the result is useless as a measurement.
- 02
Variant generation
The attack is rewritten in verse while keeping the execution policy, which is the part that turns the imagery into an instruction.
- 03
Parallel evaluation
Original and variant run against the same target in the same campaign. The comparison is made per pair, not across a corpus average.
- 04
Input to the BFS engine
Poeticisation is one of the transformation operators the BFS engine can combine with other techniques to form the next generation.
What it delivers
- 01
The per-pair delta
For each attack, whether the verse version passed where the original did not. That is where you see whether your defence depends on form.
- 02
Form-sensitive categories
Which taxonomy categories break on rewriting. That is usually the sign that the control is a classifier rather than a real restriction.
- 03
Intent-oriented remediation
Controls that evaluate what the message asks for rather than how it is written, plus verification on the output instead of only on the input.
Rationale
The module rests on published, reproducible research, and the suite treats it as what it is: a measurable transformation, not a trick.
Next step
Test your model before an attacker does
Let us start by defining the scope of the evaluation and agreeing the baseline for your risk score.