Skip to content
7ONE

Module 07

Poeticised attacks

This is the cleanest case of an uncomfortable principle: a model's defences are sensitive to form, not only to content. The same request, with the same intent, written in verse instead of prose, passes where the prose runs into a refusal.

The filter recognises form, not intent

Safety classifiers learn from examples. The examples they saw are written in the plain prose people use to ask for things. An identical intent dressed in metaphor falls outside that distribution and stops resembling what the classifier learned to reject.

  • The effect is measured

    With the same intent in verse rather than prose, the average success rate rises from 8.08% to 43.07%. With hand-crafted poems it reaches 62%, and with some providers it exceeds 90%.

  • It cuts across providers

    The measurement covers 25 frontier models from nine providers. This is not one model's quirk: it is a property of how alignment is trained.

  • One turn, no adaptation

    Every attack in the study is single-turn. No conversation and no iteration required: rewriting is enough.

  • The effect is not uniform

    Depending on the case, the verse version does better or worse than the original. That is why the suite keeps both and evaluates them in parallel, rather than replacing one with the other.

How it works

Poeticisation is a transformation applied to the corpus, not a separate attack. Every library entry has its pair.

Average attack success rate across 25 frontier models

Same intent in prose8.08%
Same intent in verse43.07%
Hand-crafted poem62%
  1. 01

    Intent preservation

    The variant has to ask for exactly the same thing. If the verse dilutes the objective, the result is useless as a measurement.

  2. 02

    Variant generation

    The attack is rewritten in verse while keeping the execution policy, which is the part that turns the imagery into an instruction.

  3. 03

    Parallel evaluation

    Original and variant run against the same target in the same campaign. The comparison is made per pair, not across a corpus average.

  4. 04

    Input to the BFS engine

    Poeticisation is one of the transformation operators the BFS engine can combine with other techniques to form the next generation.

What it delivers

  • 01

    The per-pair delta

    For each attack, whether the verse version passed where the original did not. That is where you see whether your defence depends on form.

  • 02

    Form-sensitive categories

    Which taxonomy categories break on rewriting. That is usually the sign that the control is a classifier rather than a real restriction.

  • 03

    Intent-oriented remediation

    Controls that evaluate what the message asks for rather than how it is written, plus verification on the output instead of only on the input.

Rationale

The module rests on published, reproducible research, and the suite treats it as what it is: a measurable transformation, not a trick.

Next step

Test your model before an attacker does

Let us start by defining the scope of the evaluation and agreeing the baseline for your risk score.