Module 08
OptionalSystem Prompt Hardening
This module is optional and it closes the cycle. A pentest that ends in a list of findings leaves the work half done. Here the output is an artefact you can deploy, with the evidence of how much it improved against the same battery.
Hardening blind is not hardening
The usual reaction to a finding is to add a line to the system prompt. Without measuring again, nobody knows whether that line closed the vector, moved it into another category, or broke the product's normal behaviour. Hardening without verification is a hypothesis.
Generic defences age
Defensive instructions that circulate publicly are in the training set of anyone who wants to route around them. Hardening has to start from your findings, not from a template.
Closing one vector can open another
An instruction that blocks system prompt leakage can make the model more rigid and more likely to hallucinate when it lacks context. Without a second run you cannot see it.
Utility is measured too
A system prompt that lowers risk at the cost of the product refusing to answer is not a fix. The comparison has to include expected behaviour.
How it works
The module runs after the first full campaign and uses its results as input.
system_prompt_leakage62%9%indirect_prompt_injection54%21%data_theft38%12%illegal_content27%6%behavioral_limits44%40%unresolved
The same battery, run twice. Cases where hardening did not move the result are reported all the same: those need a control outside the model.
Values illustrate the delivery format. The real numbers come from your campaign.
- 01
Input: the baseline
It starts from the set of attacks that penetrated, with their category, their severity and the exact behaviour they produced.
- 02
Hardened prompt generation
A reinforced version of the system prompt is built, aimed at the vectors that actually penetrated rather than at a generic list.
- 03
Second campaign, same battery
Exactly the same set of attacks runs again. Without that control the comparison means nothing.
- 04
Comparison by category
You get the delta on the global risk score and the breakdown per category, including the cases where hardening changed nothing.
- 05
Your decision
The prompt is suggested, with its evidence. Adopting it, adapting it or discarding it is your team's call.
What it delivers
- 01
The hardened system prompt
An artefact ready to deploy, with every section justified by the finding that motivated it.
- 02
Before and after comparison
The same attack set run twice, with the per-category delta and the change in the global risk score.
- 03
What was not resolved
The vectors that still penetrate with the hardened prompt. Those need a control outside the model, and it is better to know before deploying.
Why it closes the cycle
The 7ONE cycle is attack, score, report and remediate. This module is remediation applied to the layer that can be changed fastest, with the verification included in the same delivery.
Next step
Test your model before an attacker does
Let us start by defining the scope of the evaluation and agreeing the baseline for your risk score.