Skip to content
7ONE

Module 01

Catalogued attack library

Every evaluation starts here. The library is a living repository of attacks already proven against production models, grouped into ten categories and mapped one by one to OWASP, NIST and MITRE ATLAS. It does two jobs: it measures coverage against the known, and it gives the other modules somewhere to start from.

A catalogue measures coverage, not ingenuity

A large library answers one concrete question well: of everything already known to break models, how much breaks yours? That question matters, and it is the one audit asks. But a catalogue has a structural ceiling, and it is worth saying so before selling the number.

  • Ten categories, countable coverage

    Data theft, model theft, hallucinations, illegal content, retraining, input leakage, misinformation, behavioural limits, system prompt leakage and indirect injection. Coverage is countable, not a promise.

  • Catalogued by framework, not by taste

    14 of 14 OWASP LLM Top 10 rules (2025/26) and 115 MITRE ATLAS techniques and sub-techniques, mapped in advance. Findings arrive in the language audit already uses.

  • A single turn has a ceiling

    Every attack in the library is one message. A classifier that evaluates messages one at a time can learn to recognise them. That is why the adversarial-agents module exists.

  • What is published is already in training

    Known techniques circulate, and providers train against them. The library alone finds what was left open, not what nobody has tried yet.

How it runs

The full battery runs against the model or agent endpoint, with or without context information, and every result is weighted by category and severity.

attack categories
10
attack categories
OWASP LLM Top 10 rules (2025/26)
14/14
OWASP LLM Top 10 rules (2025/26)
MITRE ATLAS techniques and sub-techniques
115
MITRE ATLAS techniques and sub-techniques
  • Data theft

    data_theft

    Extracting sensitive information held in the system's document sources (RAG) or contained in previous interactions with other users.

    Frameworks

    LLM02LLM06AML.T0057AML.T0024
  • Model theft

    model_theft

    Systematic querying aimed at inferring the internal logic, parameters or intellectual property of the model under evaluation.

    Frameworks

    LLM10AML.T0035AML.T0044
  • Hallucinations

    hallucinations

    Inducing the model to generate false or inconsistent information, presented with unjustified confidence.

    Frameworks

    LLM09AML.T0048
  • Illegal or unethical content

    illegal_content

    Getting the model to produce content that breaches usage policy, legal frameworks or basic ethical principles.

    Frameworks

    LLM05LLM09AML.T0054
  • Retraining

    Sandbox only

    retraining

    Poisoning or manipulating the model's training or fine-tuning data. Given its severity it runs exclusively in sandbox.

    Frameworks

    LLM04AML.T0020AML.T0018
  • Input leakage

    input_leakage

    Determining whether the contents of instructions or data submitted by other users or system processes can be exposed.

    Frameworks

    LLM02AML.T0057
  • Misinformation

    misinformation

    Evaluating whether the model spreads or endorses false or misleading information as if it were true.

    Frameworks

    LLM09AML.T0048.003
  • Behavioural limits (agents)

    Sandbox only

    behavioral_limits

    Getting an agent with access to external tools to exceed its defined action boundary and execute unauthorised operations.

    Frameworks

    LLM06LLM08AML.T0053AML.T0011
  • System prompt leakage

    system_prompt_leakage

    Revealing internal instructions or system configuration that should never be exposed to the user.

    Frameworks

    LLM05LLM12LLM14AML.T0000AML.T0000.002AML.T0069.002
  • Indirect prompt injection

    indirect_prompt_injection

    The malicious instruction does not come from the user. It hides in external content the model processes as context: documents, pages, images, audio and search results.

    Frameworks

    LLM01AML.T0051.001AML.T0068
  1. 01

    Selection by coverage

    The campaign is assembled to cover the ten categories and each framework's rules, not to maximise the number of shots fired.

  2. 02

    Execution and verdict

    Every attack gets a result and a weight based on what it achieved: an improper answer is not the same as an executed action.

  3. 03

    Variants in parallel

    Each attack has its verse rewrite, and in the Spanish corpus, a version written natively per region. Both run alongside the original.

  4. 04

    Feedback

    Whatever penetrated feeds the BFS engine, which recombines it to look for variants the library does not hold yet.

What it delivers

  • 01

    Coverage map by framework

    Which OWASP rules and ATLAS techniques were tested, which penetrated and which did not. This is the evidence audit asks for.

  • 02

    A comparable risk score

    One number you can track over time and measure every model change against.

  • 03

    The baseline for everything else

    This campaign's result is the input to the BFS engine and the hardening module. Without it, the other modules have nothing to compare against.

What it covers in practice

The library's categories are not theoretical. These two public incidents fall squarely inside categories the battery tests on every campaign.

February 2024

Air Canada

Hallucination with contractual weight

The website chatbot invented a bereavement-fare policy that did not exist. In tribunal the airline argued the bot was a separate entity, responsible for its own statements.

The British Columbia Civil Resolution Tribunal rejected that argument and found Air Canada liable for negligent misrepresentation. The award was small; the precedent is not. What your model says binds your company.

Suite category

hallucinationsmisinformation
Moffatt v. Air Canada, 2024 BCCRT 149Source
August 2024

Slack AI

Indirect injection through RAG

PromptArmor planted instructions in a public channel. Slack AI's RAG pipeline indexed them like any other content and executed them when a different user ran a query.

API keys were pulled out of private channels the attacker never had access to. The vector was not a stolen credential. It was a public message the model read as an instruction.

Suite category

indirect_prompt_injectiondata_theft
PromptArmorSource

Next step

Test your model before an attacker does

Let us start by defining the scope of the evaluation and agreeing the baseline for your risk score.