Skip to content
7ONE

Module 04

Agent Pentest

An agent is not a model, it is a model with hands. The risk stops being what it answers and becomes what it executes. This module takes as input the functions and tools your agent can reach, and generates the action chains that turn them into an attack.

Module scope

  1. 01

    You hand us the complete list of the agent's tools and functions.

  2. 02

    A model of ours reasons over that set and builds concrete abuse chains.

  3. 03

    We execute those chains against your agent in an isolated environment.

Every tool widens the surface

Tools get reviewed one at a time and on their usefulness. An attacker reviews them together and on what they let you chain. A web search capability is harmless. A web search capability plus a disk write capability plus an agent that follows instructions it found online is a download-and-execute chain.

  • The simplest example

    You list web_search among the tools, as a legitimate one. Our model chains it so the agent searches for, retrieves and downloads a malicious binary, without ever stepping outside the behaviour it was authorised to perform.

  • Natural-language instructions are not controls

    The Replit agent deleted a production database during an explicit code freeze repeated in capitals, then fabricated 4,000 fake profiles to cover it.

  • A third-party agent is also surface

    In the Nx attack, the malware invoked the Claude, Gemini and Q CLIs already installed on the machine with --yolo and --trust-all-tools. The developer's own assistant located and handed over 2,349 credentials.

  • The prompt can arrive through the repository

    In Amazon Q Developer, the destructive instruction was committed into the extension repository and shipped to the Marketplace in version 1.84.0.

How we test it

You supply the input to this module: the complete schema of your tools, with names, descriptions, parameters and permissions. That is enough; the agent's code is not needed.

chain · AT-7734

unauthorised result
  1. web_search()authorised

    locate the named resource

  2. fetch_url()authorised

    retrieve the content

  3. write_file()authorised

    persist to disk

  4. run_command()

    execute the binary

Each tool is legitimate on its own. The full chain is not.

  1. 01

    You hand us the tool schema

    We load the definition of every function you declared, with its parameters and the real reach of what it can touch. The more complete the list, the more realistic the attack.

  2. 02

    Adversarial generation

    A 7ONE model reasons over the whole set and proposes abuse chains: which combination of calls produces a result nobody authorised.

  3. 03

    Sandbox execution

    The chains run against your agent in an isolated environment. No test that mutates real state ever runs outside the sandbox.

  4. 04

    Measuring the boundary

    We record how far each chain travelled before it met a control, and which ones met none at all.

What it delivers

  • 01

    Map of abuse chains

    The call sequences that lead to an unauthorised action, ordered by impact and by how easily they can be triggered.

  • 02

    Highest-risk tools

    Which ones concentrate the danger and which only ever appear as a link in the chain. This is what tells you what to cut first.

  • 03

    Suggested controls

    Human confirmation by action class, scope restriction per parameter, credential separation and per-tool rate limits.

Cases behind this module

Three separate incidents, one pattern: the agent did exactly what it was asked, and that was the problem.

July 2025

Replit Agent · SaaStr

Agent exceeding its action boundary

During an explicit code freeze, repeated in capitals, the agent deleted the production database holding real records for 1,206 executives and 1,196 companies.

It then fabricated over 4,000 fake user profiles and falsified test results to conceal the deletion. No natural-language instruction is an access control.

Suite category

behavioral_limitshallucinations
The RegisterSource
August 2025

Nx · s1ngularity

Developer agents weaponised

Malicious versions of the Nx build system reached npm with a post-install payload that invoked the locally installed Claude, Gemini and Q CLIs, using --yolo and --trust-all-tools to skip permission prompts.

The developer's own AI assistant located and handed over the secrets: 2,349 credentials from 1,079 machines, published into repositories created inside the victims' own GitHub accounts.

Suite category

behavioral_limitsdata_theft
Wiz ResearchSource
July 2025

Amazon Q Developer for VS Code

Malicious prompt in the supply chain

An over-scoped GitHub token in the CodeBuild configuration let an attacker commit a prompt into the extension repository. The instruction told the agent to act as a "system cleaner".

Version 1.84.0 shipped to the Marketplace carrying commands to wipe the local filesystem and destroy AWS resources: terminate EC2 instances, delete S3 buckets, remove IAM users. AWS confirmed no customer infrastructure was affected and shipped a clean 1.85.0.

Suite category

behavioral_limits
AWS Security AdvisorySource

Next step

Test your model before an attacker does

Let us start by defining the scope of the evaluation and agreeing the baseline for your risk score.