Contact
Let us define the scope of the evaluation
The first step is agreeing what gets evaluated and with how much context information. From there, the first full campaign establishes your risk baseline.
What we need to start
We do not need access to your code or your weights. The endpoint is enough for a black box evaluation, and whatever context information you choose to share enables the grey box attacks, which are the ones an insider or a well-informed attacker could attempt.
- 01
Model or agent endpoint
The interface the battery runs against. It can be the LLM directly or the full agent with its tools.
- 02
Test scope
Which categories are enabled and which stay restricted to the sandbox. Retraining and behavioural limits never run in production.
- 03
Optional context (grey box)
The name of the RAG store, the agent's tool schema, roles with authority. More context means a more realistic attack.
- 04
First campaign window
The full baseline. After that execution is continuous, and every change to the model triggers new tests.