Skip to the content

A rig for testing how well AI agents hold up

We built a lab where our own and our clients’ agents are run through a catalogue of attacks, with a regression suite kept alongside it.

Client
IT&SOFT internal project
Industry
Research
Year
2026
Duration
ongoing
The problem

Testing agents by hand is expensive and not reproducible. We needed a rig that runs the same set of attacks after every change of prompt or model.

Solution

What we did
and why we did it that way.

01

The attack catalogue

Direct and indirect injections, system prompt extraction, guardrail bypasses, tool abuse.

02

Indirect injection through documents

A separate set of documents with instructions hidden inside—the most productive class in our testing.

03

Regression in CI

The test set is handed to the client so they can run it on every model update.

04

An open method

We write the findings up on the blog with no reference to clients. What we find should get fixed across the industry; keeping it to ourselves buys us nothing.

Result

What changed
and what we measured it with.

2,400+
attacks in the catalogue
8 of 10systems tested gave up access to other people’s data
~40min
for a full run against one agent
3
write-ups published on the blog
Stack
PythonGarakPyRITClaudeGPTLangfuseMITRE ATLAS
Timeline

How it went
day by day and week by week.

  1. Stage 1

    Catalogue

    We collected the known attack classes and mapped them to MITRE ATLAS.

  2. Stage 2

    Automation

    Running them, collecting responses, picking out leads for manual work.

  3. Stage 3

    Regression

    Exporting the test set into a format that drops into someone else’s CI.

Screens

How it looks
in schematics.

These are screen schematics. We do not publish client interfaces without permission

Next

A similar problem
on your side?

Describe it in the brief. In working hours we come back with an estimate of time and cost within two hours.