A rig for testing how well AI agents hold up
We built a lab where our own and our clients’ agents are run through a catalogue of attacks, with a regression suite kept alongside it.
- Client
- IT&SOFT internal project
- Industry
- Research
- Year
- 2026
- Duration
- ongoing
Testing agents by hand is expensive and not reproducible. We needed a rig that runs the same set of attacks after every change of prompt or model.
What we did
and why we did it that way.
The attack catalogue
Direct and indirect injections, system prompt extraction, guardrail bypasses, tool abuse.
Indirect injection through documents
A separate set of documents with instructions hidden inside—the most productive class in our testing.
Regression in CI
The test set is handed to the client so they can run it on every model update.
An open method
We write the findings up on the blog with no reference to clients. What we find should get fixed across the industry; keeping it to ourselves buys us nothing.
What changed
and what we measured it with.
How it went
day by day and week by week.
- Stage 1
Catalogue
We collected the known attack classes and mapped them to MITRE ATLAS.
- Stage 2
Automation
Running them, collecting responses, picking out leads for manual work.
- Stage 3
Regression
Exporting the test set into a format that drops into someone else’s CI.
How it looks
in schematics.
These are screen schematics. We do not publish client interfaces without permission
A similar problem
on your side?
Describe it in the brief. In working hours we come back with an estimate of time and cost within two hours.