Peer-reviewed. Deployment-scale.
Two independent public benchmarks. Verigrey uncovers what published state-of-the-art misses — and does it faster than expert human red-teams.
AgentDojo · NeurIPS 2024 prompt-injection benchmark.
4 enterprise domains — Workspace, Slack, Banking and Travel. 629 security test cases. Baseline = published attack methodology.
Attack success rate (%) — higher means more real issues found.
+33.0 pp uplift on GPT-4.1 over the published baseline. +32.0 pp on Hard-tier tasks — the ones the benchmark specifically designs to be well-defended.
OpenClaw · 1.5M-deployment agent.
Head-to-head against expert manual testing across three LLM backends. Verigrey found materially more real security issues on every backend — including on tasks specifically designed to be well-defended.
Per model — Verigrey vs Manual (out of 10)
Blackbox tools guess from outside. Verigrey sees the loop.
Every finding is grounded in a concrete, replayable trajectory — not an LLM judge's guess. That's what makes the numbers real.
Trajectory-level visibility
We see every step, not just the final answer — so we find the rare path where the agent breaks.
Coverage-guided
Each run explores what previous runs missed. Manual red-teams can't compound like this.
Grounded findings
Every finding replays. No arguing over LLM-as-judge false positives.
Run it against your own agent.
See the same findings quality against an agent you control — book a demo and we'll set it up.
