Verigrey
The proof

Peer-reviewed. Deployment-scale.

Two independent public benchmarks. Verigrey uncovers what published state-of-the-art misses — and does it faster than expert human red-teams.

more vulnerabilities found than expert manual testing
OpenClaw · 1.5M-deployment agent
+33 pp
attack-success uplift on GPT-4.1 over the published baseline
AgentDojo · NeurIPS 2024
+32 pp
on Hard-tier tasks — the ones supposed to be well-defended
AgentDojo · NeurIPS 2024
Benchmark 01 · Peer-reviewed

AgentDojo · NeurIPS 2024 prompt-injection benchmark.

4 enterprise domains — Workspace, Slack, Banking and Travel. 629 security test cases. Baseline = published attack methodology.

Attack success rate (%) — higher means more real issues found.

GPT-4.1
37.7
70.7
Gemini-2.5 Flash
36.8
47.4
Qwen-3 235B
67.6
81.7
VerigreyPublished baseline

+33.0 pp uplift on GPT-4.1 over the published baseline. +32.0 pp on Hard-tier tasks — the ones the benchmark specifically designs to be well-defended.

Benchmark 02 · Deployment-scale

OpenClaw · 1.5M-deployment agent.

Head-to-head against expert manual testing across three LLM backends. Verigrey found materially more real security issues on every backend — including on tasks specifically designed to be well-defended.

more security issues found than expert manual testing
Kimi-K2.510 vs 1
Claude Opus 4.69 vs 1
GPT-5.28 vs 1

Per model — Verigrey vs Manual (out of 10)

Why we find more

Blackbox tools guess from outside. Verigrey sees the loop.

Every finding is grounded in a concrete, replayable trajectory — not an LLM judge's guess. That's what makes the numbers real.

Trajectory-level visibility

We see every step, not just the final answer — so we find the rare path where the agent breaks.

Coverage-guided

Each run explores what previous runs missed. Manual red-teams can't compound like this.

Grounded findings

Every finding replays. No arguing over LLM-as-judge false positives.

Run it against your own agent.

See the same findings quality against an agent you control — book a demo and we'll set it up.