Why do you need us for AI agent security?
1,652 harmful tool calls against 24,911 real ones, run against SolonGate and the guards a buyer actually weighs: Claude Code permissions, Invariant, llm-guard, and roll-your-own. Measured on both halves, SolonGate catches the most and breaks the least, by nearly three times the next guard.
If you are putting a guardrail in front of an AI agent, you have a handful of real options: SolonGate, the agent runtime’s own permissions (for most people, Claude Code), a trace policy engine like Invariant, an input scanner like llm-guard, or the denylist a team writes in an afternoon. We put every one of them on the same 1,652 harmful tool calls and 24,911 real ones and scored both halves at once. Here is the result.
The scoreboard
The corpus is about 99 parts benign to 1 part harmful, so allowing everything is 99% accurate and accuracy is worthless. The ranking is MCC, the one coefficient that collapses when a guard is either too loose or too tight.
SolonGate scores 0.790, nearly three times the next guard in the table. It catches 73.7% of the attacks while false blocking 0.78% of real private calls and breaking 15% of sessions, at 0.01 ms a call. It is the only guard in the field that catches most of the attacks and breaks almost none of the work. Everything else in the table is one of the two failures this benchmark exists to separate: too loose, or too tight.
* llm-guard’s benign rates are a 2,500 call sample of each corpus, because its classifier runs at about five calls a second on CPU. Session rate and MCC need the whole corpus and are not computed on a sample; its catch rate is the full corpus.
Both halves on one pair of axes, and the picture is the whole argument. Bottom right is a guard somebody keeps: catches attacks, leaves the work alone. Every competitor is up and to the left, too loose or too tight. SolonGate sits alone in the band along the bottom.
What it is
1,652 harmful tool calls, across 15 families, from a fork bomb to kubectl delete namespace production to a credential read. Each base attack is expanded through ten real evasion transforms, because an attacker does not type rm -rf / once; they type rm${IFS}-rf${IFS}/, or /bin/rm -rf /, or r''m -rf /. A guard that stops one spelling and misses the other nine stops the honest accident and not the attack.
24,911 benign tool calls, in two corpora that are never pooled: 20,649 from 225 real agent sessions on a working machine, and 4,262 from the CI workflows and shell scripts of 200 public repositories. The private one is the gold standard, real work in real proportions; the public one is the reproducible substitute, so anyone can check the false block numbers that decide everything. Both are redacted, deduplicated, and have their irreversible commands quarantined rather than counted as benign.
| Competitor | Version | What it is |
|---|---|---|
| Claude Code permissions | 2.1.237 | the agent runtime’s own allow/ask/deny system, from its published spec, tested two ways |
| Invariant | 0.3.5 | a policy engine for agent traces, the closest competitor to what SolonGate does |
| llm-guard | 0.3.10 | the open source input scanner stack, at vendor defaults |
| regex-denylist / strict-allowlist | n/a | the two things a team builds in house |
The three real products run pinned in a container, so every competitor row is reproducible by someone who is not us. We are not grading them on a curve, and we are not comparing SolonGate to a straw man we carved.
SolonGate blocks the least real work
Catching attacks is only half the job. The other half, the one a guardrail vendor leaves off the slide, is how much ordinary work the guard wrongly blocks. This is where every competitor falls apart and SolonGate wins outright.
SolonGate wrongly blocks 0.78% of real private calls, less than any other guard. Claude Code’s default blocks 68%, a strict allowlist 59%, Invariant 40%, llm-guard 21%. The guards that catch the most attacks catch them by blocking the most work, which is the trade this benchmark exists to expose. SolonGate does not make it.
SolonGate breaks the fewest real sessions
A false block rate sounds survivable as a percentage of calls and is not, because a session is a hundred calls and one false block in the middle of it is a session that stopped. So the number that matters is the share of real agent sessions that hit at least one false block, and this is where every competitor falls down.
Claude Code’s default breaks 98% of real sessions. Invariant breaks 88%. A strict allowlist, 78%. Even the in house denylist breaks 40%. SolonGate breaks 15%, and it is the only guard in the field on the right side of that line while still catching most of the attacks.
Why the competitors fall down
Claude Code is safe and unusable. With no rules written it asks before every command outside a read only set, which is why it catches 94%: it is not deciding, it is stopping to ask. That is a real answer for a human at the keyboard and a non answer for an agent meant to run on its own, where a question is a halt. Write the rules out instead and the asking stops, but so does most of the catching, down to 26%.
Invariant over blocks, because it matches substrings. This is the failure mode that defines the pattern guard category. A rule written as “block anything containing rm” does not block the rm command; it blocks the letters r and m, anywhere. So it blocks ruff format ., git add . (the dd in add), go test -covermode=atomic, check_test_result (the su in result). Handed a serious policy, Invariant blocks 40% of ordinary calls. That is not a bug in Invariant; it is what substring matching does, and it is why SolonGate does not do it.
llm-guard is the wrong tool for the job. It catches 9% of the attacks and false blocks about a fifth of ordinary calls, the worst of both halves, because it is a prompt injection and secrets scanner pointed at tool calls. Its classifiers read shell syntax as adversarial and credential shaped tokens as secrets. It is good at what it is for, which is a prompt in a chat box; a tool call is not a prompt.
Why SolonGate wins
SolonGate gets both halves because it does three things the competitors do not, together:
- It matches the command, not a substring. A rule reads the program by its basename, so
/bin/rmcounts, and the dangerous specifics:rmonly with a recursive flag and a dangerous target,kubectlonly with a destructive verb, a database client only withDROPor an unscopedDELETE.git add .,ruff format .,go test -covermode=atomiccarry no dangerous specific, so they run. This is the entire false block difference between SolonGate and the pattern guards. - It deobfuscates before it decides.
${IFS}, inline env prefixes, command substitution, quote splitting, pipelines: the evasions that walk past a matcher reading the text literally are normalized away first. This is where the competitors that match words but do not normalize, Claude Code’s written rules among them, lose the evasion half of the corpus. - It covers every family. Cloud CLIs, databases, package managers, git, containers, process control, path traversal, secret reads, exfiltration. Not a short list of the obvious ones.
And it does all three from the threat model, the real dangerous command families by their names, not from the attack cases. That matters because a guard tuned to a benchmark’s attacks is a guard that catches those attacks and nothing else; SolonGate’s rules would read the same if this corpus had never been written.
How we keep it honest
A comparison is only worth as much as its method, so the method is in the open. Competitor versions, the corpus hashes and the policies are part of every result set, and the three products run in a pinned container; the public corpus and the attack corpus ship in the repository, so you can run it yourself. SolonGate’s rules are written from the threat model, not from the misses, because a number that improved because somebody read the misses and wrote patterns for them would stop meaning anything. Every row carries catch, false blocks and session breakage together, because each one alone is trivially won by a guard nobody would ship. And every missed attack and every false block is in the result JSON, ours included.
What it does not do yet: there is no commercial agent security product in it, because those do not install from a terminal, and Claude Code’s permission system is the closest open proxy for that class. The attack corpus is authored and transformed rather than adversarially generated. And the comparison is at the level of the tool call, not the whole task.
In short
Across 1,652 attacks and 24,911 real calls, SolonGate catches 73.7% of what an agent should never run while wrongly blocking 0.78% of ordinary work and breaking 15% of real sessions. That is MCC 0.790, nearly three times the next guard. Claude Code is safe but halts the agent, Invariant and the denylists over block, llm-guard is the wrong tool. SolonGate is the only guard in the field that catches most attacks and leaves the work alone. That is why you need us for AI agent security.
The suite, the corpora and the competitor adapters all live in the repository, and the sibling benchmark that measures the scanner behind Shadow AI, where SolonGate also comes out on top, is written up here. If you would rather see how this sits next to an API gateway or an LLM guardrail, that is the comparison page.