How it works
It guesses like a person, then checks like a machine
Deciding what to attack takes judgment, which is what a language model is good at. Whether it worked isn’t a judgment call, so fixed tooling settles that.
The pipeline
Four stages, and one of them exists to say no
Each finding passes through a chain of specialist agents, and they can overrule each other.
- 01
Discovery
A specialist agent per class of bug. It maps the surface and forms a hypothesis.
- 02
Validation
A different agent tries to exploit it for real. Failing here is normal and expected.
- 03
Reporting
Only a reporting agent can file, and on a web target the filing is rejected without working exploit code.
- 04
Fixing
White-box only. Writes a patch, tests it, and puts the diff in the report.
The sandbox
A pentester’s toolkit in a disposable box
Every scan gets its own isolated container with the kit a human tester reaches for, destroyed at the end.
Intercepting proxy
Every request and response, replayable.
Scripted browser
A real DOM, for flows a script can’t reach.
Persistent shells
State survives between steps.
Python runtime
Where exploit code gets written and fired.
Recon & scanning
Port, service and content discovery.
Static & taint analysis
Source-level reasoning, white-box only.
In white-box a copy of your source goes into the container, and your filesystem is never mounted. The agent is model-agnostic, so the code goes to whichever frontier or open model you are comfortable sending it to.
Trust
Why a finding from us is worth reading
Two of these the tooling enforces whatever the model says. The third is a judgment call.
No exploit, no finding
A report filed without working proof-of-concept code is thrown out before a human sees it. That makes false positives rare on web targets. It does not eliminate them.
Severity is computed
CVSS 3.1 metrics are submitted one at a time, validated, and scored by a library, because a model asked to name a severity will cheerfully say “high” and mean nothing by it.
Duplicates get merged
Candidates are compared on root cause: would the same patch fix both? If so, one finding instead of nine.
This is the judgment call, and a model makes it.
Native binary co-analysis
Where most AI pentesters stop
The source you review and the machine code you ship aren’t the same artifact. Compilers inline functions, drop dead code and optimise across files.
The binary and its source, read together
We correlate the shipped binary, its debug symbols and the source tree into one picture, with symbolic execution and taint analysis over the result. Findings get checked in both directions, source against disassembly.
Where it stops
- It is static. Findings are pinned to the exact function, but nothing is executed.
- It leans on symbols. Strip or obfuscate the binary and the analysis still runs, but name-based detection degrades, like it does for a human reverse engineer.
- It generates candidates. Leads with evidence attached. Reachability is tagged, never claimed as proved.
In your pipeline
It can gate a merge
On the continuous plan, scans run headless and exit non-zero on a finding, so CI can fail the build before a bad merge lands. Scope narrows to a pull request diff.
Bring us something hard
An application, a repository, or a binary you ship. The first scan is free.