Skip to content

How it works

It guesses like a person, then checks like a machine

Deciding what to attack takes judgment, which is what a language model is good at. Whether it worked isn’t a judgment call, so fixed tooling settles that.

The pipeline

Four stages, and one of them exists to say no

Each finding passes through a chain of specialist agents, and they can overrule each other.

  1. 01

    Discovery

    A specialist agent per class of bug. It maps the surface and forms a hypothesis.

  2. 02

    Validation

    A different agent tries to exploit it for real. Failing here is normal and expected.

  3. 03

    Reporting

    Only a reporting agent can file, and on a web target the filing is rejected without working exploit code.

  4. 04

    Fixing

    White-box only. Writes a patch, tests it, and puts the diff in the report.

The sandbox

A pentester’s toolkit in a disposable box

Every scan gets its own isolated container with the kit a human tester reaches for, destroyed at the end.

Intercepting proxy

Every request and response, replayable.

Scripted browser

A real DOM, for flows a script can’t reach.

Persistent shells

State survives between steps.

Python runtime

Where exploit code gets written and fired.

Recon & scanning

Port, service and content discovery.

Static & taint analysis

Source-level reasoning, white-box only.

In white-box a copy of your source goes into the container, and your filesystem is never mounted. The agent is model-agnostic, so the code goes to whichever frontier or open model you are comfortable sending it to.

Trust

Why a finding from us is worth reading

Two of these the tooling enforces whatever the model says. The third is a judgment call.

No exploit, no finding

A report filed without working proof-of-concept code is thrown out before a human sees it. That makes false positives rare on web targets. It does not eliminate them.

Severity is computed

CVSS 3.1 metrics are submitted one at a time, validated, and scored by a library, because a model asked to name a severity will cheerfully say “high” and mean nothing by it.

Duplicates get merged

Candidates are compared on root cause: would the same patch fix both? If so, one finding instead of nine.

This is the judgment call, and a model makes it.

Native binary co-analysis

Where most AI pentesters stop

The source you review and the machine code you ship aren’t the same artifact. Compilers inline functions, drop dead code and optimise across files.

The binary and its source, read together

We correlate the shipped binary, its debug symbols and the source tree into one picture, with symbolic execution and taint analysis over the result. Findings get checked in both directions, source against disassembly.

Where it stops

  • It is static. Findings are pinned to the exact function, but nothing is executed.
  • It leans on symbols. Strip or obfuscate the binary and the analysis still runs, but name-based detection degrades, like it does for a human reverse engineer.
  • It generates candidates. Leads with evidence attached. Reachability is tagged, never claimed as proved.

In your pipeline

It can gate a merge

On the continuous plan, scans run headless and exit non-zero on a finding, so CI can fail the build before a bad merge lands. Scope narrows to a pull request diff.

Bring us something hard

An application, a repository, or a binary you ship. The first scan is free.