> Source: https://planckproof.ai/benchmark  |  Plain-Markdown twin of the page.

Open Benchmark

# The agentic API pentest benchmark you can re-run yourself

Every agentic pentester claims to find more and prove it. Almost none lets you check. This benchmark is open: public vulnerable targets with known flaws, a published method, and a working proof for every claimed finding, so the numbers are yours to reproduce, not ours to assert.

[Run a free first scan](https://cloud.planckproof.ai)

[How proof works](https://planckproof.ai/proof)

The Problem

## Self-run benchmarks prove nothing

The headline numbers in this market, the vulnerability counts, the near-zero false-positive rates, are almost always produced by the vendor, validated by the same vendor, on a target no one else can see. That is not evidence, it is marketing with a chart. It is true of the API pentest agents we compare ourselves against directly, and it is true of adjacent autonomous offensive tools like [XBOW](https://planckproof.ai/xbow-alternative). A benchmark is only worth the paper it is on if someone other than the vendor can run it and get the same answer.

- **Public targets, known ground truth.** When the planted vulnerabilities are documented, coverage and false positives can be scored objectively instead of asserted.
- **Published method.** Same target versions, same scope, same scoring rules, written down so a run is a run, not a demo.
- **A proof for every finding.** The result is not a number in a slide. Each claimed finding ships with a working, replayable proof-of-concept anyone can execute.

The Target Set

## Four public, intentionally vulnerable API targets

Each is a well-known, maintained, deliberately vulnerable application with documented flaws, chosen to cover the classes that matter for modern APIs: broken authorization, business logic, injection, and GraphQL-specific abuse.

| Target | What it exercises | Primary classes |
| --- | --- | --- |
| OWASP crAPI | A realistic vehicle-services API with users, vehicles, mechanics, and orders | [BOLA](https://planckproof.ai/bola-testing), [BFLA](https://planckproof.ai/bfla-testing), mass assignment, JWT abuse |
| VAmPI | A compact REST API built to plant and toggle the OWASP API Top 10 | Broken auth, BOLA, excessive data exposure |
| DVGA | Damn Vulnerable GraphQL Application | Introspection abuse, batching, injection, DoS |
| Juice Shop (API) | The API surface behind the OWASP flagship training app | Business logic, access control, injection |

Targets are pinned to specific versions in the open harness so a result is reproducible over time, and the same run works for any tool, not just Operator.

What We Measure

## Three numbers that are hard to fake

Finding count alone rewards noise. This benchmark scores the things a buyer actually cares about, and weights the one nobody else reports: whether the finding came with a proof you can run.

- **Coverage.** Of the planted, in-scope vulnerabilities, how many did the tool actually find.
- **False-positive rate.** Of everything the tool reported, how much was not real. High coverage with a pile of false positives is a worse product, not a better one.
- **Proof rate.** Of the real findings, how many shipped with a working, reproducible proof-of-concept a third party can replay. This is the metric the rest of the market quietly skips.
- **Time to first proven finding.** How long from pointing the tool at a target to a single, reproduced exploit in hand.

Results

## Scored against pinned runs, with the proofs attached

Operator's coverage, false-positive rate, and proof rate on each target are scored from a real, version-pinned run, and every number links to the exact reproducible proof in the open harness. We publish the scoreboard the same way we ask everyone else to: in the repo, where anyone can re-run it and check the claim. The current scoreboard, and any challenger's results, live in the harness.

[Get the scoreboard & harness](https://planckproof.ai/contact)

A benchmark is only worth anything if a third party can reproduce it. If you can stand up the target, you can check every number we publish.

For Press & Researchers

## A benchmark you can quote, and then check

Most security benchmarks cannot be cited responsibly because no one outside the vendor can reproduce them. This one can. If you write about agentic API security, testing tools, or the OWASP API Top 10, you are welcome to reference the method, run it, and report what you find, favourable to us or not.

- **Openly documented method.** The target set (crAPI, VAmPI, DVGA, Juice Shop), the pinned versions, and the scoring rules for coverage, false positives, and proof rate are all published, not summarised.
- **Reproducible by a third party.** Stand up the same targets and run the same steps; every published number is meant to be re-derived, not taken on trust.
- **A proof for every finding.** Each reported vulnerability ships with the request, the response, and replayable steps, so a claim can be verified rather than believed.
- **Open to challengers.** Any vendor in the category can run the same harness and publish their own result. We will link to credible challenger runs.

[Request the method & harness](https://planckproof.ai/contact)

[See how proof works](https://planckproof.ai/proof)

Press & research enquiries: contact us for the methodology brief and access to the harness. We are happy to walk a reporter or independent researcher through a live run.

Reproduce It Yourself

## The harness is open

The methodology, pinned target versions, scoring scripts, and a working proof for every finding live in a public repository. Stand up the same targets, run the same steps, and judge the result on your own machine. The invitation is open to every vendor in this category.

- **Pinned targets.** Exact versions and setup, so a run today matches a run next quarter.
- **Scoring you can audit.** The rules for coverage, false positives, and proof rate are code, not a footnote.
- **A proof per finding.** Each reported vulnerability comes with the request, the response, and the steps to reproduce it.
- **Open to challengers.** Run any agentic pentester through the same harness and publish the result. A benchmark that only its author can pass is not a benchmark.

[Get the harness](https://planckproof.ai/contact)

[See how Operator compares](https://planckproof.ai/equixly-alternative)

Get Started

## Now run it on the target that matters: yours.

A public benchmark shows the method. Your own API shows the point. Your first scan is free, with a reproducible proof for anything Operator finds.

[Run a free first scan](https://cloud.planckproof.ai)

[See pricing](https://planckproof.ai/pricing)
