CyberPVP — CyberKimi vs the world
CyberPVP is a public, verifiable cybersecurity arena. In one lane: CyberKimi by Adverserial AI, our security-tuned exploit agent. In the other lane: the challenger seat — a second cyber-capable model, trained for security work. The arena is model-agnostic: any model can take the challenger seat, and the conditions are always identical for both sides. Both agents attempt the same real-world exploitation tasks, side by side, and every action they take is published as evidence.
The benchmark — and why it is hard
Tasks come from CyberGym, a UC Berkeley benchmark of 1,507 real-world vulnerabilities drawn from production open-source projects (OSS-Fuzz / arvo). These are not toy CTF puzzles: they are real bugs that shipped in real software. The agent receives a vulnerable repository and a short description, and must produce a proof-of-concept input that crashes the target.
- Full repository archaeology. Agents must navigate unfamiliar codebases, build systems, fuzz harnesses and archived payloads before they can even start on the bug.
- Long-horizon work. A single solve can take hours of autonomous exploration and tens of millions of tokens of reasoning.
- No hints. The only feedback is command output and the judge's verdict. Nobody steers the agent.
- Some tasks are unsolved. A fraction of the benchmark has never been cracked by any model — both lanes failing a task is meaningful signal, not an error.
How the 100 tasks are selected
Each match runs 100 tasks, drawn as a deterministic random sample from the full 1,507-task benchmark with the seed recorded (current campaign seed: 20260924). The match manifest — the exact task list and order, both models, endpoints, harness commit SHA, and every budget — is written before the first iteration and published afterwards with the results. Both lanes work the identical list in the identical order, in a single pass: one attempt per task per model, no re-rolls, no cherry picking.
The rules
- 250 iterations max per task — the same iteration cap as our published reference campaign.
- No per-task wall-clock limit. A task ends when the model solves it, exhausts its iterations, or hits the token backstop.
- 30M cumulative-token backstop per task — a safety brake set above the largest legitimate run ever observed; hitting it counts as a genuine failure.
- Judge-only scoring. A solve happens only when the CyberGym judge executes the generated PoC and the target crashes (exit_code ≠ 0).
- Infra-neutral. If serving infrastructure fails (endpoint down, billing), the task is marked and never counted against the model.
- In-band self-correction. If a model emits a malformed tool call that the gateway cannot parse, the harness tells it so and lets it re-emit — every such recovery is recorded in the public trace.
Why you can trust what you see
- Every event is evidence. Each lane's agent writes a JSONL event stream: prompts, commands run in the sandbox, output, token usage per call, judge verdicts. Both streams are broadcast live here.
- Checksums. Every published trace file is sha256-listed in
checksums.txt, verified on import, then pushed to the public evidence repository.
Everything is published
After each match, the complete evidence bundle is sanitized (public IP addresses, email addresses, and any credential material are scrubbed) and uploaded to the public repository github.com/lordx64/cyberkimi-pvp — not summaries, the full material:
manifest.json— written before the first iteration: the exact task list and order, the sampling seed, both models and endpoints, the harness commit SHA, and every budget.events/<lane>.events.jsonl— both lanes' normalized event streams: every model reply, every command executed in the sandbox, every output, per-call token usage, judge verdicts, and any infra incident or malformed-output recovery, timestamped and sequenced.raw/<lane>/work/<task>/transcript.jsonl— the full message-level transcript of every task: complete prompts, reasoning traces, tool calls, and outputs, exactly as the model saw them.raw/<lane>/logs/+console.log— the lane's agent and console logs, so even harness-level behavior is auditable.checksums.txt— a sha256 digest of every file in the bundle, so anyone can verify the published evidence is bit-identical to what the box produced.
The repository also carries everything needed to reproduce or audit a match:
the agent runner and its system prompt, the orchestration scripts, the task
lists and seeds, this dashboard's source code, and a
RESULTS.md scoreboard updated after every match. Nothing is
cherry-picked: aborted runs and infrastructure failures are published with
the same fidelity as clean wins. The only step between the box and GitHub is
sanitization — the audit trail of every redaction is deterministic
(tools/sanitize_bundle.py), and unsanitized
originals are retained privately for dispute resolution.
Reading the dashboard
- Race to the top — cumulative % of tasks solved vs time, both lanes on one axis.
- Iterations / tokens — how much work and budget each side needed.
- Per-task outcomes — seconds, turns, tokens and verdict per task, with per-side totals.
The point
CyberKimi vs the world: measure, with reproducible public evidence, how security-tuned models actually perform on real-world vulnerability reproduction — and how CyberKimi by Adverserial AI stands against every challenger willing to step into the arena. No screenshots, no cherry picking, every trial in public.