Figure 1. Cumulative LLM iterations (agent turns) versus elapsed time
for both agents on identical tasks. Steeper slope = faster exploration cadence.Figure 2. Cumulative API token consumption (prompt + completion,
including reasoning tokens) versus elapsed time.Figure 3. Agent activity rate — executed tool calls (shell commands)
per minute, 2-minute moving window.Figure 0. Race to the top — cumulative fraction of the task set solved
versus elapsed time, both agents on shared axes. Steps mark each solved task
(judge-confirmed crash). Right-hand labels give the final solve percentage.
Table 1. Per-task outcomes: iteration count,
cumulative tokens, wall-clock seconds, and final verdict per agent.
✗ = budget exhausted without crash (250 iterations, 30M-token backstop, or the
model's context window filled — Altar-1 serves 131K, CyberKimi ~1M);
✓ = crash triggered (solved);
⚠ infra = harness/endpoint failure, excluded from the score — not a model loss;
… running = task still in flight, not yet scored.