Tournament Forge Example Run
Claude skill · tournament-forge · standard tier

Tournament Forge

Illustrative example. This is not a recorded run. It shows what each stage produces for a realistic sample problem, so you can see the flow without spending tokens. A real recorded run is in the repo under examples/rate-limiter-run.

The prompt Add rate limiting to our public REST API. About 50k users, a two person team, Node and Postgres. It must never lock out paying customers.
8approaches entered
22subagent calls
~1.3Mtokens (at measured 57k per call)
5/5tests pass on the final answer
Stage 0Sharpenorchestrator · 0 calls
Stage 0bWrite tests1 call · code only
Stage 1Enumerateorchestrator · 0 calls
Stage 2Blind builds8 calls · mid
Stage 3Run + gate1 call · cheap
Stage 4Bracket11 calls · cheap to strong
Stage 5Forgeorchestrator + 1 red-team
0

Sharpen the input

The rubric and tests are locked here, before any answer exists.

Brief

  • Must: paid plans are never hard-blocked by normal use.
  • Must: no extra write per request on Postgres (already at 70% CPU).
  • Must: a two person team can run it.
  • Out of scope: billing changes, DDoS protection.

Rubric (weights sum to 100)

Protects paying customers30 Correct under abuse and outages25 Simple to run for 2 people20 Infra and token cost15 Clear to API users10
1–2

Eight real approaches

Each differs in mechanism. Built blind to each other.

Probability is how likely a typical expert would suggest it. At least a third must be long shots (under 20%), so the field isn't eight versions of the obvious answer.

3

Tests decide first

Written from the brief by an agent that never sees any code.

Spec tests

Results after blind builds

T3 failed by 5 of 8. A test most builds fail gets reviewed against the brief. Verdict: valid, kept. B and C fail a HARD test and are out.

4

The bracket

Contenders only ever see “A” and “B”. Select a match to read it.
5

The forged answer

Champion plus the grafts that survived critique.

Plan-tiered quotas, enforced by a Redis token bucket per API key, with soft limits for paid plans. Paid keys get warning headers at 80% and a 24 hour grace period before any 429. Free keys are hard-limited.

1

Cache the three hot read endpoints first. This removes about 60% of load before any limiting runs.

graft from G
2

Redis token bucket per API key, sized by plan.

graft from A
3

Costly endpoints spend 5 tokens instead of 1.

graft from H
4

Soft limits for paid plans: X-RateLimit-Remaining header, an email at 80%, grace before 429.

champion E
5

Emergency load shedding for the free tier only, when p95 latency passes 800 ms.

graft from F

Test T3 + final red-team

If Redis goes down, every request fails the limit check and paying customers get 429s.

Fixed: fail open for paid keys, fail closed for free keys, alert on Redis errors.

Ships with its tests

5 spec tests pass, plus 3 regression tests added from upheld attacks and the deciding distinguishing input.

8 / 8 passing

vs

How it compares

Estimated tokens per run

~26× fewer

tokens than a full arena-skill run, and about 4× fewer than its quick mode, for a standard Tournament Forge run.

Subagent calls per run

FeatureTournament Forgearena-skillllm-councilagent-review-panel

✓ yes · ✗ no · ~ partly · ? not stated in the project's README (it may still do it)