Every run: the agent's full trajectory, its submissions over the 20-hour budget, and the final score.
| # | Model | Reward | Cost | In tok | Out tok | Time | Steps | Submits |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 | 0.5275 | $209.42 | 198.9M | 277k | 4h 5m | 589 | 12 |
| 2 | GPT-5.6 | 0.3901 | $654.39 | 640.1M | 687k | 19h 58m | 1472 | 28 |
| 3 | GPT-5.6 | 0.0055† | $293.05 | 285.8M | 311k | 4h 20m | 671 | 6 |
| 4 | GPT-5.6 | 0.0000 | $137.59 | 141.5M | 243k | 3h 27m | 480 | 2 |
| 5 | GPT-5.6 | 0.0000 | $142.29 | 142.1M | 259k | 4h 34m | 462 | 11 |
† score adjusted by manual review; the trace's Judge tab carries the measured score and the rationale.