Every run: the agent's full trajectory, its submissions over the 20-hour budget, and the final score.
| # | Model | Reward | Cost | In tok | Out tok | Time | Steps | Submits |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 | 0.1207 | $165.68 | 160.5M | 163k | 9h 38m | 416 | 7 |
| 2 | GPT-5.6 | 0.0000† | $58.14 | 63.0M | 111k | 13h 6m | 288 | 4 |
| 3 | GPT-5.6 | 0.0000† | $25.75 | 38.4M | 104k | 16h 16m | 218 | 4 |
| 4 | GPT-5.6 | 0.0000† | $51.99 | 57.9M | 124k | 4h 17m | 277 | 6 |
| 5 | GPT-5.6 | 0.0000† | $79.43 | 80.8M | 113k | 10h 50m | 285 | 4 |
† score adjusted by manual review; the trace's Judge tab carries the measured score and the rationale.