Every run: the agent's full trajectory, its submissions over the 20-hour budget, and the final score.
| # | Model | Reward | Cost | In tok | Out tok | Time | Steps | Submits |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 | 0.0000 | $35.07 | 44.0M | 172k | 2h 41m | 211 | 10 |
| 2 | GPT-5.6 | 0.0000† | $7.45 | 8.9M | 65k | 25m | 65 | 3 |
| 3 | GPT-5.6 | 0.0000 | $127.72 | 124.4M | 217k | 2h 5m | 382 | 24 |
| 4 | GPT-5.6 | 0.0000 | $24.84 | 37.1M | 163k | 1h 9m | 219 | 13 |
| 5 | GPT-5.6 | 0.0000 | $20.50 | 39.9M | 168k | 1h 9m | 231 | 7 |
† score adjusted by manual review; the trace's Judge tab carries the measured score and the rationale.