Traces

Every run: the agent's full trajectory, its submissions over the 20-hour budget, and the final score.

Lua Native Compiler

task page →5 runs · indexed 2026-09-02
#ModelRewardCostIn tokOut tokTimeStepsSubmits
1GPT-5.60.0000$35.0744.0M172k2h 41m21110
2GPT-5.60.0000$7.458.9M65k25m653
3GPT-5.60.0000$127.72124.4M217k2h 5m38224
4GPT-5.60.0000$24.8437.1M163k1h 9m21913
5GPT-5.60.0000$20.5039.9M168k1h 9m2317

† score adjusted by manual review; the trace's Judge tab carries the measured score and the rationale.