score table, my own set
72 hours in
Same harness, real repos, 72 hours on the clock. Scores land here once the runs finish — cost, limits, and whether it finished the job.
Ranked apart, never averaged. ~ means provisional.
fixed suite v1
My own set, frozen. Every model gets the same tasks at each effort level; the ranking below only counts v1 runs.
No v1 scores yet — nothing baked. Check back once the runs finish.
fixed suite v2
A new fixed set gets a new table. v2 scores never mix into the v1 ranking, and v1 never props up v2.
No v2 scores yet — nothing baked. The table fills in when real v2 runs land.
harnesses, head to head
Same model, same task, different harness. Hidden tests grade every run; cost and time come from the harness's own usage.
Loading harness results…