The same 1,198-word specification for a browser game, handed once to every model, with no follow-up questions and no second attempts.
Six open-weight models were run through four different agent harnesses, to separate what the model does from what the harness does around it.
Six models × four harnesses is the experiment proper — twenty-four cells where the only thing that changes is the agent wrapped around the model.
Three closed models were pushed down the same rerouted path as a control. If an open model fails where a frontier model succeeds, the path works and the model did not.
Fourteen closed runs on their own native harness set the ceiling — Codex with GPT, Claude Code with Claude, Grok Build with Grok, Kimi Code with K3.
Filter to shape the pool; click a capture to lift it into the comparison, where it stays no matter what the filters do next. Cards grow as the set shrinks. The captures are near-black by nature — lift opens the shadows without blowing the highlights.