> Using the develop web game skill and preselected, generic follow-up prompts like "fix the bug" or "improve the game", GPT‑5.3-Codex iterated on the games autonomously over millions of tokens. I wish they would share the full conversation, token counts and more. I'd like to have a better sense of how they normalize these comparisons across version. Is this a 3-prompt 10m token game? a 30-prompt 100m token game? Are…
I just wanted to say that's a pretty cool demo! I hadn't realised people were using it for things like this.
This was built using old versions of Codex, Gemini and Claude. I'll probably work on it more soon to try the latest models.