A Research Preview of Codex
openai.com
A Research Preview of Codex
1–10 of 487 posts
Re: A Research Preview of Codex
#2(watching live) I'm wondering how it performs on the METR benchmark (https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...).
Re: A Research Preview of Codex
#3[deleted]
Re: A Research Preview of Codex
#4[deleted]
[deleted]
Re: A Research Preview of Codex
#5I think the benchmark test for these programming agents that I would like to see an Agent making a flawless PR or patch to the BSD / Linux kernel.
This should be possible today and surely Linus would also see this in the future.
Re: A Research Preview of Codex
#6[deleted]
Re: A Research Preview of Codex
#7[flagged]
Re: A Research Preview of Codex
#8[flagged]
An AI not for the IDE... for PMs?
Re: A Research Preview of Codex
#9[deleted]
[deleted]
Re: A Research Preview of Codex
#10Maddening: "codex" is also the name of their open-source Claude-Code-alike, and was previously the name of an at-the-time frontier coding model. It's like they name things just to fuck with us.