Are benchmarks the right way to measure LLMs? Not because benchmarks can be gamed, but because the most useful outputs of models aren't things that can be bucketed into "right" and "wrong." Tough problem!
GPT-5.2
41–50 of 1001 posts
Re: GPT-5.2
#42No wall yet and I think we might have crossed the threshold of models being as good or better than most engineers already.
GDPval will be an interesting benchmark and I'll happily use the new model to test spreadsheet (and other office work) capabilities. If they can going like this just a little bit further, much of the office workers will stop being useful.... I don't know yet how to feel about this.
Great for humanity probably but but for the individuals?
Re: GPT-5.2
#430: https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3f...
Re: GPT-5.2
#44I emailed support a while back to see if there was an early access program (99.99% sure the answer is yes). This is when I discovered that their support is 100% done by AI and there is no way to escalate a case to a human.
Re: GPT-5.2
#45Re: GPT-5.2
#46Re: GPT-5.2
#47From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!
Re: GPT-5.2
#48Everything is still based on 4 4o still right? is a new model training just too expensive? They can consult deepseek team maybe for cost constrained new models.
I thought whenever the knowledge cutoff increased that meant they’d trained a new model, I guess that’s completely wrong?
I don’t think it’s publicly known for sure how different the models really are. You can improve a lot just by improving the post-training set.
Re: GPT-5.2
#49Re: GPT-5.2
#50For me the last remaining killer feature of ChatGPT is the quality of the voice chat. Do any of the competitors have something like that?