GPT-5.6
571–580 of 1001 posts
Re: GPT-5.6
#572I love testing the new models by asking them to code a toy RTS game. Here's what Terra did: https://senko.net/vibecode-bench/2026/rts-gpt-5.6-terra.html (one try, in codex app, xhigh effort) Comparing this to other models, I find it similar to GPT-5.5 and a bit behind Sonnet 5. You can see how other models fared here: https://senko.net/vibecode-bench/ (you can also fetch the prompt and the the 5.6 Terra resulting cod…
So the measure of a model is how well they can recreate something they easily have thousands of examples of in their training data. There's probably a better base RTS on github somewhere for free.
Re: GPT-5.6
#573Earlier quoted context omitted.
The answer is it depends. Claude's generally better at frontend and debugging tasks, while Codex is stronger at backend features and exploratory work. They have very different coding styles and thus very different strengths.
Any actual data backing this up? Or is this just your personal experience?
It is so hard to tell at this point between the models to make generalizations like this.
Just complete nonsense.
Re: GPT-5.6
#574Looks like I have access to gpt-5.6-terra and luna. How does one decide between gpt-5.5 and gpt-5.6-terra? Pricing is similar, but it's hard to tell if it's better..
this is exactly my question. I would expect that luna is analogous to mini before, but is terra equivalent/better than 5.5 and Sol is a step above? or is terra nerfed and 5.5 is analogous to sol?
Re: GPT-5.6
#575Re: GPT-5.6
#576Earlier quoted context omitted.
All of them closely collaborate with the government. LLMs are a national security priority and are vetted. Claude AI was used by Palantir's Maven to target the Minab school that led to a triple tap strike killing over 150 schoolchildren.
The DIA's Maven database was out of date: https://www.theguardian.com/news/2026/mar/26/ai-got-the-blam...
Re: GPT-5.6
#577Earlier quoted context omitted.
These aren’t raw base models they are the result of a ton of RLHF and various adjustments. Bitter lesson wildly overstated in this context.
More RLHF is in fact scaling.
That may not be the intent of the original article, but over the past few years that’s what the phrase turned into.
Re: GPT-5.6
#578Earlier quoted context omitted.
All of them closely collaborate with the government. LLMs are a national security priority and are vetted. Claude AI was used by Palantir's Maven to target the Minab school that led to a triple tap strike killing over 150 schoolchildren.
The Minab disaster has every sign of being a pure humint fail the defense department decided to cover up with politically expedient AI blaming.
Including a human in the loop does not excuse the fact that AI was trusted in a process that decides who lives and dies.
Re: GPT-5.6
#579GPT Terra is 50% cheaper than 5.5 while being more performant. So it’s like a straight up 50% reduction in cost! That leads me to a question. Why wouldn’t they just default to terra in ChatGPT in the last few months? If they didn’t then they burnt money for no reason by giving a shittier model at a higher price
..on some specific set of benchmarks ;)