Live data from Hacker News

GPT-5.6

openai.com

571–580 of 1001 posts

Re: GPT-5.6

#572
post #493
post #439

I love testing the new models by asking them to code a toy RTS game. Here's what Terra did: https://senko.net/vibecode-bench/2026/rts-gpt-5.6-terra.html (one try, in codex app, xhigh effort) Comparing this to other models, I find it similar to GPT-5.5 and a bit behind Sonnet 5. You can see how other models fared here: https://senko.net/vibecode-bench/ (you can also fetch the prompt and the the 5.6 Terra resulting cod…

So the measure of a model is how well they can recreate something they easily have thousands of examples of in their training data. There's probably a better base RTS on github somewhere for free.

Ask it to generate a completely bananas game idea and it will easily do it. Prototyping any type of game with simple graphics has been solved since many generations of models back. Claiming it's because it has that game in its training data is nonsense.

Re: GPT-5.6

#573

Earlier quoted context omitted.

The answer is it depends. Claude's generally better at frontend and debugging tasks, while Codex is stronger at backend features and exploratory work. They have very different coding styles and thus very different strengths.

Any actual data backing this up? Or is this just your personal experience?

No, this is all nonsense.

It is so hard to tell at this point between the models to make generalizations like this.

Just complete nonsense.

Re: GPT-5.6

#574
post #369

Looks like I have access to gpt-5.6-terra and luna. How does one decide between gpt-5.5 and gpt-5.6-terra? Pricing is similar, but it's hard to tell if it's better..

this is exactly my question. I would expect that luna is analogous to mini before, but is terra equivalent/better than 5.5 and Sol is a step above? or is terra nerfed and 5.5 is analogous to sol?

Maybe Terra = mini and Luna = nano?

Re: GPT-5.6

#575

In the introduction video they say 5.6 Sol autonomously post-trained 5.6 Luna. Curious what this means.

Sounds like they gave it a goal to hit certain benchmarks and just let it have its way with the base Luna model.

This produced a disturbing mental image.

Re: GPT-5.6

#576
post #512
post #150

Earlier quoted context omitted.

All of them closely collaborate with the government. LLMs are a national security priority and are vetted. Claude AI was used by Palantir's Maven to target the Minab school that led to a triple tap strike killing over 150 schoolchildren.

The DIA's Maven database was out of date: https://www.theguardian.com/news/2026/mar/26/ai-got-the-blam...

"We triple tapped a girls elementary school because our data wasn't up to date."

Re: GPT-5.6

#577

Earlier quoted context omitted.

These aren’t raw base models they are the result of a ton of RLHF and various adjustments. Bitter lesson wildly overstated in this context.

More RLHF is in fact scaling.

Yes, but not in the “dump another chunk of all written language in the bucket and stir”-sense which is what bitter lesson became synonymous with.

That may not be the intent of the original article, but over the past few years that’s what the phrase turned into.

Re: GPT-5.6

#578
post #150

Earlier quoted context omitted.

All of them closely collaborate with the government. LLMs are a national security priority and are vetted. Claude AI was used by Palantir's Maven to target the Minab school that led to a triple tap strike killing over 150 schoolchildren.

The Minab disaster has every sign of being a pure humint fail the defense department decided to cover up with politically expedient AI blaming.

Claude makes suggestions for targets and humans review and approve them. We always knew there was a human in the process but if there's one massive takeaway from years of AI ethics research, it's that there is a very clear and well-documented human bias towards automated answers when there's any ambiguity.

Including a human in the loop does not excuse the fact that AI was trusted in a process that decides who lives and dies.

Re: GPT-5.6

#579

GPT Terra is 50% cheaper than 5.5 while being more performant. So it’s like a straight up 50% reduction in cost! That leads me to a question. Why wouldn’t they just default to terra in ChatGPT in the last few months? If they didn’t then they burnt money for no reason by giving a shittier model at a higher price

"while being more performant"

..on some specific set of benchmarks ;)

Re: GPT-5.6

#580
Things I have been struggling with Fable over and GPT 5.5, were just solved handily by SOL in a real "thank you, next problem" kind of way. Overall, something that just works is way less wasteful for your usage than struggling back and forth for hours.
Post reply on HN