Live data from Hacker News

Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

openai.com

161–170 of 280 posts

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#162
post #83

Earlier quoted context omitted.

I don't understand this at all. Whenever I ask Gemini 3.1 Pro Extended, or Claude 5 Max something in chat, the most I ever wait is maybe 30 seconds. Is that really so bad?

If you just want to know "When it's the next full moon", yes. Very bad as google can answer in 1s.

That I just type into Google.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#164

Earlier quoted context omitted.

I have the opposite problem. I'm not well-calibrated on when I'd want lower reasoning than what's available to me (and how to compare that to lower-tier models). OpenAI now has Luna, Terra and Sol, each at Low, Medium, High and Xhigh, with Pro/Ultra depending on harness and plan. That's ~15 possible combinations of model and reasoning level, and there isn't a satisfactory explanation of which one you want for any par…

I don’t see why they just don’t allow a smaller model to answer the question while letting the bigger one vet it. The vetting can be asynchronous and can be delivered after a few seconds (if it’s an easy query). If it’s a hard query, the UI can show the answer is currently being vetted or something.

if the bigger model can vet it fast it can also answer it fast.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#165
post #141

Earlier quoted context omitted.

I'm not sure Google has lower talent costs. What makes you think so?

You are right to question that. I am basing that statement off of the many famous engineers that have recently left and were very highly paid. My guess is Google is ceding the frontier and the high salaries that go with it and letting Anthropic/OpenAI fight over the high priced talent that exit. The remaining non-famous engineers will not be able to command celebrity salaries. But this is just speculation and I don't…

Agreed. One effect I had in mind was much simpler: if you are the cool new company, people will want to work for you and might even take a hit in salary to do so.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#166
post #105

Earlier quoted context omitted.

How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).

Yes there are a few sites with IQ test benchmarks. The frontier models come up around 130 or 140 depending on which model/test.

It also looks like they're saturating the test, with one LLM hitting the maximum possible score. (https://www.trackingai.org/home)

The test wasn't made to accurately measure IQs that high.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#167
post #73

Earlier quoted context omitted.

And after half an hour using it he'd just admit that his test was way too simple as these models are still way too dumb

How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).

I wonder if it’s a version of Dunning-Kruger effect to call AI models dumb. I haven’t seen a “dumber than me” model since years. Also the smartest people known in the world use them in their fields so I don’t know what is meant by a “too dumb” model.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#168
post #73

Earlier quoted context omitted.

And after half an hour using it he'd just admit that his test was way too simple as these models are still way too dumb

How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).

I just asked Fable 5 max to create a 2d game about caterpillar climbing a tree and eating fruits. The game looks good - animations, 8-bit aesthetics, procedural tree branching, but the tree's branches are dead ends. You can't go back once you started climbing a branch. Yes, LLM doesn't have a reliable way to test it's game yet. All the screenshots, and playwright tests will never be enough to test even a simple game. But can we call a machine doing such mistakes a general intelligence? It has no embodied intelligence. No way to experience time the way we do. All it has is text. Yes, they can have images, sound too, but no big models (at least those we are supposed to use for coding) currently are native with video as far as I am concerned. And I am not sure that just video without embodied experience is enough to understand the world the way humans do. Of course we can get incredible results from machines that have a very different experience of the world than we do. But is this a general intelligence? I guess "general" is supposed to mean being able to do everything any human can do (minus the skills requiring a body)?

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#169
post #73

Earlier quoted context omitted.

And after half an hour using it he'd just admit that his test was way too simple as these models are still way too dumb

How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).

I think it's Karpathy who coined the term “jagged intelligence”. LLMs are both extraordinary smart in domain they have been explicitly trained on (like Math) and positively dumb on things they haven't.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#170
post #105

Earlier quoted context omitted.

Yes there are a few sites with IQ test benchmarks. The frontier models come up around 130 or 140 depending on which model/test.

That tracks, I wonder what the people who say that the models are too dumb expect to see. Miracles?

Not making mistakes my four years old would not.

They are definitely smart enough to be useful, but dumb enough in their weak spots not to deserve the "general intelligence" qualifier.

Post reply on HN