Live data from Hacker News

Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

openai.com

181–190 of 281 posts

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#181

> For Plus and Pro users, we’re updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers. I was not impressed with 5.6 and this hits exactly why. Also, this instant, medium, and high slider situation we now have everywhere is batshit crazy. It’s a major step back in technology. I’ll put money on the table there will be a surprise in revenue because users have no clue what to choose…

It seems even worse than that: when you're calling the model, you have this strange two-dimensional thing with reasoning effort (specified in a small handful of random strings like "medium", "xhigh", "max") and "pro" vs "not pro" which is also somehow increasing reasoning budget, but with an interaction that is entirely opaque. How does "medium", "pro" compare with "max", "not-pro" for instance?

If you're going to force people to specify manually, at least make it 0.0 - 1.0 normalised such that 0.5 is the default.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#182
post #171
post #23

> Our mission is to ensure that artificial general intelligence benefits all of humanity. We’re introducing updates to ChatGPT that improve everyday conversations while expanding access for Free users. This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud. Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help f…

It’s interesting that they haven’t declared AGI yet, even as a PR stunt. It can be like a “pre-revenue” tactic. They are pre-AGI so investors can still pour money.

Public models are months behind what the labs have internally, and I believe OAI's upcoming model is so advanced they consider it AGI.

OAI is not going to announce anything until their next model officially releases (allegedly later this month).

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#183

For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show. It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?

And it hides again and reverts to instant when you are away from the app for a while. Definitely annoying.

It's getting seriously annoying. This is obviously in an attempt to push people to use the cheaper models but I find them appallingly stupid. There was a worrying comment by Tibo recently where he implied that they're aiming to meter chat queries against one's subscription quota.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#185
post #178
post #23

> Our mission is to ensure that artificial general intelligence benefits all of humanity. We’re introducing updates to ChatGPT that improve everyday conversations while expanding access for Free users. This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud. Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help f…

ok so first the industry takes the AI term, uses it with a very vague resemblance to what it used to mean, for marketing purpose, then they latch on to the AGI term to talk about what people used to consider AI. Now there is no sign of actual AGI happening any time soon, so we're going to reinterpret AGI to mean something diminutive like a chatbot? Don't you see a problem here? Terms are used to describe the world an…

> Now there is no sign of actual AGI happening any time soon

Do you genuinely hold this position, or do you not realize how far the goalposts have shifted?

In 2022, prominent AI critic Gary Marcus offered to bet $100,000 that we wouldn't have AGI by 2029. https://garymarcus.substack.com/p/dear-elon-musk-here-are-fi... Because the definition of AGI is unclear, he defined that AGI would be achieved if an AI model could do THREE of the five following tasks:

- In 2029, AI will not be able to watch a movie and tell you accurately what is going on (what I called the comprehension challenge in The New Yorker, in 2014). Who are the characters? What are their conflicts and motivations? etc.

- In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI.

- In 2029, AI will not be able to work as a competent cook in an arbitrary kitchen (extending Steve Wozniak’s cup of coffee benchmark).

- In 2029, AI will not be able to reliably construct bug-free code of more than 10,000 lines from natural language specification or by interactions with a non-expert user. [Gluing together code from existing libraries doesn’t count.]

- In 2029, AI will not be able to take arbitrary proofs from the mathematical literature written in natural language and convert them into a symbolic form suitable for symbolic verification.

Today's AI models can do FOUR of these five.

Using 2022 goalposts, we already have AGI. We blew past these goalposts months ago, and nobody noticed.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#186

Earlier quoted context omitted.

Claude's rate limits are absolutely shit for free users. 3 chats and its gone for 24 hours.

I dont see that, and I have a lot of chats, long messages, and is not 3 chats.

It has limits when there's an attachment in that chat session. Switch to a new chat and you're good to go

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#187

> For Plus and Pro users, we’re updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers. I was not impressed with 5.6 and this hits exactly why. Also, this instant, medium, and high slider situation we now have everywhere is batshit crazy. It’s a major step back in technology. I’ll put money on the table there will be a surprise in revenue because users have no clue what to choose…

It seems even worse than that: when you're calling the model, you have this strange two-dimensional thing with reasoning effort (specified in a small handful of random strings like "medium", "xhigh", "max") and "pro" vs "not pro" which is also somehow increasing reasoning budget, but with an interaction that is entirely opaque. How does "medium", "pro" compare with "max", "not-pro" for instance? If you're going to fo…

OpenRouter map this into their API in a slightly different way, making -pro and normal different model slugs, so you have openai/gpt-5.6-luna-pro vs openai/gpt-5.6-luna vs openai/gpt-5.6-sol[-pro], etc. This makes reasoning effort one-dimensional again (albeit with the arbitrary sequence of strings), but now model choice in a given generation is two-dimensional. Either way, it's hard to make any kind of informed decision.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#188

For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show. It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?

I think thats a difference between Chat and Work. In work mode you can select a model and chat should feel "instant". But I also don't like that differentiation. It's already hard to know which thinking level is necessary, how should users know which mode to use?

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#189
post #167

Earlier quoted context omitted.

How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).

I wonder if it’s a version of Dunning-Kruger effect to call AI models dumb. I haven’t seen a “dumber than me” model since years. Also the smartest people known in the world use them in their fields so I don’t know what is meant by a “too dumb” model.

You need to be smarter (or rather: more knowledgeable in the problem domain) than the model to be able to use it efficiently. Hallucinations are still a problem occasionally but a bigger one is failure of imagination. Even Claude Fable lacks a holistic understanding of many domains it wasn't obviously trained on. The biggest problem with AI (if we assert that LLMs can be the basis of AI) is that these models will make mistakes that exist in an entirely different category of the kind of mistakes humans will make.

As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.

Re: Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

#190

I dont get why they'd make this available to free users when they're probably dealing with compute constraints considering their competition with the Chinese models. Giving a more capable model to a huge free user base looks like an expensive choice to me

Even though it seems implicit that all of these companies are losing money hand-over-fist, it's still a competitive market.

Improving the free offerings is meant to increase visibility and therefore market share. It's not about goodwill, and it never will be. :)

The free stuff is primarily marketing and marketing always has costs.

In terms of compute, I have no way to really look behind the curtain and see what goes on back there. But I know with codex CLI, in terms of weekly quota: I can get a ton of work done with luna and usually get reasonable results. It feels very compute-light in this way.

I'm amazed by the work luna on xhigh can do for the price I pay (just $20, every month). It has the presentation of something that is very efficient to run, while also being something that can actually produce OK results. It's also fairly quick.

It differs from many previous smaller offerings of yore in this way. Like, I mean: I found stuff like the -mini models and 4o to be utterly useless wastes of my time. Luna isn't like that at all; it can get some stuff done.

So far for me, luna is the most impressive part of the 5.6 rollout. Not because it is best, but because it is useful and cheap.

So if luna is decent (it seems that it is), and if it is in fact light (which seems to be true from what I can observe), and it is offered for free, then it may very well be better, faster, and cheaper than the competition is.

And that's good for visibility. Marketing is all about buying eyeballs.

Post reply on HN