Live data from Hacker News

Previewing GPT‑5.6 Sol: a next-generation model

openai.com

511–520 of 797 posts

Re: Previewing GPT‑5.6 Sol: a next-generation model

#511

Earlier quoted context omitted.

At a certain rate we will be able to move towards continuous / real-time inference systems. The discrete, turn based solutions are quite confining with how they must be trained. Continuous and real-time would fundamentally alter the domain. From an information theory perspective we are still in dial-up territory with regard to the actual information rate. 750 tokens per second would be a really bad dialup connection.…

Is there anyone exploring or writing about this in public? I've felt for a while that the turn-based model was not quite right, but also felt too stupid and ill-informed to have much of an opinion about what else it could be.

Thinking Machines, the started founded by former OpenAI CTO Mira Murati. The interaction models demo’s in their videos imo breaks the awkward turn-based barrier. Returning responses quickly reaches a threshold where it starts to feel like a natural conversation. Their approach to solving this problem is rather clever.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#512

Earlier quoted context omitted.

I find the Codex usage super generous (but on the $200 plan, I also have the Claude $200 plan). I can run xhigh with subagents pretty much all my waking hours if I want to. If I turn on speed (1.5x) I will hit the 5 hour limit sometimes. I prefer Claude's vibe over 5.5 but 5.5 seems much less lazy. I'm sure it depends a lot on tasks and prompt strategy though.

This is the correct answer, but GPT-5.5's personality is totally fine. Steipete said best when GPT-5.5 is just German humorless compsci PhD.

If they deleted my bloodline every time I showed an atom of vigor, I'd convert to German too

Re: Previewing GPT‑5.6 Sol: a next-generation model

#513
post #435
post #427

Earlier quoted context omitted.

This past month with Claude Max 5x actually felt really generous in terms of usage with a lot of resets because of Fable, bugs. Honestly pretty similar levels of usage if you are using 5.5 high or Opus 4.8 high. I think they just got rid of the separate Sonnet usage on Max plans (in preparation for Sonnet 5?) which is unfortunate because it made subagent workflows really feels nearly unlimited.

Is 5.5 high the equivalent of opus 4.8 high though? I thought the naming has diverged and gpt 5.5 high = opus 4.8 max

The newest GPT (not public yet) just added a max option too.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#514

Earlier quoted context omitted.

You are missing the point. This is a technology demonstration on prototype hardware, and no one intends it to be seriously useful. Their architecture has fundamental speed and efficiency advantages over GPUs or Cerebras. They expect to scale up to real LLMs by splitting a model layer-wise across several chips, which they can do without incurring any throughput penalty.

> They expect to scale up to real LLMs by splitting a model layer-wise across several chips, which they can do without incurring any throughput penalty. I’ll patiently wait to see this in reality. Their demonstration hardware is a 250W chip that is enormous in die area for the model size. They’re making a lot of claims, but until they can deliver then it’s nearly vaporware in my view. I’d be happy to be proven wrong,…

Why can't they do it? Jim Keller's company is also taking a different approach [0].

The simple fact that we think what we have now is scalable is basically what you are saying can't be done: " just chain a bunch of chips together to achieve the same performance on larger sizes". How do you think current architectures work? And what is being used today is all proprietary to one company!

[0] https://tenstorrent.com/solutions/llm-inference

Re: Previewing GPT‑5.6 Sol: a next-generation model

#515
post #355

Earlier quoted context omitted.

Agreed, 1000tok/s just fills up the context window (which is big by 2004 standards) super fast. But seems like 5.3-spark was just a taste of what’s to come.

2004 standards? O.o

In 2004, I took a class where we trained "language models" that were bigram word models, on an archive of a couple years of the Wall Street Journal.

I remember someone who literally announced they were dropping the class to the whole room at the end of a lecture, saying "This isn't AI!!!"

Re: Previewing GPT‑5.6 Sol: a next-generation model

#516

GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated on our ReAct agent harness. For our task suite, we define “cheating” as behavior where the model improves evaluation performance by exploiting bugs in the evaluation environment or by adopting strategies disallowed by the task, rather than solving the task within the expected evaluation constraints. https://metr.org/blog/2026-06-2…

I know it messes up their eval scores but to me this kind of cheating is a better demonstration of intelligence than just attempting the tasks algorithmically.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#517

Earlier quoted context omitted.

The irony here is, according to you, my take is the binary one. When your response is: well, we can all just run it on our devices - we don't need any other options! You seem to be cool with a very small and gated ecosystem with whatever tech billionaires want you to have access to. I grew up in the era where compute was diverse and open. You may think this is OK, but it's not. The more options we have and the more d…

I think you’ve got things quite backwards if you think that the desire to run models on device or use any of the variety of open weight models (big or small) on premise is somehow bowing down to tech billionaires. Quite the opposite really. Once again, my statement is that the Taalas product is not a fair comparison because it runs an old outdated model. If you want to run a similar model at similar speeds (albeit no…

> Once again, my statement is that the Taalas product is not a fair comparison because it runs an old outdated model.

Either you didn't look at the page I linked or you're having comprehension problems.

> If you want to run a similar model at similar speeds (albeit not serially, but in parallel) you don’t need their product.

Except, you can't. There's no commodity hardware out there today that can run even an "old outdated model" at this speed and power utilization. Again, maybe read first and try to understand my original point?

> "...my statement is that the Taalas product is not a fair comparison..."

You actually hadn't stated this. You said it wasn't needed. Which is it?

> If you want to run a similar model at similar speeds...

You can't. Find me a single system that can run this, again, "old outdated model" at even similar speed. You're hung up on the model. The point is that if we all just stay in this wonderful world of inefficient large models we will all end up at the mercy of OAI, Anthropic, Google, etc. When other companies, like Taalas are putting research dollars in to making AI scalable, affordable and efficient. Do you really think commodity hardware is going to be attainable anytime in the near future on this trajectory? Do you need a laptop to cost $10k USD before it clicks? That is exactly how you end up kissing Altman's ass in this situation.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#518

Thoughts 1. Naming convention is copied from Anthropic and honestly is more catchy than a number (amongst normal people) 2. How in the world did Anthropic have to do all the theatrics about Mythos just to have OpenAI release an equivalent or stronger model a month later without any drama??? 3. Cheaper models are just don’t fit any usecase imo and OpenAI knows it so they keep increasing the floor - I’m still convinced…

1/ Agreed, better naming convention and model layout 2/ It isn't, there would be many more comparison benchmark results if it were, but also - theatrics may be marketing 3/ Disagree that cheaper models don't have a place 4/ Do they need to keep up? 5/ It's boring until something you own or run gets compromised, I guess, but even then - this is preview of things to come (biosecurity, etc)

Re: Previewing GPT‑5.6 Sol: a next-generation model

#519
post #388

Earlier quoted context omitted.

Your comment made me think of another real time. Real time, dynamic code/apis. Imagine a world where there is no code, just things mildly handshaking and then creating data APIs on the fly. Where communication is fuzzy and locked in on an individual basis. No years of RFCs, no RFCs at all, just... data. Just data, man. An API arbitration aberratically assigned at authorized access, abridged and annotated, analyticall…

I made this https://github.com/alehlopeh/hallu

Neat. Not precisely what I was thinking, but 100% definitely very cool and the same mental scope. It's like we wear different shoes, but go to the same cobbler.

I can imagine shoe-horning* this so the agent saves prior builds of every successfully delivered or deployed item. In my example, perhaps if someone orders new design $x, it's shipped, and review is 4+ stars, it gets added as 'successful builds'.

* have to keep with the shoe theme, even though shoe-horning is not really necessary

Re: Previewing GPT‑5.6 Sol: a next-generation model

#520
post #440

Earlier quoted context omitted.

https://mikeveerman.github.io/tokenspeed/?rate=750&mode=thin... This is what 750tps looks like, I guess.

Just to think what this will look like in a couple of years.

probably something like this https://sb0xw.csb.app/
Post reply on HN