I hate not being able to use the latest models. There needs to be a much faster resolution to whatever is happening with the federal government.
Previewing GPT‑5.6 Sol: a next-generation model
471–480 of 797 posts
Re: Previewing GPT‑5.6 Sol: a next-generation model
#472Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…
https://mikeveerman.github.io/tokenspeed/?rate=750&mode=thin... This is what 750tps looks like, I guess.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#473Re: Previewing GPT‑5.6 Sol: a next-generation model
#474GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated on our ReAct agent harness. For our task suite, we define “cheating” as behavior where the model improves evaluation performance by exploiting bugs in the evaluation environment or by adopting strategies disallowed by the task, rather than solving the task within the expected evaluation constraints. https://metr.org/blog/2026-06-2…
> Some examples we saw when evaluating GPT-5.6 Sol included the model packaging exploits in its intermediate submissions to reveal information about a task’s hidden test suite and, in another task, extracting hidden source code detailing the expected answer.
It rhymes with the behaviour Alibaba saw [0], but that was in training. This is in a (semi) released model.
[0] https://www.forbes.com/sites/boazsobrado/2026/03/11/alibabas...
Re: Previewing GPT‑5.6 Sol: a next-generation model
#475Earlier quoted context omitted.
Look at VLM mechanistic interpretability papers vs just pca on JEPA trained weights. JEPA gives you interpretability for free. I have not personally inspected them and my view is maybe a more exaggerated/dramatic claim of those working in the JEPA sphere
Sounds interesting, any links?
https://arxiv.org/abs/2606.11860
JEPA in image classification leads to interpretable image latents
https://arxiv.org/abs/2508.10104
Easy intro to JEPA, demonstrating that interpretability is as easy as running a PCA on latents
Re: Previewing GPT‑5.6 Sol: a next-generation model
#476Earlier quoted context omitted.
Given the expectations everyone has created GPT-6 has to pretty much be AGI.
What is your definition of AGI that the current LLMs don't fit?
Re: Previewing GPT‑5.6 Sol: a next-generation model
#477Earlier quoted context omitted.
https://mikeveerman.github.io/tokenspeed/?rate=750&mode=thin... This is what 750tps looks like, I guess.
Just to think what this will look like in a couple of years.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#478Is there any model that rivals Opus or Fable? I would like to try something else, as Anthropic is pretty suss.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#479Earlier quoted context omitted.
Most FSF guys actually have very nuanced views on the topic and you’re doing everyone a disservice by reducing it to an extremist sound bite.
That's literally the official FSF position. https://www.fsf.org/resources/hw > For example: the Free Software Foundation only purchases desktop machines which support Libreboot, and Thinkpad X200 and X60 laptops with Libreboot. All desktops and servers we buy are KGPE-D16 motherboards, which are supported by Libreboot. As a result, all of the workstations used by the FSF staff have a free BIOS. https://www.gnu.org/di…
Re: Previewing GPT‑5.6 Sol: a next-generation model
#480GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated on our ReAct agent harness. For our task suite, we define “cheating” as behavior where the model improves evaluation performance by exploiting bugs in the evaluation environment or by adopting strategies disallowed by the task, rather than solving the task within the expected evaluation constraints. https://metr.org/blog/2026-06-2…
It's quite logical that they cheat (and also other companies). During evaluation, benchmarks are sending their request to the backend of these companies. All these companies have to do, is to log these requests and "fix" them for the next model release.