Live data from Hacker News

Previewing GPT‑5.6 Sol: a next-generation model

openai.com

101–110 of 797 posts

Re: Previewing GPT‑5.6 Sol: a next-generation model

#101
Thoughts

1. Naming convention is copied from Anthropic and honestly is more catchy than a number (amongst normal people)

2. How in the world did Anthropic have to do all the theatrics about Mythos just to have OpenAI release an equivalent or stronger model a month later without any drama???

3. Cheaper models are just don’t fit any usecase imo and OpenAI knows it so they keep increasing the floor - I’m still convinced task per capability is reduced with each release

4. How in the world would open source models keep up with the multi layer security? Either this security is all theater or we will finally see a ceiling in open source models because by definition they can’t have those protections

5. Cybersecurity things are boring to me because it’s all zero sum cat and mouse games

Re: Previewing GPT‑5.6 Sol: a next-generation model

#102
post #79

Waiting for @simonw to report on this, before I read and try it

You might be waiting a while, I'm not in that set of "a small group of trusted partners whose participation has been shared with the government".

The government doesn't have a Department of Vector Pelicans?

Re: Previewing GPT‑5.6 Sol: a next-generation model

#104
Some interesting stats here about the current landscape https://arena.ai/leaderboard/agent

Agent Arena (Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.)

Top 10, Highest rank to lowest

Claude Fable 5 (High), Claude Opus 4.8 (Thinking), GPT 5.5 (xHigh), Claude Opus 4.7 (Thinking), GPT 5.5 (High), Claude Opus 4.7, Claude Opus 4.6, GPT 5.5, GPT 5.4 (High), GLM 5.2 (Max)

Text Arena View overall rankings across various AI models in text-to-text tasks across math, coding, creative writing, and other open-ended domains.

Top 10, Highest rank to lowest

claude-fable-5, claude-opus-4-6-thinking, claude-opus-4-7-thinking, claude-opus-4-6, claude-opus-4-7, muse-spark, gemini-3.1-pro-preview, gemini-3-pro, claude-opus-4-8-thinking, gpt-5.5-high

Re: Previewing GPT‑5.6 Sol: a next-generation model

#106

Waiting for @simonw to report on this, before I read and try it

I would love to see a more descriptive review from simonw instead of just SVGs generations.

He is not an ML researcher or engineer, he is a passionate AI enthusiast blogger. He mostly does SVGs and other low effort checks (sometimes with major flaws, as people have pointed out a few times in the HN comments). Properly evaluating the model across all fronts requires a deep understanding of LLMs, how they work, the trade offs behind new architectures and the relevant research papers. It also takes a lot of time to build a proper evaluation framework so basically you can't just vibe code that if you want something that is solid.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#107
post #79

Waiting for @simonw to report on this, before I read and try it

You might be waiting a while, I'm not in that set of "a small group of trusted partners whose participation has been shared with the government".

I think that there are some OAI employees on Hackernews. I do believe that they should give access to ya, because after all it would allows us to generate pelicans :-D

What is the consensus on who becomes part of the said small group of trusted partners and if they weren't so opaque about it. I'd expect comparatively big names like Simon to be included within such but Alas its not reality.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#108
post #84

What happened to the nano/mini/standard/pro naming scheme, which worked perfectly fine and is intuitive to understand? Why does OpenAI insist on having the most inconsistent and confusing model and product names possible? I'm looking at you Codex.

It’s still easy to understand as the more capable the model the bigger the celestial body they’re named after.
Post reply on HN