Live data from Hacker News

GPT-6 Astra

openai.com

731–740 of 1001 posts

Re: GPT-6 Astra

#731

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks

I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamental, lie or give bad or dangerous advice.

Re: GPT-6 Astra

#732
These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.

Re: GPT-6 Astra

#733

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month?

I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and the ability to prioritize correctly, i.e. tell which of the n things it could do next is the one that actually matters.

Can't say for sure that those are enough, but not having them seems to be most of why I still have to "babysit" these incredible tools.

Re: GPT-6 Astra

#734

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

If only we could agree on what AGI is.

Re: GPT-6 Astra

#735
I guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.

Re: GPT-6 Astra

#736

The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.

I'm not sure how you call a model AGI without it learning new information at the model level and not introducing nasty surprises (both from adversarial users and unintentional badness)

I guess there is fine tuning (and RAG) for those that need something bigger than just the knowledge contained in the context.

Re: GPT-6 Astra

#737

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Let me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.

Per the Fireship video, it was less that the answers were in the training set and more that the ability to calculate the flag on Exploitbench was left in from the previous test that went awry.

Scoring 100% is easy if noone checks your work

https://youtu.be/0Rp9KJCEIvg

Re: GPT-6 Astra

#738

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

To me AGI has always meant sentience. And only since we’ve discovered that you can have something that is intelligent without it being apparently sentient that we’ve changed the definition to being, I suppose, more exactly aligned with the namesake. A real AGI, like the ones from science fiction, would make Astra look like a child’s toy. And I guess more concretely I would expect it to inhibit the following propertie…

> To me AGI has always meant sentience.

Sentience and intelligence are different things. Many humans and animals are pretty dumb, yet they are sentient. AI is intelligent, but non-sentient (hopefully!).

Re: GPT-6 Astra

#739
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

lol have you actually tried to build anything useful e2e? Leaving the AI to itself gives horrendous results.

Re: GPT-6 Astra

#740
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.

How do you imagine a post work society where people can have inequality? I don't want to be equal. I want to do more.
Post reply on HN