Live data from Hacker News

GPT-6 Astra

openai.com

711–720 of 1001 posts

Re: GPT-6 Astra

#711
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.

Re: GPT-6 Astra

#712
post #227

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

People really believe in this AGI marketing?

Re: GPT-6 Astra

#715

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint. I wouldn't w…

There was amazon dash button for this

Re: GPT-6 Astra

#716

AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me. It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.

AGI to me means intuition and I don't think that's ever going to happen with a LLM.

Re: GPT-6 Astra

#717

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

Codex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.

Put it in an isolated container and set it to YOLO

Re: GPT-6 Astra

#718
post #507

I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?

They're gaming benchmarks HTH

Re: GPT-6 Astra

#719

I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.

Wise decision.

Re: GPT-6 Astra

#720

So they’re copying Gemini with the whole star motif? I guess it makes sense they are unoriginal. like Zuck, @sama never invented anything or innovated at all - just took other people’s ideas

"distilling"... ;)
Post reply on HN