Live data from Hacker News

GPT-6 Astra

openai.com

721–730 of 1001 posts

Re: GPT-6 Astra

#722

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint. I wouldn't w…

There was amazon dash button for this

Re: GPT-6 Astra

#723

AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me. It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.

AGI to me means intuition and I don't think that's ever going to happen with a LLM.

Re: GPT-6 Astra

#724

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

Codex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.

Put it in an isolated container and set it to YOLO

Re: GPT-6 Astra

#725
post #513

I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?

They're gaming benchmarks HTH

Re: GPT-6 Astra

#726

I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.

Wise decision.

Re: GPT-6 Astra

#727

So they’re copying Gemini with the whole star motif? I guess it makes sense they are unoriginal. like Zuck, @sama never invented anything or innovated at all - just took other people’s ideas

"distilling"... ;)

Re: GPT-6 Astra

#728

openai vs anthropic. that's it right? anyone else?

It's more like OpenAI vs no one, at this point. Anthropic has shown they don't care about general consumers or small/med businesses. You can't even use their models without it giving refusals on the most mundane tasks.

Re: GPT-6 Astra

#729
“Humanity’s Last Exam”?

“ARC-AGI-3”?

Is your bullshit detector going wild? Good, it’s working!

How is this not the most cringe marketing strat in history???

Post reply on HN