Live data from Hacker News

GPT-6 Astra

openai.com

761–770 of 1001 posts

Re: GPT-6 Astra

#762

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

>what would make you think Astra is yet to be AGI...

Can it detect if I feed bullshit (by bullshit I mean stuff that contradicts with its own existing "knowledge") in its training data? If not, then I think it is a good indicator that it is not intelligent at all, let alone AGI...

And I think discussions on whether these models are AGI or not are AI marketing triggered. And that is exactly what these statements are targeting....HN appear to have fallen for it, as usual...

Re: GPT-6 Astra

#763

For people skeptical of AGI. Consider the following: 15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role. I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-fi…

For people delirious about AGI, you don’t get to it by redefining it favorably.

Re: GPT-6 Astra

#764

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?

I feel like I'm completely missing something with ARC-AGI. The tasks are so limited in scale and very black and white, which do not at all map to real-life challenges.

I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.

Re: GPT-6 Astra

#765
post #662

- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelli…

> Humanity's Last Exam (w/ tools)

This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.

Re: GPT-6 Astra

#767

The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.

I agree, but I also long considered llm's stochastic parrots. Then this year happened.

Opus/Sol are easily far smarter programmers than I, and this thing supposedly blows them out of the water. Once an LLM is a better doctor, researcher, biologist, chemist, mathematician, physicist than any human is that not AGI?

It didn't arrive in the form I would have ever imagined, but it's hard to say its not (imo).

Re: GPT-6 Astra

#768
post #686

> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively. So the closed source application should open its source in near future? [15] https://arxiv.or…

Not if OpenAI considers reverse engineering an offensive cybersecurity skill.

Surprisingly, I've had really good luck with reverse engineering on frontier models (without being part of CVP or similar). It's more the exploit development/PoC that it locks up on, which, as I use `pi`, I just switch the model to Kimi K3 to finish up making the PoC.

Ironically, due to the stringent guardrails on American models that exist to avoid giving adversaries a leg-up in cybersecurity, I end up feeding dozens of 0days straight to the CCP lol

Re: GPT-6 Astra

#769

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

And I could bet money on that in short time after someone builds that sort of system the next step is to make it worse. Push worse and more expensive options to user. Or at least those from highest bidder... Anyone involved just can't keep themselves honest so it is doomed to be exploitative.

Re: GPT-6 Astra

#770
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

> Like, what's the point, if the next AI can do it in 5 seconds? I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total. The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think…

Me too! (although I am hoping to release it, but it was primarily for me)

I was thinking about writing something and hopefully starting a community/resource area around hyper-personalized software, it's so easy to make now. I was thinking about how it'd be useful to have a place to share these, for ideation, sharing techniques and the ability for LLMs to riff on something already existing. It's not quite like open source's advantage of having many people contribute to the same project, but rather something that is closer to evolution, giving the next generation a place to start modifying.

Would you have any interest in sharing what you've made?

Post reply on HN