Live data from Hacker News

GPT-6 Astra

openai.com

831–840 of 1001 posts

Re: GPT-6 Astra

#831

The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.

I agree, but I also long considered llm's stochastic parrots. Then this year happened.

Opus/Sol are easily far smarter programmers than I, and this thing supposedly blows them out of the water. Once an LLM is a better doctor, researcher, biologist, chemist, mathematician, physicist than any human is that not AGI?

It didn't arrive in the form I would have ever imagined, but it's hard to say its not (imo).

Re: GPT-6 Astra

#832
post #746

> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively. So the closed source application should open its source in near future? [15] https://arxiv.or…

Not if OpenAI considers reverse engineering an offensive cybersecurity skill.

Surprisingly, I've had really good luck with reverse engineering on frontier models (without being part of CVP or similar). It's more the exploit development/PoC that it locks up on, which, as I use `pi`, I just switch the model to Kimi K3 to finish up making the PoC.

Ironically, due to the stringent guardrails on American models that exist to avoid giving adversaries a leg-up in cybersecurity, I end up feeding dozens of 0days straight to the CCP lol

Re: GPT-6 Astra

#833

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

And I could bet money on that in short time after someone builds that sort of system the next step is to make it worse. Push worse and more expensive options to user. Or at least those from highest bidder... Anyone involved just can't keep themselves honest so it is doomed to be exploitative.

Re: GPT-6 Astra

#834
post #367

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

> Like, what's the point, if the next AI can do it in 5 seconds? I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total. The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think…

Me too! (although I am hoping to release it, but it was primarily for me)

I was thinking about writing something and hopefully starting a community/resource area around hyper-personalized software, it's so easy to make now. I was thinking about how it'd be useful to have a place to share these, for ideation, sharing techniques and the ability for LLMs to riff on something already existing. It's not quite like open source's advantage of having many people contribute to the same project, but rather something that is closer to evolution, giving the next generation a place to start modifying.

Would you have any interest in sharing what you've made?

Re: GPT-6 Astra

#835

Anthropic in tears today.

Anthropic don't care, they don't want their products to be used by general audiences in any serious manner. Their interest is in selling to megacorps and using the models for themselves internally to swallow industry, and drumming up AI fear to regulatory capture to shut down the businesses that do want to make AI accessible to the people.

Yeah, but if OpenAI provides a cheaper model with the same capabilities the megacorps will buy from OpenAI. If you want to target enterprises you need to have some competitive advantage, it's not enough that "I wanna target them".

Re: GPT-6 Astra

#836

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

"Somebody is wrong on the internet" will get you more, and more reliable, results on what people actually want from AI. You are now participating in a very high-value survey.

Re: GPT-6 Astra

#837
post #367

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

So this what a deflation economy is like. People don't want to do anything because they feel like whatever they do will be worthless in the near future.

I've retreated to doing stuff with my hands. Wokdwork, DIY, that kind of thing. At least for now and the foreseeable future that doesn't seem pointless. Only problem is it's hard work yet nobody would pay me for it.

Re: GPT-6 Astra

#838

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Maybe because the people making them are workaholic types who really don't care? I've certainly been in situations where I didn't really care what shows up for a meal. Someone was tasked with getting food and we leave it up to them. Sure, I can imagine such a meal being bad and it has been a few times in my life but 24 of 25 times, maybe more, it's fine. Further, the AI knows your preferences.

Re: GPT-6 Astra

#839
post #619

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Chollet writes he expects AGI now sooner than 2030, "given progress is happening faster than I expected." https://x.com/fchollet/status/2095607046129463577

How is that an argument to the comment you replied to?

Re: GPT-6 Astra

#840

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries

Maybe this is why they NEED AGI (does it come with a soul?)
Post reply on HN