Live data from Hacker News

GPT-6 Astra

openai.com

771–780 of 1001 posts

Re: GPT-6 Astra

#771
post #561

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Chollet writes he expects AGI now sooner than 2030, "given progress is happening faster than I expected." https://x.com/fchollet/status/2095607046129463577

How is that an argument to the comment you replied to?

Re: GPT-6 Astra

#772

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries

Maybe this is why they NEED AGI (does it come with a soul?)

Re: GPT-6 Astra

#773

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and nev…

- Amazon is full of AI books, and they're clearly making money. AI has won multiple literary and artist awards.

- Okay, it's not a "new company" idea, but VendingBench is all about ability to run a company

- Plenty of people disagree with you on conversational quality; see "AI Boyfriends" etc.. (and it's not hard to find people who consider it uniquely valuable for discussing mental health)

- "come up with its own ideas or theories that nobody else has presented" C'mon, seriously? Solving a half-dozen hard open math problems wasn't enough there? What the heck counts as "it's own ideas or theories" at this point?

- plenty of evidence that custom models are starting to do well on the stock market, although I'll admit we're a year or so from any solid proof, since you need a track record to really make the claim

- LLMs have been capable of being a GM for a TTRPG for over a year (although like humans, they make mistakes)

- Okay, conceded, but humans tend to take years and large teams to make a game. Even if the capability existed today, it would take a while to actually build, test, market, etc.. - all made much more complicated by gamers being largely opposed to AI art styles, etc..

- "be able to sort through research and come to conclusions on complex geopolitical/sociological topics" - uh... did you mean to say something else, because "come to conclusions" is... like, LLM 101?

- Hahaha, have you met humans? We definitely cannot do that.

- Uh... thinking about it's own thinking is trivial. Most LLMs these days are built using LLMs, so uh, self-optimization seems nailed, too? We just don't let them do it unsupervised.

- LLMs fucking love to wonder about things

- "observe contradictions and ironies in the social-consciousness" really seriously have you actually used an LLM recently? I think you would find it remarkably enlightening.

Re: GPT-6 Astra

#774

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

Plan a holiday, definitely. There are many esoteric things one has to research to properly plan a holiday that I’d rather just not. Some things require reservations months in advance and I’d rather an AI just figure all that out for me ahead of time.

Re: GPT-6 Astra

#775

AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me. It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.

AGI to me means intuition and I don't think that's ever going to happen with a LLM.

How do you define intuition?

Re: GPT-6 Astra

#776

> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively. So the closed source application should open its source in near future? [15] https://arxiv.or…

I was listening to the Lex Fridman / DHH podcast last night [0], and DHH was saying that this is a new era for open source software. I'd agree, and also extend it open hardware.

Recently I've seen quite a few posts from people using AI to reverse engineer the Bluetooth protocol or such on devices that need a proprietary app. The same thing for firmware is surely coming, which is great as you can get lots of fun hardware from China, but it often has shitty firmware. Once that becomes the norm there's no reason not to make it open in the first place.

[0] - https://open.spotify.com/episode/45lhw2Adbrsw0xSCOgIeg3?si=r...

Re: GPT-6 Astra

#778

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

And without this harness it scores about 62%+, a dramatic improvement over even Fable 5.1 at 30%. I thought it just bears saying for context.

Re: GPT-6 Astra

#779

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries

There's a lot of truth to this: Execs want their own use case covered. In addition, everyone wants to be better than the competition at basic use cases, because that's what people compare first - even though we all know they are not a good reflection of reality. I'm currently working on exactly a project like this.

Re: GPT-6 Astra

#780

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and nev…

To me all this makes the label of AGI completely meaningless.

What AGI has always meant (eg. in 2019) is Artifical General Intelligence.

Artificial -- something made by humans instead of occurring naturally

General -- not confined by specialization or careful limitation

Intelligence -- the capacity to learn, reason, solve problems, think abstractly, and adapt to new situations

Basically, the metric was that any healthy adult human on the planet represents a general intelligence. This has certainly long been reached.

Also some of the stuff you're listing has long been solved as well, such as listing what it knows and what it doesn't know, and what information it would need. Other is just poorly defined: "be able to argue persuasively". AI can certainly write an argument on almost any topic that would pass any University homework in 2019.

Post reply on HN