Live data from Hacker News

GPT-6 Astra

openai.com

731–740 of 1001 posts

Re: GPT-6 Astra

#731

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Its AGI when it can fit years of information in the context window.

Re: GPT-6 Astra

#733
post #373

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

Make cool stuff because the process is fun and makes you learn?

Re: GPT-6 Astra

#734
post #373

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

The process was for you, the product was for the world.

Now it's just the product for the world, which was where most of the value was anyways.

It's a big paradigm shift and the industry is quickly going to shed people who needed the process to care about the product and we'll be left with people whose motivation to build the product (or money) is enough.

Re: GPT-6 Astra

#735

Pelicans please

Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.

Hah, this is a new one: first time there's been a complaint about the pelican before I've even posted one!

(I don't have access yet.)

Re: GPT-6 Astra

#736
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.

So the closed source application should open its source in near future?

[15] https://arxiv.org/abs/2608.11469v1

Re: GPT-6 Astra

#737

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

Plan, sure, many people ask this of AI already, but not actually ask it to buy autonomously.

Re: GPT-6 Astra

#738

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

[dead]

Re: GPT-6 Astra

#739
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

> If this is truly AGI (subject to one's definition of AGI still) Scoring well in a benchmark that's called AGI does not make an LLM AGI.

If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.

Re: GPT-6 Astra

#740
Through various comments here there is a clear confusion on what AGI means.

Can someone point to a definite clarification?

Is it:

A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)

B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?

C) “The Singularity” (whatever that is?) so that AI can now do ____?

Someone please clarify for me!

Post reply on HN