Live data from Hacker News

GPT-6 Astra

openai.com

711–720 of 1001 posts

Re: GPT-6 Astra

#711
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.

So the closed source application should open its source in near future?

[15] https://arxiv.org/abs/2608.11469v1

Re: GPT-6 Astra

#712

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

Plan, sure, many people ask this of AI already, but not actually ask it to buy autonomously.

Re: GPT-6 Astra

#713

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

[dead]

Re: GPT-6 Astra

#714
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

> If this is truly AGI (subject to one's definition of AGI still) Scoring well in a benchmark that's called AGI does not make an LLM AGI.

If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.

Re: GPT-6 Astra

#715
Through various comments here there is a clear confusion on what AGI means.

Can someone point to a definite clarification?

Is it:

A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)

B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?

C) “The Singularity” (whatever that is?) so that AI can now do ____?

Someone please clarify for me!

Re: GPT-6 Astra

#716

Earlier quoted context omitted.

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and nev…

By this definition, even most humans would not qualify as having AGI though.

They can't, by definition, have the A part btw.

Re: GPT-6 Astra

#717

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

>I am reasonably confident that there's essentially nothing that I am better than Fable at

While this may be true, it’s a pretty poor indicator of whether or not it’s AGI.

Re: GPT-6 Astra

#718

Earlier quoted context omitted.

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and nev…

You want a computer program to be able to take a single phrase and execute decade long journies? Who will be responsible for the outputs and side effects of such a closed loop system? Half of those the agent fleet systems can do right now. These are things it cant do and will not be able to do without human labor and long running human vision: https://rcsnyder.github.io/open-frontier-curriculum/05-front... https://rc…

>You want a computer program to be able to take a single phrase and execute decade long journies? > Who will be responsible for the outputs and side effects of such a closed loop system?

Itself. That's the point. We can do it. Until it can met that bar, it ain't AGI. That's always been the bar.

Re: GPT-6 Astra

#719
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

I think we're getting to the point where it is difficult to identify the goal post of AGI.

Is it rapid skill acquisition? -> ARC benchmarks are saturated Is it breadth of knowledge? -> See many ... many benchmarks Is it ability to do hard tasks? -> see terminal-bench and released outputs.

We are at the point where the starting point for most tasks should be "send your agent to work on it."

So where do we draw the line in a way that doesn't move every 6 months?

Re: GPT-6 Astra

#720

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Watch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3 Then realize LLMs have zero of what anyone would consider intelligence.

I decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics.

Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness.

I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited.

So I don’t know why it can track fib algo, but no chess concepts.

Post reply on HN