Live data from Hacker News

GPT-6 Astra

openai.com

691–700 of 1001 posts

Re: GPT-6 Astra

#691

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

[dead]

Re: GPT-6 Astra

#692
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

> If this is truly AGI (subject to one's definition of AGI still) Scoring well in a benchmark that's called AGI does not make an LLM AGI.

If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.

Re: GPT-6 Astra

#693
Through various comments here there is a clear confusion on what AGI means.

Can someone point to a definite clarification?

Is it:

A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)

B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?

C) “The Singularity” (whatever that is?) so that AI can now do ____?

Someone please clarify for me!

Re: GPT-6 Astra

#694

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

>I am reasonably confident that there's essentially nothing that I am better than Fable at

While this may be true, it’s a pretty poor indicator of whether or not it’s AGI.

Re: GPT-6 Astra

#695
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

I think we're getting to the point where it is difficult to identify the goal post of AGI.

Is it rapid skill acquisition? -> ARC benchmarks are saturated Is it breadth of knowledge? -> See many ... many benchmarks Is it ability to do hard tasks? -> see terminal-bench and released outputs.

We are at the point where the starting point for most tasks should be "send your agent to work on it."

So where do we draw the line in a way that doesn't move every 6 months?

Re: GPT-6 Astra

#696

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Watch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3 Then realize LLMs have zero of what anyone would consider intelligence.

I decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics.

Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness.

I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited.

So I don’t know why it can track fib algo, but no chess concepts.

Re: GPT-6 Astra

#697

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint.

I wouldn't want it to pick food for me from a place I've never been, though to be honest with enough order history it could probably do a decent job at it.

Re: GPT-6 Astra

#698

“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

It's worse than that, someone else generated it using and then put it on a static page. We just have to take their word for it that GPT6 can do this. It probably can. It's not really an impressive test anymore. Claude Fable can do it. Opus can do it. I've been making one-shotted games with models for a while now, to test out their capabilities, and they all pretty much come out like this - generic bland and basic, us…

I love your game. It's wonderful and exactly the sort of thing that would showcase something interesting as opposed to just copying what's already out there. It is something I could share with my kids, and exactly the right note of fun and exploratory in a unique and even natural way. It could be extended and played with.

I usually roll my eyes when I see a comment like this because rarely do they make the points they claim to make, but I see what you're getting at. They just chose to clone someone elses work and do it in a boring way. I like OpenAI's models a lot, but they should do better.

edit - just a sidenote that I hadn't looked at the games, I just took the comment about "super-mario cart" at face value. I stand by my points 110% (even moreso perhaps), what they're showing is more polished than I expected, I assume they spent a lot of tokens on it. It is a legit shame they couldn't have spent time thinking of a better idea to illustrate something just as polished, but more interesting.

Re: GPT-6 Astra

#699
post #367

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

Make cool stuff because the process is fun and makes you learn?

Before it was fun because I was learning useful things for the future.

Now it feels like whatever I learn will be obsolete in 2 months.

Re: GPT-6 Astra

#700

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

The rich fucks who run the show do.
Post reply on HN