Live data from Hacker News

GPT-6 Astra

openai.com

921–930 of 1001 posts

Re: GPT-6 Astra

#921

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

and that's exactly what most people don't want. they want some ultra-intelligent being that can do marvelous things, and they can claim the credit on it

Re: GPT-6 Astra

#922
The anthropomorphic delusion about LLMs being a coherent entity, with memory and continuity (which is also a delusion even when it comes to humans), quickly evaporates when you realize that each message in a chat is a completely separate stateless API request, and the only trick fueling the illusion is that the previous messages are sent to the server along with your new message. So AGI or not, it’s just a process that exists literally for the duration of one API request.

Re: GPT-6 Astra

#923

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamenta…

Can you share some examples of this as shared conversations? I’d genuinely love to see this in practice.

Re: GPT-6 Astra

#924

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…

> Canceling my Anthropic Max sub when this ships.

At this point, it reads like people are cancelling old ones and getting new subscriptions every two to three days, whenever a new ,model drops, and quite possibly by the end of the week they are back to the old provider while still having active subscriptions with at least two to three others. Interesting times.

Re: GPT-6 Astra

#925

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI Give someone 10 remote employees for a few months, 5 of them human, 5 of them AI. After a few months, check to see if the humans (manager, other coworkers) can figure out who is AI and who isn't. Would that be sufficient? I'd have to think about it. But AGI is supposed have human level capabilities, so this would be a necessary prerequ…

Being as good as a human at everything isn’t the same as being able to masquerade convincingly as a human at everything.

A better benchmark would be seeing which cohort of employees the manager prefers employing after a few months.

Re: GPT-6 Astra

#926
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

„The depressing thing about tennis is that no matter how good I get, I'll never be as good as a wall.“ -Mitch Hedberg

Well people will still pay to watch a human player play tennis even if a robot could play infinitely better than that human. Try something like that with an employer and software engineer combo (and no, you don't even have to imagine that).

Re: GPT-6 Astra

#927
post #752

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

>what would make you think Astra is yet to be AGI... Can it detect if I feed bullshit (by bullshit I mean stuff that contradicts with its own existing "knowledge") in its training data? If not, then I think it is a good indicator that it is not intelligent at all, let alone AGI... And I think discussions on whether these models are AGI or not are AI marketing triggered. And that is exactly what these statements are t…

Anthropic's J-Lens research indicates so.

Like when web results are fake it would show things like 'FAKE PROMPT INJECTION'.

Or in a safety evaluation with a contrived scenario it was 'FAKE FICTIONAL'.

Re: GPT-6 Astra

#928

Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.

Its still human made frontier progress.

And for sure it has a tremendes amount of implications, but its not the fault of the technology (we found, not invented).

And i'm only living once, my main motivation is not to just live day in day out the same stuff, i'm quite happy to see progress.

Am i worried about the future of our planet? For sure.

Re: GPT-6 Astra

#929
post #571

Seems like only yesterday that gpt 5 was supposed to mark our downfall

Its hard to understand how exactly these AI people mean it.

When i say "AI is changning the world" its more like "I can already see how this technology will continue to become better and better and has more impact every single day. It already affects people and it will have fundamentally changed A LOT in 3-15 years"

But lets be very realistic and clear: My computer systems at home were exploitable a lot more often this year than any year before JUST because of GPT or Claude.

Post reply on HN