Live data from Hacker News

GPT-6 Astra

openai.com

911–920 of 1001 posts

Re: GPT-6 Astra

#911
post #333

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

„The depressing thing about tennis is that no matter how good I get, I'll never be as good as a wall.“ -Mitch Hedberg

Well people will still pay to watch a human player play tennis even if a robot could play infinitely better than that human. Try something like that with an employer and software engineer combo (and no, you don't even have to imagine that).

Re: GPT-6 Astra

#912
post #737

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

>what would make you think Astra is yet to be AGI... Can it detect if I feed bullshit (by bullshit I mean stuff that contradicts with its own existing "knowledge") in its training data? If not, then I think it is a good indicator that it is not intelligent at all, let alone AGI... And I think discussions on whether these models are AGI or not are AI marketing triggered. And that is exactly what these statements are t…

Anthropic's J-Lens research indicates so.

Like when web results are fake it would show things like 'FAKE PROMPT INJECTION'.

Or in a safety evaluation with a contrived scenario it was 'FAKE FICTIONAL'.

Re: GPT-6 Astra

#913

Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.

Its still human made frontier progress.

And for sure it has a tremendes amount of implications, but its not the fault of the technology (we found, not invented).

And i'm only living once, my main motivation is not to just live day in day out the same stuff, i'm quite happy to see progress.

Am i worried about the future of our planet? For sure.

Re: GPT-6 Astra

#914
post #560

Seems like only yesterday that gpt 5 was supposed to mark our downfall

Its hard to understand how exactly these AI people mean it.

When i say "AI is changning the world" its more like "I can already see how this technology will continue to become better and better and has more impact every single day. It already affects people and it will have fundamentally changed A LOT in 3-15 years"

But lets be very realistic and clear: My computer systems at home were exploitable a lot more often this year than any year before JUST because of GPT or Claude.

Re: GPT-6 Astra

#916

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

They can do new tasks with in-context learning but its obviously limited by context window

But what does it mean in practice? Obviously we humans also have a limited cognitive capacity.

Let me offer a thought experiment: Let's say that tomorrow we discover Atlantis, with a treasure trove of books about their culture and science, written in a dialect of ancient Greek that we know how to start to analyze, but no one can read fluently. And let's say that you are a billionaire really curious about their culture and want to converse with an "Atlantean expert" as soon as possible. Would you invest your money in a "we-hate-ai-slop(tm)" group of researchers who would abhor AI and instead delegate the books to a massive number of human grad students? Or in a small group of researchers who are willing to use AI agents to go over these? Or maybe just open a chat session with GPT-6 yourself immediately? What would most effectively assuage your curiosity?

Re: GPT-6 Astra

#917
post #163

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…

yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall). Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.

I still use Opus 4.8 for a lot of tasks because I can't stand the way it talks.

Re: GPT-6 Astra

#918

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

> That's what the AI's really are terrible at -- creativity. But I'd argue the vast majority of humans aren't very creative, with truly out-of-the-box ideas. Given that this is HN, and a non-trivial number of us have ADHD, creativity is positively correlated with ADHD.

I think it's hard to define creativity in the context of AI because they seemingly just make up new hyphenated terms for everything. Is that creativity? If not, what about when they do the same thing different ideas in the latent space?

If we say that simply nailing one concept to another isn't creativity, then AIs are incapable of creativity, while the vast majority of humans are incapable of creativity. This is just a long way of saying "0 AIs have creativity, 0.00001% of humans have creativity", and the difference between zero and a very small number is infinity.

Re: GPT-6 Astra

#919

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

Hi thanks for the insight. Do you see a role for Control Systems(i.e. ones analogus to Instrumentation engineering) playing a role to modulate certain parts of continual learning? One very important way we learn are lived experiences, it's like telling memory:this part is more important( for emotional or social utility values), pay attention. Good or bad lived experiences both count. I guess is that a path that practical research is considering?

Re: GPT-6 Astra

#920

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

I think our AI systems are essentially massive Central Executive Networks. But novel ideas (creativity) come from the Default Mode Network.

These are the difference in what Kahneman called System 2&1 thinking and what the ancients called the Ratio and the Intellect.

LLMs are all ratio. They depend on our intellect for guidance.

Post reply on HN