Live data from Hacker News

GPT-6 Astra

openai.com

921–930 of 1001 posts

Re: GPT-6 Astra

#921

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

Hi thanks for the insight. Do you see a role for Control Systems(i.e. ones analogus to Instrumentation engineering) playing a role to modulate certain parts of continual learning? One very important way we learn are lived experiences, it's like telling memory:this part is more important( for emotional or social utility values), pay attention. Good or bad lived experiences both count. I guess is that a path that practical research is considering?

Re: GPT-6 Astra

#922

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

I think our AI systems are essentially massive Central Executive Networks. But novel ideas (creativity) come from the Default Mode Network.

These are the difference in what Kahneman called System 2&1 thinking and what the ancients called the Ratio and the Intellect.

LLMs are all ratio. They depend on our intellect for guidance.

Re: GPT-6 Astra

#923

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Agree. Token predicting machines will continue to be token predicting machines by nature. Continued size and tuning will have the effect of making them more and more perfect at being average.

They learned generic concepts like our brain does to optimize for this particular surprisngly perfect task:

You have to be able to respond to a very generic question in a way that the other entity thinks this is good, comprehensive, etc.

You can call us situation predicting machines as well if you want.

But you undermine what the latent space of an LLM is representing.

Re: GPT-6 Astra

#924
Why are OpenAI so keen to start calling things AGI?

Isn't there some corporate/legal shenanigans where they become a real non-profit at that point? Or does it just let them cut Microsoft and other investors out?

There's got to be a business reason for it unrelated to the model capabillities.

Re: GPT-6 Astra

#925
post #861

I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right d…

AI models do not live and learn - it's worse. They actually get DUMMER if you don't start with a clean slate. This is important. One has to curate the context carefully.

Dumber.

Re: GPT-6 Astra

#926

That hero video is interesting. A projector and speech. Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now. The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for m…

>this could bring us closer to the dream of more natural, social computing What I saw was multiple people living alone in a small box in a warehouse (probably filled with other boxes) with all of their natural, social interactions directed at a wall. I wonder if this is foreshadowing for the future of work, at least it is what work will look like as envisioned by OpenAI.

this observation is insightful and worth calling out. this paints the picture of a future I'm not excited about

Re: GPT-6 Astra

#927
I firmly believed OpenAI would take back the lead. It's healthy to have competition with Anthropic and, hopefully, other LLM providers. I'm excited to test their model. Their attitude towards developers/builders has been nice and appreciated for the last couple of months.

Re: GPT-6 Astra

#928
post #640

- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelli…

Used curiously fewer tokens, however.

Re: GPT-6 Astra

#929
post #721

These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.

not a chance its gonna be realtime lol. It's fun marketing though, I'll allow it

Re: GPT-6 Astra

#930

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…

Why would I ask a LLM a rhetorical question
Post reply on HN