Live data from Hacker News

GPT-6 Astra

openai.com

891–900 of 1001 posts

Re: GPT-6 Astra

#891

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…

Interesting. I actually prefer when agents don't commit on my behalf unless I explicitly say so, I even had to add a custom instruction for Claude to stop doing it (Codex never does it). Even if I don't read all the code line-by-line, I at least want to see the changes at glance and commit myself. Git Fork[0] is a great tool for that, by the way.

In general, I don't like when I have to prompt models to NOT do something. It's probably difficult for the AI companies to get this right, they should understand ambiguity but still not over-do simple instructions.

[0] https://git-fork.com/

Re: GPT-6 Astra

#892

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

> That's what the AI's really are terrible at -- creativity

Good thought piece here "We Are Losing the Ability to Discover What We Didn’t Know to Ask[1]" By Anne-Laure Le Cunff

It keeps playing on my mind as I see people at work follow some predetermined AI workflow to get their jobs done, the art of being curious and exploring around the problem is so important to the really big innovations. Been thinking about how to address this through some of the harnesses we are developing in the knowledge working space.

[1] https://archive.is/IAxf9

Re: GPT-6 Astra

#893

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

This comment and the one above from astrobiased feel like coming into a messy codebase, and it’s more work to sort it out than it would have been to write it from scratch.... And since I actually do intelligence testing as a clinical psychologist, I have experience with this in both practice and theory. So now I’m going to waste an hour because I just have to respond to “something is wrong on the internet.”.

Chollet's distinction is useful. High performance on known tasks is not the same thing as efficient adaptation to a novel task. Prior knowledge and training data can buy skill. That is a central point of On the Measure of Intelligence. But it does not follow that current frontier progress is only "coverage-driven competence." That is a hypothesis. It is not a result established by Chollet's framework.

"Overfitting at scale" is also the wrong term. A model that learns broad representations and applies them successfully to unseen examples is generalizing. The relevant concern is whether apparent novelty is actually inside the effective training distribution, not whether the model is "overfit."

There is also an unstated premise here: that adding broad knowledge and skills cannot improve the machinery used for novel problem solving. I do not see a basis for assuming that. Learned representations, abstractions, reasoning patterns, and cross-domain analogies can themselves support transfer to new tasks. Whether this becomes sufficient for general intelligence is an open question with insufficient data. But its a perfectly valid hypothesis right now that, given enough domain knowledge and symbolic reasoning examples, LLM COULD maybe "Grok" AGI at a certain critical threshold.

And ARC-AGI-3 was specifically designed around novel abstract environments that require exploration and adaptation. Astra scores 99.9% with OpenAI's context-preserving Provider Adapter, and ARC reports that Astra constructed compact symbolic models of unfamiliar environments. That does not prove AGI, but it points in that direction more so than the other way around.

Gc roughly maps to acquired knowledge. Gf roughly maps to reasoning in relatively novel situations. Naming those two categories does not tell us whether increasing acquired knowledge and learned abstractions in an AI can improve Gf-like behavior. That causal question is exactly what is disputed.

And "Frontier models probably have maxed out crystallized intelligence" is just obviously wrong, unless you think they have been able to dig up every a scrap of paper with knowledge/information on it in the entire world, AND that there is no more useful knowledge to be generated left in the universe.

And the statement that intelligence and creativity are independent is simply wrong. A meta-analysis of 112 studies and 34k participants found a positive correlation of about r .25 between intelligence and divergent thinking. It also found that using g, Gf, or Gc did not eliminate that relationship. Creative achievement has a smaller but still positive meta-analytic association with intelligence, around r = .16. These are distinct constructs, not independent constructs.

And this is just a bad take: "AIs are terrible at creativity". At best that depends on which creativity, and I think its straight up wrong. On divergent thinking tasks, the operationalization behind every ADHD study you could cite, LLMs score above most humans, with the top humans still ahead. If you means Big-C, paradigm-shifting creativity, that is a different construct and none of the ADHD evidence transfers to it.

And if I where to say what I subjectively feel and see.... I have ABSOLUTELY no idea how people can say that we are not seeing sparks of creativity from AIs already. If a PERSON produced some of the music, solutions or deductions that I have seen AIs do, people would have NO problem celebrating it as extremely creative.

And finally, the ADHD claim is also, at best, overstated and just as often debunked. There is some evidence that higher subclinical ADHD trait scores, often survey studies only, are associated with better performance on some divergent-thinking measures. But a review of 31 studies did not find a consistent creativity advantage for people with clinical ADHD, and it found no evidence of better convergent thinking.

Okay, I’m done… And nobody noticed that I’m not doing my job here.

Re: GPT-6 Astra

#894
post #882

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

They do new stuff all the time. Ask your AI to draw a gerbil riding a unicycle around Pluto and you'd get an image that hasn't been there before. If by genuinely new you mean without any help from past culture, do humans do that? For significant pushing the boundaries of knowledge stuff you maybe need different algorithms like AlphaGo move 37 or Alpha Fold protein folding. Though again how often do humans do that?

Humans are not prompted.

Re: GPT-6 Astra

#895

I thought intelligence was going to be democratized but you have to stay behind a 200$ plan. The trend hints that open models can also catch up

Democratized in the Classical Greece sense, where you had to be male, not a slave, and own land to get a vote.

Re: GPT-6 Astra

#896
post #882

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

They do new stuff all the time. Ask your AI to draw a gerbil riding a unicycle around Pluto and you'd get an image that hasn't been there before. If by genuinely new you mean without any help from past culture, do humans do that? For significant pushing the boundaries of knowledge stuff you maybe need different algorithms like AlphaGo move 37 or Alpha Fold protein folding. Though again how often do humans do that?

Humans do that by inventing algorithms like AlphaGo :)

In a more serious note, I think the person you are responding to meant "new" more in line with "novel".

Take a microfluidic chip, for example. Current AI systems can create new flow cell geometry, but cannot come up with the idea for a microfluidic flow cell itself.

Re: GPT-6 Astra

#897
post #329

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

They can't create innovative stuff. They can only create what they are trained on.

Re: GPT-6 Astra

#898
post #329

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

A project I've just started is to follow the Roguelikedev tutorial on a 32-year-old PC, using the period dev tools. I'm going to have to write everything myself, and I'm currently researching how to poke the graphics adaptor.

Why bother doing that when it would be vastly easier on a modern computer?

Re: GPT-6 Astra

#899

Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.

My boss is already joking that I'll be out of job.

Re: GPT-6 Astra

#900

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Agree. Token predicting machines will continue to be token predicting machines by nature. Continued size and tuning will have the effect of making them more and more perfect at being average.

Next-token prediction is a general paradigm, though. In principle, there isn't really anything a sufficiently advanced token predictor couldn't do.
Post reply on HN