Live data from Hacker News

GPT-6 Astra

openai.com

941–950 of 1001 posts

Re: GPT-6 Astra

#941

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Define novel intelligence in a way that would not exclude 95% of humans, yourself included.

It gets kind of out there, but what i often hear peopel refer to is that frontier models lacks the visdom component. Which I guess is in the realm of intuition, i.e. i have a feeling it might be a problem with X based on some vague signs, maybe something a colleague mentioned offhand, something that was out of alignment etc.

Re: GPT-6 Astra

#942

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Define novel intelligence in a way that would not exclude 95% of humans, yourself included.

I remember vaguely from a presentation by Yann LeCun "Intelligence is not what you know, but what you do when you don´t know". I find it helpful when trying to build an intuition for how to understand the LLM tool.

Re: GPT-6 Astra

#943

“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

Have you seen ads for mobile games, where there are seemingly 100 different versions of the same type of game (like tower defense types)? And they're all obviously the worst type of pay-to-play traps? I think this kind of solves that. Or at least it is the start of it. Most such games are kind of trivial. If people can easily just get AI to generate such games on the fly, then that'll hopefully be the end of predator…

The problem is, no one wants to play a game they designed themselves

Re: GPT-6 Astra

#945

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Agree. Token predicting machines will continue to be token predicting machines by nature. Continued size and tuning will have the effect of making them more and more perfect at being average.

Re: GPT-6 Astra

#946
post #609

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Chollet writes he expects AGI now sooner than 2030, "given progress is happening faster than I expected." https://x.com/fchollet/status/2095607046129463577

Ridiculous

Re: GPT-6 Astra

#947
post #939

I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right d…

AI models do not live and learn - it's worse. They actually get DUMMER if you don't start with a clean slate. This is important. One has to curate the context carefully.

Re: GPT-6 Astra

#948
post #939

I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right d…

> It does not matter how smart the model is, it will make mistakes

It does matter, otherwise why are we all using GPT 5.6 rather than GPT 3.5? Because it's way smarter, makes less mistakes and therefore finishes tasks faster.

The smarter the model is, the faster it can complete what you actually wanted.

> Regardless, working on the wrong things is time wasted

Agreed. And "smart" for me, would mean understanding what is the right thing to work on vs the wrong thing, so a smart model would waste less time, thinking like this.

Re: GPT-6 Astra

#949

ASTRA means tool ( for war or attack specifically ) in hindi

It's basically impossible to pick a word that doesn't mean something unintended in 20+ major languages in the world.

The recent OpenAI model names are obviously based on Latin: Luna (Moon), Terra (Earth), Sol (Sun), Astra (Star).

Re: GPT-6 Astra

#950
On the ScreenSpot-Pro benchmark, all effort levels achieve roughly the same score. I wonder if that is just a limitation of the benchmark, or if the effort levels actually do not make a difference for purely visual tasks.

> ScreenSpot-Pro tests whether models can locate the correct interface element in high-resolution screenshots of professional software.

Post reply on HN