Live data from Hacker News

GPT-6 Astra

openai.com

871–880 of 1001 posts

Re: GPT-6 Astra

#871

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

> why do so many of these demos include people buying things autonomously? > Even if I did trust an AI to get everything right, it's not like the AI can read my mind.

On one hand I'm with you. On the other I thought making ecommerce purchases on your phone is absolutely idiotic idea that will never catch on.

Re: GPT-6 Astra

#872

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Define novel intelligence in a way that would not exclude 95% of humans, yourself included.

Could I hook up a SOTA model the 2D computer puzzle game Gruntz (1999) so that it can read it from screenshots and act on it through keyboard and mouse inputs in a way where it would learn how to play and progress through the game? I don't think so. I doubt we'd see any sign of progress in building an internal model of how the game works and the win states in its "thinking" tokens.

Anyone that can read English could do that though.

Re: GPT-6 Astra

#873

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

> Somewhat analogous to overfitting at scale.

Sounds like entirety of human education.

Re: GPT-6 Astra

#874
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.

This Kool-Aid supply never runs out.

Re: GPT-6 Astra

#875

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

This is spot on. A collaborator is exactly what real AGI is. It will figure out the perfect questions to ask, in the perfect order, by intelligently assessing the entire solution and problem space upfront, so when you leave it to go off on its own it isn't making stupid decisions for you. They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.

> They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.

I think this already exists in Codex? If you use "/plan" and something is unclear or ambiguous, Codex will ask you and present choices, and let you enter your own custom answer. Then it'll iterate like this until the plan is clear and ambiguous. Isn't this what you're talking about? If so, it has existed for a long time in Codex.

Overall I agree with you though, all the models currently don't have the right hunches nor the right approach about when things are clear enough or not.

Re: GPT-6 Astra

#876
I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right direction, not the skills of the developer (with many exceptions of course).

Regardless, working on the wrong things is time wasted. And again, I'm procrastinating here while waiting for Fable to run a benchmark on a few solutions to a problem I have. We can guess what would work, but we only know after the benchmark. A faster model, with fewer capabilities, would've been a much better choice this time... well, "git gud" they said... and live and learn! Faster model = less time for procrastination.

PS. AI models don't live and learn; the discussion about AGI is pretty pointless imo. It's a tool. Does it matter if it is AGI or not if it does what you want it to do? Does the IQ of your colleague matter if he's good at what he's supposed to do? Or bad? Well... I guess it does matter, as many people are up in arms about whether Astro is AGI or not. Personally, I think we're past the point for that debate. These are amazing tools.

Re: GPT-6 Astra

#877

That hero video is interesting. A projector and speech. Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now. The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for m…

>this could bring us closer to the dream of more natural, social computing

What I saw was multiple people living alone in a small box in a warehouse (probably filled with other boxes) with all of their natural, social interactions directed at a wall. I wonder if this is foreshadowing for the future of work, at least it is what work will look like as envisioned by OpenAI.

Re: GPT-6 Astra

#878

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Define novel intelligence in a way that would not exclude 95% of humans, yourself included.

It gets kind of out there, but what i often hear peopel refer to is that frontier models lacks the visdom component. Which I guess is in the realm of intuition, i.e. i have a feeling it might be a problem with X based on some vague signs, maybe something a colleague mentioned offhand, something that was out of alignment etc.

Re: GPT-6 Astra

#879

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

Define novel intelligence in a way that would not exclude 95% of humans, yourself included.

I remember vaguely from a presentation by Yann LeCun "Intelligence is not what you know, but what you do when you don´t know". I find it helpful when trying to build an intuition for how to understand the LLM tool.

Re: GPT-6 Astra

#880

“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

Have you seen ads for mobile games, where there are seemingly 100 different versions of the same type of game (like tower defense types)? And they're all obviously the worst type of pay-to-play traps? I think this kind of solves that. Or at least it is the start of it. Most such games are kind of trivial. If people can easily just get AI to generate such games on the fly, then that'll hopefully be the end of predator…

The problem is, no one wants to play a game they designed themselves
Post reply on HN