Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

661–670 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#661

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

compute gets cheaper and cheaper every year. This model will be in your phone by 2030 if we continue at the pace we've been at the last few years.

These models are nearing 2+ trillion parameters. At 4 bits each, we're talking about somewhere around 1tb of RAM.

The problem is that RAM stopped scaling a long time ago now. We're down to the size where a single capacitor's charge is held by a mere 40,000 or so electrons and all we've been doing is making skinnier, longer cells of that size because we can't find reliable ways to boost even weaker signals, but this is a dead end because as the math shows, if the volume is consistent and you are reducing X and Y dimensions, that Z dimension starts to get crazy big really fast. The chemistry issues of burning a hole a little at a time while keeping wall thickness somewhat similar all the way down is a very hard problem.

Another problem is that Moore's law hit a wall when Dennard Scaling failed. When you look at SRAM (it's generally the smallest and most reliable stuff we can make), you see that most recent shrinks can hardly be called shrinks.

Unless we do something very different like compute in storage or have some radical breakthrough in a new technology, I don't know that we will ever get a 2T parameter model inside a phone (I'd love for someone in 10 years to show up and say how wrong I was).

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#662

Earlier quoted context omitted.

Have we really watered down the definition of AGI that much? LLMs aren't really capable of "learning" anything outside their training data. Which I feel is a very basic and fundamental capability of humans. Every new request thread is a blank slate utilizing whatever context you provide for the specific task and after the tread is done (or context limit runs out) it's like it never happened. Sure you can use database…

> LLMs aren't really capable of "learning" anything outside their training data. ChatGPT has had for some time the feature of storing memories about its conversations with users. And you can use function calling to make this more generic. I think drawing the boundary at “model + scaffolding” is more interesting.

Calling the sentence or two it arbitrarily saves when you statd your preferences and profile info "memories" is a stretch.

True equivalent to human memories would require something like a multimodal trillion token context window.

RAG is just not going to cut it, and if anything will exacerbated problems with hallucinations.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#663
post #537

I just graduated college, and this was a major blow. I studied Mechanical Engineering and went into Sales Engineering because cause I love technology and people, but articles like this do nothing but make me dread the future. I have no idea what to specialize in, what skills I should master, or where I should be spending my time to build a successful career. Seems like we’re headed toward a world where you automate s…

You're actually positioned to have an amazing career.

Everyone needs to know how to either build or sell to be successful. In a world where the ability to the former is rapidly being commoditised, you will still need to sell. And human relationships matter more than ever.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#664
post #537

I just graduated college, and this was a major blow. I studied Mechanical Engineering and went into Sales Engineering because cause I love technology and people, but articles like this do nothing but make me dread the future. I have no idea what to specialize in, what skills I should master, or where I should be spending my time to build a successful career. Seems like we’re headed toward a world where you automate s…

LLMs are mostly hype. They're not going to change things that much.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#665
post #427

How can there be "private" taks when you have use the OpenAI API to run queries? OpenAI sees everything.

We worked with ARC to run inference on the semi-private tasks last week, after o3 was trained, using an inference only API that was sent the prompts but not the answers & did no durable logging.

What's your opinion on the veracity of this benchmark - given o3 was fine-tuned and others were not? Can you give more details on how much data was used to fine-tune o3? It's hard to put this into perspective given this confounder.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#667
post #537

I just graduated college, and this was a major blow. I studied Mechanical Engineering and went into Sales Engineering because cause I love technology and people, but articles like this do nothing but make me dread the future. I have no idea what to specialize in, what skills I should master, or where I should be spending my time to build a successful career. Seems like we’re headed toward a world where you automate s…

It's a tool. You learn to master it or not. I have greybeard coworkers that dissed the technology as a fad 3 years ago. Now they are scrambling to catch up. They have to do this while sustaining a family with pets and kids and mortgages and full time senior jobs. You're in a position to invest substantial amounts of time compared to your seniors. Leverage that opportunity to your advantage. We all have access to thes…

> I have greybeard coworkers that dissed the technology as a fad 3 years ago. Now they are scrambling to catch up. They have to do this while sustaining a family with pets and kids and mortgages and full time senior jobs.

I want to criticize Art’s comment on the grounds of ageism or something along the lines of “any amount life outside of programming is wasted”, but regardless of Art’s intention there is important wisdom here. Use your free time wisely when you don’t have much responsibilities. It is a superpower.

As for whether to spend it on AI, eh, that’s up to you to decide.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#668

Sad to see everyone so focused on compute expense during this massive breakthrough. GPT-2 originally cost $50k to train, but now can be trained for ~$150. The key part is that scaling test-time compute will likely be a key to achieving AGI/ASI. Costs will definitely come down as is evidenced by precedents, Moore’s law, o3-mini being cheaper than o1 with improved performance, etc.

[deleted]

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#669
post #537

I just graduated college, and this was a major blow. I studied Mechanical Engineering and went into Sales Engineering because cause I love technology and people, but articles like this do nothing but make me dread the future. I have no idea what to specialize in, what skills I should master, or where I should be spending my time to build a successful career. Seems like we’re headed toward a world where you automate s…

Don’t worry. This thing only knows how to answer well structured technical questions.

99% of engineering is distilling through bullshit and nonsense requirements. Whether that is appealing to you is a different story, but ChatGPT will happily design things with dumb constraints that would get you fired if you took them at face value as an engineer.

ChatGPT answering technical challenges is to engineering as a nailgun is to carpentry.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#670

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

o3-mini (high) uses 1/3rd of the compute of o1, and performs about 200 Elo higher than o1 on Codeforces.

o1 is the best code generation model according to Livebench.

So how is this not a breakthrough? It's a genuine movement of the frontier.

Post reply on HN