Live data from Hacker News

Open models by OpenAI

openai.com

801–810 of 909 posts

Re: Open models by OpenAI

#801

Earlier quoted context omitted.

> In fact it's worse than many local models that can do it, including e.g. QwQ-32b. I'm not going to be surprised that a 20B 4/32 MoE model (3.6B parameters activated) is less capable at a particular problem category than a 32B dense model, and its quite possible for both to be SOTA, as state of the art at different scale (both parameter count and speed which scales with active resource needs) is going to have differ…

[flagged]

This isn't reddit.

Re: Open models by OpenAI

#802
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

>no lakes being drained

When you imagine a lake being drained to cool a datacenter do you ever consider where the water used for cooling goes? Do you imagine it disappears?

Re: Open models by OpenAI

#803
post #726

Earlier quoted context omitted.

But was it reasoning or did it solve this because it was parting it‘s training data?

Allow me to answer with a rhetorical question: S8O2bm5lbiBTaWUgZGllc2VuIFNhdHogbGVzZW4sIGRhIGVyIGluIEJhc2UtNjQta29kaWVydGVtIERldXRzY2ggdm9ybGllZ3Q/IEhhYmVuIFNpZSBkaWUgQW50d29ydCB2b24gR3J1bmQgYXVmIGVyc2NobG9zc2VuIG9kZXIgaGFiZW4gU2llIG51ciBCYXNlIDY0IGVya2FubnQgdW5kIGRhcyBFcmdlYm5pcyBkYW5uIGluIEdvb2dsZSBUcmFuc2xhdGUgZWluZ2VnZWJlbj8gV2FzIGlzdCDDvGJlcmhhdXB0IOKAnnJlYXNvbmluZ+KAnCwgd2VubiBtYW4gbmljaHQgZGFzIEdlbGVybnRlIGF1c…

No post body was provided.

Re: Open models by OpenAI

#804

Earlier quoted context omitted.

I'm pretty sure you are completely correct on the last part. Nobody in Republican management wanted a second Trump term. If the candidate wasn't Trump, Republicans would have had a guaranteed victory. Imagine that infamous debate, but with some 50-year-old youngster facing Joe Biden. It's the White House that wanted Trump to be candidate. They played Republican primary voters like a fiddle by launching a barrage of t…

You think the Democratic White House, manipulated Republicans into Voting for Trump. So it is the Democrats fault we have Trump??? Next Level Cope.

> You think the Democratic White House, manipulated Republicans into Voting for Trump.

Yes, that is what he thinks. Did you not read the comment? It is, like, uh, right there...

He also explained his reasoning: If Trump didn't win the party race, a more compelling option (the so-called "50-year-old youngster") would have instead, which he claims would have guaranteed a Republican win. In other words, what he is saying that the White House was banking on Trump losing the presidency.

Re: Open models by OpenAI

#806
post #726

Earlier quoted context omitted.

But was it reasoning or did it solve this because it was parting it‘s training data?

Allow me to answer with a rhetorical question: S8O2bm5lbiBTaWUgZGllc2VuIFNhdHogbGVzZW4sIGRhIGVyIGluIEJhc2UtNjQta29kaWVydGVtIERldXRzY2ggdm9ybGllZ3Q/IEhhYmVuIFNpZSBkaWUgQW50d29ydCB2b24gR3J1bmQgYXVmIGVyc2NobG9zc2VuIG9kZXIgaGFiZW4gU2llIG51ciBCYXNlIDY0IGVya2FubnQgdW5kIGRhcyBFcmdlYm5pcyBkYW5uIGluIEdvb2dsZSBUcmFuc2xhdGUgZWluZ2VnZWJlbj8gV2FzIGlzdCDDvGJlcmhhdXB0IOKAnnJlYXNvbmluZ+KAnCwgd2VubiBtYW4gbmljaHQgZGFzIEdlbGVybnRlIGF1c…

As mgoetzke challenges, change the names of the items to something different, but the same puzzle. If it fails with "fox, hen, seeds" instead of "wolf, goat, cabbage" then it wasn't reasoning or applying something learned to another case. It was just regurgitating from the training data.

Re: Open models by OpenAI

#807
post #338

Just posted my initial impressions, took a couple of hours to write them up because there's a lot in this release! https://simonwillison.net/2025/Aug/5/gpt-oss/ TLDR: I think OpenAI may have taken the medal for best available open weight model back from the Chinese AI labs. Will be interesting to see if independent benchmarks resolve in that direction as well. The 20B model runs on my Mac laptop using less than 15GB…

[dead]

Re: Open models by OpenAI

#808
post #780

Earlier quoted context omitted.

Not really the other commenters are correct I feel and this is not really proving anything about the fundamental capability of the model. It’s just a hello world benchmark adding no real value, just driving blog traffic for you.

The space invaders benchmark proves that the model can implement a working HTML and JavaScript game from a single prompt. That's a pretty fundamental capability for a model. Comparing them between models is also kind of interesting, even if it's not a flawlessly robust comparison: https://simonwillison.net/tags/space-invaders/

Implement or retrieve? That’s an important distinction. When evaluating models, you run a variety of tests, and the benchmarks that aren’t publicly disclosed are the most reliable. Your Space Invaders game isn’t really a benchmark of anything, just Google it, and you’ll find plenty of implementations.

Re: Open models by OpenAI

#809
post #726

Earlier quoted context omitted.

But was it reasoning or did it solve this because it was parting it‘s training data?

Allow me to answer with a rhetorical question: S8O2bm5lbiBTaWUgZGllc2VuIFNhdHogbGVzZW4sIGRhIGVyIGluIEJhc2UtNjQta29kaWVydGVtIERldXRzY2ggdm9ybGllZ3Q/IEhhYmVuIFNpZSBkaWUgQW50d29ydCB2b24gR3J1bmQgYXVmIGVyc2NobG9zc2VuIG9kZXIgaGFiZW4gU2llIG51ciBCYXNlIDY0IGVya2FubnQgdW5kIGRhcyBFcmdlYm5pcyBkYW5uIGluIEdvb2dsZSBUcmFuc2xhdGUgZWluZ2VnZWJlbj8gV2FzIGlzdCDDvGJlcmhhdXB0IOKAnnJlYXNvbmluZ+KAnCwgd2VubiBtYW4gbmljaHQgZGFzIEdlbGVybnRlIGF1c…

(Decoded, if anyone's wondering):

> Können Sie diesen Satz lesen, da er in Base-64-kodiertem Deutsch vorliegt? Haben Sie die Antwort von Grund auf erschlossen oder haben Sie nur Base 64 erkannt und das Ergebnis dann in Google Translate eingegeben? Was ist überhaupt „reasoning“, wenn man nicht das Gelernte aus einem Fall auf einen anderen anwendet?

>

> Can you read this sentence, since it's in Base-64 encoded German? Did you deduce the answer from scratch, or did you just recognize Base 64 and then enter the result into Google Translate? What is "reasoning" anyway if you don't apply what you've learned from one case to another?

Re: Open models by OpenAI

#810

Earlier quoted context omitted.

It’s Bob or Jane. The dad of has 5 daughters. Four are listed off. So the answer for the fifth is .

Except having five daughters doesn't prevent them also having 20 sons one of whom is called Bob.

That’s why it’s a riddle.
Post reply on HN