Live data from Hacker News

Open models by OpenAI

openai.com

731–740 of 909 posts

Re: Open models by OpenAI

#731

Tried an English to Greek translation with the smaller one. Results were hideous. Mistral small is leaps and bounds better. Also I don't get why the 4-bit quantization by default. In my experience anything below 8-bit and the model fails to understand long prompts. They gutted their own models.

They used quantization-aware training, so the quality loss should be negligible. Doing anything with this model's weights would be a different story, though.

The model is clearly heavily finetuned towards coding and math, and is borderline unusable for creative writing and translation in particular. It's not general-purpose, excessively filtered (refusal training and dataset lobotomy is probably a major factor behind lower than expected performance), and shouldn't be compared with Qwen or o3 at all.

Re: Open models by OpenAI

#732
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

I’m still trying to understand what is the biggest group of people that uses local AI (or will)? Students who don’t want to pay but somehow have the hardware? Devs who are price conscious and want free agentic coding? Local, in my experience, can’t even pull data from an image without hallucinating (Qwen 2.5 VI in that example). Hopefully local/small models keep getting better and devices get better at running bigger…

At the company where I currently work, for IP reasons (and with the advice of a patent lawyer), nobody is allowed to use any online AIs to talk about or help with work, unless it's very generic research that doesn't give away what we're working on.

That rules out coding assistants like Claude, chat, tools to generate presentations and copy-edit documents, and so forth.

But local AI are fine, as long as we're sure nothing is uploaded.

Re: Open models by OpenAI

#733
post #727

Earlier quoted context omitted.

More likely is that there's a lot of source material having to very stridently assert that Trump didn't win in 2020, and it's generalising to a later year. That's not political bias.

It's also extremely weird that Trump did win in 2024. If I'd been in a coma from Jan 1 2024 to today, and woke up to people saying Trump was president again, I'd think they were pulling my leg or testing my brain function to see if I'd become gullible.

You’re in a bubble. It was no surprise to folks who touch grass on the regular.

Re: Open models by OpenAI

#734
post #338

Just posted my initial impressions, took a couple of hours to write them up because there's a lot in this release! https://simonwillison.net/2025/Aug/5/gpt-oss/ TLDR: I think OpenAI may have taken the medal for best available open weight model back from the Chinese AI labs. Will be interesting to see if independent benchmarks resolve in that direction as well. The 20B model runs on my Mac laptop using less than 15GB…

Nice write up!

One test I do is to give a common riddle but word it slightly to see if it can actually reason.

For example:

"Bobs dad has five daughters, Lala, Lele, Lili, Lolo and ???"

The 20B model kept picking the answer of the original riddle, even after explaining extra information to it.

The original riddle is:

"Janes dad has five daughters, Lala, Lele, Lili, Lolo and ???"

Re: Open models by OpenAI

#735

Earlier quoted context omitted.

A ton of EMR systems are cloud-hosted these days. There’s already patient data for probably a billion humans in the various hyperscalers. Totally understand that approaches vary but beyond EMR there’s work to augment radiologists with computer vision to better diagnose, all sorts of cloudy things. It’s here. It’s growing. Perhaps in your jurisdiction it’s prohibited? If so I wonder for how long.

In the US, HIPAA requires that health care providers complete a Business Associate Agreement with any other orgs that receive PHI in the course of doing business [1]. It basically says they understand HIPAA privacy protections and will work to fulfill the contracting provider's obligations regarding notification of breaches and deletion. Obviously any EMR service will include this by default. Most orgs charge a huge…

I worked a big health care company recently. We were using Azure's private instances of the GPT models. Fully industry compliant.

Re: Open models by OpenAI

#736
post #338

Just posted my initial impressions, took a couple of hours to write them up because there's a lot in this release! https://simonwillison.net/2025/Aug/5/gpt-oss/ TLDR: I think OpenAI may have taken the medal for best available open weight model back from the Chinese AI labs. Will be interesting to see if independent benchmarks resolve in that direction as well. The 20B model runs on my Mac laptop using less than 15GB…

Nice write up! One test I do is to give a common riddle but word it slightly to see if it can actually reason. For example: "Bobs dad has five daughters, Lala, Lele, Lili, Lolo and ???" The 20B model kept picking the answer of the original riddle, even after explaining extra information to it. The original riddle is: "Janes dad has five daughters, Lala, Lele, Lili, Lolo and ???"

I don’t get it. Wouldn’t it be Lulu in both cases?

Re: Open models by OpenAI

#737
post #727

Earlier quoted context omitted.

More likely is that there's a lot of source material having to very stridently assert that Trump didn't win in 2020, and it's generalising to a later year. That's not political bias.

It's also extremely weird that Trump did win in 2024. If I'd been in a coma from Jan 1 2024 to today, and woke up to people saying Trump was president again, I'd think they were pulling my leg or testing my brain function to see if I'd become gullible.

Unfortunately, it was predictable given the other "choices"

Re: Open models by OpenAI

#738
post #567

Earlier quoted context omitted.

What does failing those two questions look like? I don't really know Japanese, so I'm not sure whether I'm missing any nuances in the responses I'm getting...

The free-beer commercial ChatGPT or Gemini can read them and point out major errors. Larger Gemma models and huge Chinese models like full DeepSeek or Kimi K2 may work too. Sometimes the answer is odd enough that some 7B models can notice it. Technically there are no guarantee that models with same name in different sizes like Qwen 3 0.6B and 27B uses the same dataset, but it kind of tells a bit about quality and com…

Thanks for the detailed response.

I'm guessing the issue is just the model size. If you're testing sub-30B models and finding errors, well they're probably not large enough to remember everything in the training data set, so there's inaccuracies and they might hallucinate a bit regarding factoids that aren't very commonly seen in the training data.

Commercial models are presumably significantly larger than the smaller open models, so it sounds like the issue is just mainly model size...

PS: Okra on curry is pretty good actually :)

Re: Open models by OpenAI

#740
post #726

Earlier quoted context omitted.

But was it reasoning or did it solve this because it was parting it‘s training data?

Allow me to answer with a rhetorical question: S8O2bm5lbiBTaWUgZGllc2VuIFNhdHogbGVzZW4sIGRhIGVyIGluIEJhc2UtNjQta29kaWVydGVtIERldXRzY2ggdm9ybGllZ3Q/IEhhYmVuIFNpZSBkaWUgQW50d29ydCB2b24gR3J1bmQgYXVmIGVyc2NobG9zc2VuIG9kZXIgaGFiZW4gU2llIG51ciBCYXNlIDY0IGVya2FubnQgdW5kIGRhcyBFcmdlYm5pcyBkYW5uIGluIEdvb2dsZSBUcmFuc2xhdGUgZWluZ2VnZWJlbj8gV2FzIGlzdCDDvGJlcmhhdXB0IOKAnnJlYXNvbmluZ+KAnCwgd2VubiBtYW4gbmljaHQgZGFzIEdlbGVybnRlIGF1c…

In case of the river puzzle there is a huge difference between repeating an answer that you read somewhere and figuring it out on your own, one requires reasoning the other does not. If you swap out the animals involved, then you need some reasoning to recognize the identical structure of the puzzles and map between the two sets of animals. But you are still very far from the amount of reasoning required to solve the puzzle without already knowing the answer.

You can do it brute force, that requires again more reasoning than mapping between structurally identical puzzles. And finally you can solve it systematically, that requires the largest amount of reasoning. And in all those cases there is a crucial difference between blindly repeating the steps of a solution that you have seen before and coming up with that solution on your own even if you can not tell the two cases apart by looking at the output which would be identical.

Post reply on HN