Live data from Hacker News

Open models by OpenAI

openai.com

721–730 of 909 posts

Re: Open models by OpenAI

#721
post #616
post #535

Earlier quoted context omitted.

Pretty much all the large players in healthcare (provider and payer) have model access (OpenAI, Gemini, Anthropic)

That access is over a limited API and usually under heavy restrictions on the healthcare org side (e. g., only use a dedicated machine, locked up software, tracked responses and so on). Running a local model is often much easier: if you already have data on a machine and can run a model without breaching any network one could run it without any new approvals.

What? It’s a straight connect to the models api from azure, aws, or gcp.

I am literally using Claude opus 4.1 right now.

Re: Open models by OpenAI

#722
post #338

Just posted my initial impressions, took a couple of hours to write them up because there's a lot in this release! https://simonwillison.net/2025/Aug/5/gpt-oss/ TLDR: I think OpenAI may have taken the medal for best available open weight model back from the Chinese AI labs. Will be interesting to see if independent benchmarks resolve in that direction as well. The 20B model runs on my Mac laptop using less than 15GB…

Hasn't nailed the strawberry test yet

I am starting to get the impression the strawberry test is an OpenAI watermark, more than an actual problem.

It is a good way to detect if another model was trained on your data for example, or is a distillation/quant/ablation.

Re: Open models by OpenAI

#723
post #494

Earlier quoted context omitted.

Why does it need knowledge when it can just call tools to get it?

Why do I need "AI" when I can just (theoretically, in good old times Google) Google it?

Because now the model can do it for you and you can focus on other more sophisticated tasks.

I am aware that there’s a huge group of people who justify their salary by being able google.

Re: Open models by OpenAI

#724
Tried an English to Greek translation with the smaller one. Results were hideous. Mistral small is leaps and bounds better. Also I don't get why the 4-bit quantization by default. In my experience anything below 8-bit and the model fails to understand long prompts. They gutted their own models.

Re: Open models by OpenAI

#725
post #567

Here's a pair of quick sanity check questions I've been asking LLMs: "家系ラーメンについて教えて", "カレーの作り方教えて". It's a silly test but surprisingly many fails at it - and Chinese models are especially bad with it. The commonalities between models doing okay-ish for these questions seem to be Google-made OR >70b OR straight up commercial(so >200B or whatever). I'd say gpt-oss-20b is in between Qwen3 30B-A3B-2507 and Gemma 3n E4b(w…

What does failing those two questions look like? I don't really know Japanese, so I'm not sure whether I'm missing any nuances in the responses I'm getting...

The free-beer commercial ChatGPT or Gemini can read them and point out major errors. Larger Gemma models and huge Chinese models like full DeepSeek or Kimi K2 may work too. Sometimes the answer is odd enough that some 7B models can notice it. Technically there are no guarantee that models with same name in different sizes like Qwen 3 0.6B and 27B uses the same dataset, but it kind of tells a bit about quality and compositions of dataset that their creator owns.

I don't actually need accurate answers to those questions, it's just an expectation adjuster for me, so to speak. There should be better questions for other languages/use cases, but these seem to correlate better with model sizes and scales of companies than flappy birds.

0: https://gist.github.com/numpad0/abdf0a12ad73ada3b886d2d2edcc...

1: https://gist.github.com/numpad0/b1c37d15bb1b19809468c933faef...

Re: Open models by OpenAI

#726
post #699

Earlier quoted context omitted.

The 20b solved the wolf, goat, cabbage river crossing puzzle set to high reasoning for me without needing to use a system prompt that encourages critical thinking. It managed it using multiple different recommended settings, from temperatures of 0.6 up to 1.0, etc. Other models have generally failed that without a system prompt that encourages rigorous thinking. Each of the reasoning settings may very well have think…

But was it reasoning or did it solve this because it was parting it‘s training data?

Allow me to answer with a rhetorical question:

S8O2bm5lbiBTaWUgZGllc2VuIFNhdHogbGVzZW4sIGRhIGVyIGluIEJhc2UtNjQta29kaWVydGVtIERldXRzY2ggdm9ybGllZ3Q/IEhhYmVuIFNpZSBkaWUgQW50d29ydCB2b24gR3J1bmQgYXVmIGVyc2NobG9zc2VuIG9kZXIgaGFiZW4gU2llIG51ciBCYXNlIDY0IGVya2FubnQgdW5kIGRhcyBFcmdlYm5pcyBkYW5uIGluIEdvb2dsZSBUcmFuc2xhdGUgZWluZ2VnZWJlbj8gV2FzIGlzdCDDvGJlcmhhdXB0IOKAnnJlYXNvbmluZ+KAnCwgd2VubiBtYW4gbmljaHQgZGFzIEdlbGVybnRlIGF1cyBlaW5lbSBGYWxsIGF1ZiBlaW5lbiBhbmRlcmVuIGFud2VuZGV0Pw==

And yes, that's a question. Well, three, but still.

Re: Open models by OpenAI

#727

Earlier quoted context omitted.

It's the political bias in the training material. No surprise there.

More likely is that there's a lot of source material having to very stridently assert that Trump didn't win in 2020, and it's generalising to a later year. That's not political bias.

It's also extremely weird that Trump did win in 2024.

If I'd been in a coma from Jan 1 2024 to today, and woke up to people saying Trump was president again, I'd think they were pulling my leg or testing my brain function to see if I'd become gullible.

Re: Open models by OpenAI

#728
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

Now to embrace jevon's paradox and expand usage until we're back to draining lakes so that your agentic refrigerator can simulate sentience.

Why is your laptop (or phone, or refrigerator) plumbed directly into a lake?

Re: Open models by OpenAI

#729
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

I’m still trying to understand what is the biggest group of people that uses local AI (or will)? Students who don’t want to pay but somehow have the hardware? Devs who are price conscious and want free agentic coding? Local, in my experience, can’t even pull data from an image without hallucinating (Qwen 2.5 VI in that example). Hopefully local/small models keep getting better and devices get better at running bigger…

People like myself that firmly believe there will come a time, possibly very soon that all these companies (OpenAI, Anthropic etc) will raise their prices substantially. By then no one will be able to do their work to the standard expected of them without AI, and by then maybe they charge $1k per month, maybe they charge $10k. If there is no viable alternative the sky is the limit.

Why do you think they continue to run at a loss? From the goodness of their heart? Their biggest goal is to discourage anyobe from running local models. The hardware is expensive... The way to run models is very difficult (for example I have dual rtx 3090 for vram and running large heavily quantized models is a real pain in the arse, no high quantisation library supports two GPUs for example, and there seems to be no interest in implementating it by the guys behind the best inference tools).

So this is welcome, but let's not forget why it is being done.

Re: Open models by OpenAI

#730
post #699

Earlier quoted context omitted.

I tried 20b locally and it couldn't reason a way out of a basic river crossing puzzle with labels changed. That is not anywhere near SOTA. In fact it's worse than many local models that can do it, including e.g. QwQ-32b.

The 20b solved the wolf, goat, cabbage river crossing puzzle set to high reasoning for me without needing to use a system prompt that encourages critical thinking. It managed it using multiple different recommended settings, from temperatures of 0.6 up to 1.0, etc. Other models have generally failed that without a system prompt that encourages rigorous thinking. Each of the reasoning settings may very well have think…

Try changing the names of the objects. eg fox, hen, seeds for examples
Post reply on HN