Live data from Hacker News

Open models by OpenAI

openai.com

651–660 of 909 posts

Re: Open models by OpenAI

#651

Earlier quoted context omitted.

I tried 20b locally and it couldn't reason a way out of a basic river crossing puzzle with labels changed. That is not anywhere near SOTA. In fact it's worse than many local models that can do it, including e.g. QwQ-32b.

Well river crossings are one type of problem. My real world problem is proofing and minor editing of text. A version installed on my portable would be great.

I heard the OSSmodels are terrible at anything other than math, code etc.

Re: Open models by OpenAI

#652

Earlier quoted context omitted.

Add healthcare. Cannot send our patients data to a cloud provider

A ton of EMR systems are cloud-hosted these days. There’s already patient data for probably a billion humans in the various hyperscalers. Totally understand that approaches vary but beyond EMR there’s work to augment radiologists with computer vision to better diagnose, all sorts of cloudy things. It’s here. It’s growing. Perhaps in your jurisdiction it’s prohibited? If so I wonder for how long.

Even if it's possible, there is typically a lot of paperwork to get that stuff approved.

There might be a lot less paperwork to just buy 50 decent GPU's and have the IT guy self-host.

Re: Open models by OpenAI

#653
post #316

The 120B model badly hallucinates facts on the level of a 0.6B model. My go to test for checking hallucinations is 'Tell me about Mercantour park' (a national park in south eastern France). Easily half of the facts are invented. Non-existing mountain summits, brown bears (no, there are none), villages that are elsewhere, wrong advice ('dogs allowed' - no they are not).

Others have already said it, but it needs to be said again: Good god, stop treating LLMs like oracles. LLMs are not encyclopedias. Give an LLM the context you want to explore, and it will do a fantastic job of telling you all about it. Give an LLM access to web search, and it will find things for you and tell you what you want to know. Ask it "what's happening in my town this week?", and it will answer that with the…

To be coherent and useful in general-purpose scenarios, LLM absolutely has to be large enough and know a lot, even if you aren't using is as an oracle.

Re: Open models by OpenAI

#654
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

The model is good and runs fine but if you want to be blown away again try Qwen3-30A-A3B-2507. It's 6gb bigger but the response is comparable or better and much faster to run. Gpt-oss-20B gives me 6 tok/sec while Qwen3 gives me 37 tok/sec. Qwen3 is not a reasoning model tho.

Re: Open models by OpenAI

#655

Earlier quoted context omitted.

I wouldn’t stop at 16GB right now. 24 is the lowest I would go. Buy a used 3090. Picked one up for $700 a few months back, but I think they were on the rise then. The 3000 series can’t do FP8fast, but meh. It’s the OOM that’s tough, not the speed so much.

Are there any 24GB cards/3090s which fit in ~300mm without an angle grinder?

https://skinflint.co.uk/?cat=gra16_512&hloc=uk&v=e&hloc=at&h...

5070 Ti Super will also have 24GB.

Re: Open models by OpenAI

#656

> To improve the safety of the model, we filtered the data for harmful content in pre-training, especially around hazardous biosecurity knowledge, by reusing the CBRN pre-training filters from GPT-4o. Our model has a knowledge cutoff of June 2024. This would be a great "AGI" test. See if it can derive biohazards from first principles

Not possible without running real-life experiments, unless they still memorized it somehow.

Re: Open models by OpenAI

#657
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

I’m still trying to understand what is the biggest group of people that uses local AI (or will)? Students who don’t want to pay but somehow have the hardware? Devs who are price conscious and want free agentic coding? Local, in my experience, can’t even pull data from an image without hallucinating (Qwen 2.5 VI in that example). Hopefully local/small models keep getting better and devices get better at running bigger…

Maybe I am too pessimistic, but as an EU citizen I expect politics (or should I say Trump?) to prevent access to US-based frontier models at some point.

Re: Open models by OpenAI

#658
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

Where did you get the top ten from? https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro Are you discounting all of the self reported scores?

Came here to say this. It's behind the 14b Phi-reasoning-plus (which is self-reported).

I don't understand why "TIGER-LAb"-sourced scores are 'unknown' in terms of model size?

Re: Open models by OpenAI

#659

Earlier quoted context omitted.

Wondering about the same for my M4 max 128 gb

It should fly on your machine

Yeah, was super quick and easy to set up using Ollama. I had to kill some processes first to avoid memory swap though (even with 128gb memory). So a slightly more quantized version is maybe ideal, for me at least.

Edit: I'm talking about the 120B model of course

Re: Open models by OpenAI

#660
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

The environmentalist in me loves the fact that LLM progress has mostly been focused on doing more with the same hardware, rather than horizontal scaling. I guess given GPU shortages that makes sense, but it really does feel like the value of my hardware (a laptop in my case) is going up over time, not down. Also, just wanted to credit you for being one of the five people on Earth who knows the correct spelling of "le…

> Also, just wanted to credit you for being one of the five people on Earth who knows the correct spelling of "lede."

Not in the UK it isn’t.

Post reply on HN