Earlier quoted context omitted.
I tried 20b locally and it couldn't reason a way out of a basic river crossing puzzle with labels changed. That is not anywhere near SOTA. In fact it's worse than many local models that can do it, including e.g. QwQ-32b.
Well river crossings are one type of problem. My real world problem is proofing and minor editing of text. A version installed on my portable would be great.
Open models by OpenAI
651–660 of 909 posts
Re: Open models by OpenAI
#652Earlier quoted context omitted.
Add healthcare. Cannot send our patients data to a cloud provider
A ton of EMR systems are cloud-hosted these days. There’s already patient data for probably a billion humans in the various hyperscalers. Totally understand that approaches vary but beyond EMR there’s work to augment radiologists with computer vision to better diagnose, all sorts of cloudy things. It’s here. It’s growing. Perhaps in your jurisdiction it’s prohibited? If so I wonder for how long.
There might be a lot less paperwork to just buy 50 decent GPU's and have the IT guy self-host.
Re: Open models by OpenAI
#653The 120B model badly hallucinates facts on the level of a 0.6B model. My go to test for checking hallucinations is 'Tell me about Mercantour park' (a national park in south eastern France). Easily half of the facts are invented. Non-existing mountain summits, brown bears (no, there are none), villages that are elsewhere, wrong advice ('dogs allowed' - no they are not).
Others have already said it, but it needs to be said again: Good god, stop treating LLMs like oracles. LLMs are not encyclopedias. Give an LLM the context you want to explore, and it will do a fantastic job of telling you all about it. Give an LLM access to web search, and it will find things for you and tell you what you want to know. Ask it "what's happening in my town this week?", and it will answer that with the…
Re: Open models by OpenAI
#654The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…
Re: Open models by OpenAI
#655Earlier quoted context omitted.
I wouldn’t stop at 16GB right now. 24 is the lowest I would go. Buy a used 3090. Picked one up for $700 a few months back, but I think they were on the rise then. The 3000 series can’t do FP8fast, but meh. It’s the OOM that’s tough, not the speed so much.
Are there any 24GB cards/3090s which fit in ~300mm without an angle grinder?
5070 Ti Super will also have 24GB.
Re: Open models by OpenAI
#656> To improve the safety of the model, we filtered the data for harmful content in pre-training, especially around hazardous biosecurity knowledge, by reusing the CBRN pre-training filters from GPT-4o. Our model has a knowledge cutoff of June 2024. This would be a great "AGI" test. See if it can derive biohazards from first principles
Re: Open models by OpenAI
#657The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…
I’m still trying to understand what is the biggest group of people that uses local AI (or will)? Students who don’t want to pay but somehow have the hardware? Devs who are price conscious and want free agentic coding? Local, in my experience, can’t even pull data from an image without hallucinating (Qwen 2.5 VI in that example). Hopefully local/small models keep getting better and devices get better at running bigger…
Re: Open models by OpenAI
#658The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…
Where did you get the top ten from? https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro Are you discounting all of the self reported scores?
I don't understand why "TIGER-LAb"-sourced scores are 'unknown' in terms of model size?
Re: Open models by OpenAI
#659Earlier quoted context omitted.
Wondering about the same for my M4 max 128 gb
It should fly on your machine
Edit: I'm talking about the 120B model of course
Re: Open models by OpenAI
#660The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…
The environmentalist in me loves the fact that LLM progress has mostly been focused on doing more with the same hardware, rather than horizontal scaling. I guess given GPU shortages that makes sense, but it really does feel like the value of my hardware (a laptop in my case) is going up over time, not down. Also, just wanted to credit you for being one of the five people on Earth who knows the correct spelling of "le…
Not in the UK it isn’t.