Earlier quoted context omitted.
A current major outstanding problem with thinking models is how to get them to think an appropriate amount.
The providers disagree. You pay per token. Verbacious models are the most profitable. Have fun!
Open models by OpenAI
561–570 of 909 posts
Re: Open models by OpenAI
#562Re: Open models by OpenAI
#563Re: Open models by OpenAI
#564The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…
Re: Open models by OpenAI
#565OpenAI/Claude are censored in China without a VPN.
Re: Open models by OpenAI
#566Model cards, for the people interested in the guts: https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7... In my mind, I’m comparing the model architecture they describe to what the leading open-weights models (Deepseek, Qwen, GLM, Kimi) have been doing. Honestly, it just seems “ok” at a technical level: - both models use standard Grouped-Query Attention (64 query heads, 8 KV heads). The card talks about how…
I don't know how to ask this without being direct and dumb: Where do I get a layman's introduction to LLMs that could work me up to understanding every term and concept you just discussed? Either specific videos, or if nothing else, a reliable Youtube channel?
Re: Open models by OpenAI
#567Here's a pair of quick sanity check questions I've been asking LLMs: "家系ラーメンについて教えて", "カレーの作り方教えて". It's a silly test but surprisingly many fails at it - and Chinese models are especially bad with it. The commonalities between models doing okay-ish for these questions seem to be Google-made OR >70b OR straight up commercial(so >200B or whatever). I'd say gpt-oss-20b is in between Qwen3 30B-A3B-2507 and Gemma 3n E4b(w…
I don't really know Japanese, so I'm not sure whether I'm missing any nuances in the responses I'm getting...
Re: Open models by OpenAI
#568Earlier quoted context omitted.
Estimated 1.5 billion vehicles in use across the world. Generous assumptions: a) they're all IC engines requiring 16 liters of water each. b) they are changing that water out once a year That gives 24m cubic meters annual water usage. Estimated ai usage in 2024: 560m cubic meters. Projected water usage from AI in 2027: 4bn cubic meters at the low end.
what does water usage mean? is that 4bn cubic meters of water permanently out of circulation somehow? is the water corrupted with chemicals or destroyed or displaced into the atmosphere to become rain?
Re: Open models by OpenAI
#569Earlier quoted context omitted.
if you're going to get that kind of hardware, you need a larger case. IMHO this is not an unreasonable thing if you are doing heavy computing
Noted for my next build - I am aware this is a problem I've made for myself, otherwise I like the mini-ITX form factor a lot
Re: Open models by OpenAI
#570Earlier quoted context omitted.
A current major outstanding problem with thinking models is how to get them to think an appropriate amount.
The providers disagree. You pay per token. Verbacious models are the most profitable. Have fun!