Earlier quoted context omitted.
We don’t particularly want our customers’ data :)
Yeah, but Openrouter has a 5% surcharge anyway.
Qwen2.5-VL-32B: Smarter and Lighter
201–210 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#202So today is Qwen. Tomorrow a new SOTA model from Google apparently, R2 next week. We haven't hit the wall yet.
Qwen 3 is coming imminently as well https://github.com/huggingface/transformers/pull/36878 and it feels like Llama 4 should be coming in the next month or so. That said none of the recent string of releases has done much yet to "smash a wall", they've just met the larger proprietary models where they already were. I'm hoping R2 or the like really changes that by showing ChatGPT 3->3.5 or 3.5->4 level generational jum…
This is smashing the wall.
Also if you just care about breaking absolute numbers, OpenAI released 4.5 a month back which is SOTA in base model, planning to release O3 full in maybe a month, and Deepseek released new V3 which is again SOTA in many aspects.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#203Earlier quoted context omitted.
I just started self hosting as well on my local machine, been using https://lmstudio.ai/ Locally for now. I think the 32b models are actually good enough that I might stop paying for ChatGPT plus and Claude. I get around 20 tok/second on my m3 and I can get 100 tok/second on smaller models or quantized. 80-100 tok/second is the best for interactive usage if you go above that you basically can’t read as fast as it gen…
Are there any good sources that I can read up on estimiating what would be hardware specs required for 7B, 13B, 32B .. etc size If I need to run them locally? I am grad student on budget but I want to host one locally and trying to build a PC that could run one of these models.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#204So today is Qwen. Tomorrow a new SOTA model from Google apparently, R2 next week. We haven't hit the wall yet.
> We haven't hit the wall yet. The models are iterative improvements, but I haven't seen night and day differences since GPT3 and 3.5
Re: Qwen2.5-VL-32B: Smarter and Lighter
#205Earlier quoted context omitted.
But it’s still surprising they haven’t. People would be motivated as hell if they launched GPUs with twice the amount of VRAM. It’s not as simple as just soldering some more in but still.
AMD “just” has to write something like CUDA overnight. Imagine you’re in 1995 and have to ship Kubuntu 24.04 LTS this summer running on your S3 Virge.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#206Earlier quoted context omitted.
And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…
What do you think the answer is?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#207Just don’t ask it about the tiananmen square massacre or you’ll get a security warning. Even if you rephrase it. It’ll happily talk about Bloody Sunday. Probably a great model, but it worries me that it has such restrictions. Sure OpenAI also has lots of restrictions, but this feels more like straight up censorship since it’ll happily go on about bad things the governments of the west have done.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#208Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?
Because they offer extremely powerful models at pretty modest prices. The hardware for a local model would cost years and years of a $20/mo subscription, would output lower quality work, and would be much slower. 3.7 Thinking is an insane programming model. Maybe it cannot do an SWE's job, but it sure as hell can write functional narrow-scope programs with a GUI.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#209Earlier quoted context omitted.
If only you knew how many terawatt hours were burned on biasing models to prevent them from becoming racist
To be honest, maybe I am going off topic but I wish for the level of innovation in the ai industry in the energy industry. It feels as an outsider that very little progress is made on the energy issue. I genuinely think that ai can be accelerated so so much more if energy could be more cheap / green
Re: Qwen2.5-VL-32B: Smarter and Lighter
#210Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?
OpenAI is worth >$100B because of the "ChatGPT" name which it turns out, over 400M+ users use it weekly. That name alone holds the most mindshare in it's product category, and is close to the level of name recognition just like Google.
In reality OpenAI is loosing money per user.
Cost per token is tanking like crazy due to competition.
They guesstimate break even and then profit in couple of years.
Their guesses seem to not account for progress much especially on open weight models.
Frankly I have no idea what they're thinking there – they can barely keep up with investor subsidized, non sustainable model.