Earlier quoted context omitted.
I've been waiting since November for 1, just 1*, model other than Claude than can reliably do agentic tool call loops. As long as the Chinese open models are chasing reasoning and benchmark maxxing vs. mid-2024 US private models, I'm very comfortable with somewhat ignoring these models. (this isn't idle prognostication hinging on my personal hobby horse. I got skin in the game, I'm virtually certain I have the only A…
You mean like https://manusai.ai/ is supposed to function?
Qwen2.5-VL-32B: Smarter and Lighter
291–300 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#292Earlier quoted context omitted.
Right: I could give you a recipe that tells you to first create a Python virtual environment, then install mlx-vlm, then make sure to downgrade to numpy 1.0 because some of the underlying libraries don't work with numpy 2.0 yet... ... or I can give you a one-liner that does all of that with uv.
python-specific side question -- is there some indication in the python ecosystems that Numpy 2x is not getting adoption? numpy-1.26 looks like 'stable' from here
Re: Qwen2.5-VL-32B: Smarter and Lighter
#293Earlier quoted context omitted.
The average user won't self-host a model.
...yet
If I try other models, I basically end up with a very bad version of AI. Even if I'm someone who uses Anthropic APIs a lot, it's absolutely not worth it to try and self host it. The APIs are much better and you get much cheaper results.
Self hosting for AI might be useful for 0.001% of people honestly.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#294Earlier quoted context omitted.
My reasoning for proprietary AI to be associated with communism is that they aren't competing in a free market way where everyone does one thing and do its best. They are simultaneously trying to do all things internally. For example , Chatgpt etc. self hosts them on their own gpu and they can generate 10tk/s or something. Now there exists groq , cerebras who can do token generation of 4000 tk/s but they kind of requ…
It seems to me like they are acting like true capitalists; they seem very happy with the idea that capital (rather than labor) gives them the right to profit. But, they don’t seem to be too attached to free-market-ism.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#295Earlier quoted context omitted.
China has allowed quite a bit of market liberalism, so it isn’t that surprising if their AI stuff is responding to the market. But, I don’t really see the connection on the flip side. Why should proprietary AI be associated with communism? If anything I guess a communist handling of AI would also be to share the model.
My reasoning for proprietary AI to be associated with communism is that they aren't competing in a free market way where everyone does one thing and do its best. They are simultaneously trying to do all things internally. For example , Chatgpt etc. self hosts them on their own gpu and they can generate 10tk/s or something. Now there exists groq , cerebras who can do token generation of 4000 tk/s but they kind of requ…
That seems based on a very weird idea of what capitalism and communism are; idealized free markets have very little to do with the real-world economic system for which the name “capitalism” was coined, and dis-integration where “everyone does one thing” has little to do with either capitalism or free markets, though it might be a convenient assumption for 101-level discussions of market competition where you want to avoid dealing with real-world issues like partially-overlapping markets and imperfect substitutes to assume every good exists in an isolated market of goods which compete only and exactly with the other groups in that same market in a simple way.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#296Earlier quoted context omitted.
Right: I could give you a recipe that tells you to first create a Python virtual environment, then install mlx-vlm, then make sure to downgrade to numpy 1.0 because some of the underlying libraries don't work with numpy 2.0 yet... ... or I can give you a one-liner that does all of that with uv.
python-specific side question -- is there some indication in the python ecosystems that Numpy 2x is not getting adoption? numpy-1.26 looks like 'stable' from here
Similar thing happened when Pydantic upgraded from 1 to 2.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#297Earlier quoted context omitted.
> Right now DeepSeek models hosted in china are having very high latency. If you are talking about DeepSeek's own hosted API service. It's because they deliberately decided to run the service in heavily overloaded conditions and have very aggressive batching policy to extract more out of their (limited) H800s. Yes, for some reason (the reason I heard is "our boss don't want to run such a business" which sounds absurd…
> the reason I heard is "our boss don't want to run such a business" which sounds absurd Liang gave up the No.1 Chinese hedge fund position to create AGI, he has very good chance to short the entire US share market and pocket some stupid amount of $ when R2 is released, he has pretty much unlimited support from local and central Chinese government. Trying to make some pennies from hosting models is not going to susta…
Re: Qwen2.5-VL-32B: Smarter and Lighter
#298Earlier quoted context omitted.
For coding and other integrations people pay per token on api key, not subscription. Claude code costs few $ per task on your code - it gets expensive quite quickly.
But something comparable to a local hosted model in the 32-70b range costs pennies on the dollar compared to Claude, will be 50x faster than your gpu, and with a much larger context window. Local hosting on GPU only really makes sense if you're doing many hours of training/inference daily.
Also "many hours of inference daily" may mean you're doing your usual stuff daily while running some processing in the background that takes hours/days or you've put together some reactive automation that runs often all the time.
ps. local training rarely makes sense.
ps. 2. not sure where you got 50x slower from; 4090 is actually faster than A100 for example and 5090 is ~75% faster than 4090
Re: Qwen2.5-VL-32B: Smarter and Lighter
#299Earlier quoted context omitted.
You mean like https://manusai.ai/ is supposed to function?
Yes, exactly, and no trivially: Manus is Sonnet with tools
Re: Qwen2.5-VL-32B: Smarter and Lighter
#300Earlier quoted context omitted.
There are lots of "alliterated" versions of models too, which is where people will essentially remove the models ability to reject responding to a prompt. The huihui r1 14b alliterated had some trouble telling me about tiananmen square, basically dodging the question by telling me about itself, but after some coaxing I was able to get the info out of it. I say this because I think that the Perplexity model is tuned o…
Abliterated? Alliterated LLMs might be fun though…