Earlier quoted context omitted.
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?
Since we are on HN here, I can highly recommend open-webui with some OpenAI-compatible provider. I'm running with Deep Infra for more than a year now and am very happy. New models are usually available within one or two days after release. Also have some friends who use the service almost daily.
Qwen2.5-VL-32B: Smarter and Lighter
131–140 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#132Earlier quoted context omitted.
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?
Since we are on HN here, I can highly recommend open-webui with some OpenAI-compatible provider. I'm running with Deep Infra for more than a year now and am very happy. New models are usually available within one or two days after release. Also have some friends who use the service almost daily.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#133Just don’t ask it about the tiananmen square massacre or you’ll get a security warning. Even if you rephrase it. It’ll happily talk about Bloody Sunday. Probably a great model, but it worries me that it has such restrictions. Sure OpenAI also has lots of restrictions, but this feels more like straight up censorship since it’ll happily go on about bad things the governments of the west have done.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#134Earlier quoted context omitted.
Is OpenRouter planning on distilling models off the prompts and responses from frontier models? That's smart - a little gross - but smart.
COO of OpenRouter here. We are simply stating the WE can’t vouch for the behavior of the upstream provider’s retention and training policy. We don’t save your prompt data, regardless of the model you use, unless you explicitly opt-in to logging (in exchange for a 1% inference discount).
Re: Qwen2.5-VL-32B: Smarter and Lighter
#135Earlier quoted context omitted.
Since we are on HN here, I can highly recommend open-webui with some OpenAI-compatible provider. I'm running with Deep Infra for more than a year now and am very happy. New models are usually available within one or two days after release. Also have some friends who use the service almost daily.
And it’s quite easy to set up a Cloudflare tunnel to make your open-webui instance accessible online too just you
Re: Qwen2.5-VL-32B: Smarter and Lighter
#136Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#137Earlier quoted context omitted.
https://huggingface.co/spaces/NyxKrage/LLM-Model-VRAM-Calcul... That will help you quickly calculate the model VRAM usage as well as the VRAM usage of the context length you want to use. You can put "Qwen/Qwen2.5-VL-32B-Instruct" in the "Model (unquantized)" field. Funnily enough the calculator lacks the option to see without quantizing the model, usually because nobody worried about VRAM bothers running >8 bit quant…
Except when it comes to deepseek
I expect more and more worthwhile models to natively have <16 bit weights as time goes on but for the moment it's pretty much "8 bit DeepSeek and some research/testing models of various parameter width".
Re: Qwen2.5-VL-32B: Smarter and Lighter
#138Earlier quoted context omitted.
is there some reason you cant train a 1b model to just do agentic stuff?
They're real squished for space, more than I expected :/ good illustration here, Qwen2.5-1.5B trained to reason, i.e. the name it is released under is "DeepSeek R1 1.5B". https://imgur.com/a/F3w5ymp 1st prompt was "What is 1048576^0.05", it answered, then I said "Hi", then...well... Fwiw, Claude Sonnet 3.5 100% had some sort of agentic loop x precise file editing trained into it. Wasn't obvious to me until I added a…
https://chatgpt.com/share/67e1eba1-c658-800e-9161-a0b8b7b683...
Re: Qwen2.5-VL-32B: Smarter and Lighter
#139Earlier quoted context omitted.
COO of OpenRouter here. We are simply stating the WE can’t vouch for the behavior of the upstream provider’s retention and training policy. We don’t save your prompt data, regardless of the model you use, unless you explicitly opt-in to logging (in exchange for a 1% inference discount).
That 1% discount feels a bit cheap to me - if it was a 25% or 50% discount I would be much more likely to sign up for it.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#140Earlier quoted context omitted.
Is OpenRouter planning on distilling models off the prompts and responses from frontier models? That's smart - a little gross - but smart.
COO of OpenRouter here. We are simply stating the WE can’t vouch for the behavior of the upstream provider’s retention and training policy. We don’t save your prompt data, regardless of the model you use, unless you explicitly opt-in to logging (in exchange for a 1% inference discount).