Well that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.
Deepseek R1-0528
111–120 of 264 posts
Re: Deepseek R1-0528
#112Earlier quoted context omitted.
Ok https://huggingface.co/deepseek-ai/DeepSeek-R1-0528/blob/mai...
Slapping an MIT license on a compiled binary doesn't make it open source.
What they have released has been distilled into many new models that others have been using for commercial benefit and I appreciate the contributions that they have made.
Re: Deepseek R1-0528
#113Out of sheer curiosity: What’s required for the average Joe to use this, even at a glacial pace, in terms of hardware? Or is it even possible without using smart person magic to append enchanted numbers and make it smaller for us masses?
About 768 gigs of ddr5 RAM in a dual socket server board with 12 channel memory and an extra 16 gig or better GPU for prompt processing. It's a few grand just to run this thing at 8-10 tokens/s
Mobo was some kind of mining rig from AliExpress for less than $100. GPU is an inexpensive NVIDIA TESLA card that I 3D printed a shroud for (added fans). Power supply a cheap 2000 Watt Dell server PS off eBay....
[1] https://bsky.app/profile/engineersneedart.com/post/3lmg4kiz4...
Re: Deepseek R1-0528
#114Well that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.
No sign of what source material it was trained on though right? So open weight rather than reproducible from source. I remember there's a project "Open R1" that last I checked was working on gathering their own list of training material, looks active but not sure how far along they've gotten: https://github.com/huggingface/open-r1
There's a few efforts at full open data / open weight / open code models, but none of them have gotten to leading-edge performance.
Re: Deepseek R1-0528
#115Earlier quoted context omitted.
How does releasing it today affect the market compared to releasing it last week?
Hard to say exactly how it will affect the market, but IIRC when deepseek was first released Nvidia stock took a big hit as people realized that you could develop high performing LLMs without access to Nvidia hardware.
Not only DeepSeek uses a lot of Nvidia hardware for the training.
But even more so, by releasing an open weight frontier model, people around the world need more Nvidia chips than ever for inference.
Re: Deepseek R1-0528
#116No information to be found about it. Hopefully we get benchmarks soon. Reminds me of the days when Mistral would just tweet a torrent magnet link
Benchmarks seem like a fools errand at this point; overly tuning models just to specific test already published tests, rather than focusing on making them generalize. Hugging face has a leader board and it seems dominated by models that are finetunings of various common open source models, yet don't seem be broader used: https://huggingface.co/open-llm-leaderboard
Re: Deepseek R1-0528
#117Well that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.
Is there a downloadable model? (Not familiar with openrouter and not seeing the model on ollama.)
Re: Deepseek R1-0528
#118This whole “building moats” and buying competitors fascination in the US has gotten boring, obvious and dull. The world benefits when companies struggle to be the best.
Re: Deepseek R1-0528
#119Re: Deepseek R1-0528
#120Earlier quoted context omitted.
Benchmarks seem like a fools errand at this point; overly tuning models just to specific test already published tests, rather than focusing on making them generalize. Hugging face has a leader board and it seems dominated by models that are finetunings of various common open source models, yet don't seem be broader used: https://huggingface.co/open-llm-leaderboard
>overly tuning models just to specific test already published tests, rather than focusing on making them generalize. I think you just described SATs and other standardized tests