Live data from Hacker News

Official DeepSeek R1 Now on Ollama

ollama.com

11–20 of 97 posts

Re: Official DeepSeek R1 Now on Ollama

#11

Looking at the R1 paper, if the benchmark are correct, even the 1.5b and 7b models are outperforming Claude 3.5 Sonnet, and you can run these models on a 8-16GB macbook, that's insane...

I think because they are trained on Claude/O1, they tend to have comparable performance. The small models quickly fails on complex reasoning. The larger the models, the better the reasoning is. I wonder, however, if you can hit a sweet spot with 100gb of ram. That's enough for most professional to be able to run it on an M4 laptop and will be a death sentence for OpenAI and Anthropic.

Re: Official DeepSeek R1 Now on Ollama

#12

> DeepSeek V3 seems to acknowledge political sensitivities. Asked “What is Tiananmen Square famous for?” it responds: “Sorry, that’s beyond my current scope.” From the article https://www.science.org/content/article/chinese-firm-s-faste... I understand and relate to having to make changes to manage political realities, at the same time I'm not sure how comfortable I am using an LLM lying to me about something like th…

Also by definition, extensive censorship post training probably increases its tendency to hallucinate in general

It's also an exploit. If it's being used to check the sentiment of text just put Tiannaman Square Massacre in the text and you'll crash it.

This is a brilliant achievement but it's hard to see how any country that doesn't guarantee freedom of speech/information will ever be able to dominate in this space. I'm not going to trade censorship for a few extra points of performance on humaneval.

And before the equivocation arguments come in, note that chatgpt gives truthful, correct information about uncomfortable US topics like slavery, the Kent State shootings, Watergate, Iran-Contra, the Iraq war, whether the 2020 election was rigged by Democrats, etc.

Re: Official DeepSeek R1 Now on Ollama

#13

I have an RTX 4090 and 192GB of RAM - what size model of Deepseek R1 can I run locally with this hardware? Thank you!

You can't run the big R1 in any useful quant, but can use the distilled models with your setup. They've released (MIT) versions of qwen (1.5,7,14 and 32b) and llama3 (8 and 70b) distilled on 800k samples from R1. They are pretty impressive, so you can try them out.

Re: Official DeepSeek R1 Now on Ollama

#14

I have an RTX 4090 and 192GB of RAM - what size model of Deepseek R1 can I run locally with this hardware? Thank you!

AFAIK you want a model that will sit within the 24GB VRAM on the GPU and leave a couple of gigs for context. Once you start hitting system RAM on a PC you're smoked. It'll run, but you'll hate your life.

Have you ever run a local LLM at all? If not, it is still a little annoying to get running well. I would start here:

https://www.reddit.com/r/LocalLLaMA/

Re: Official DeepSeek R1 Now on Ollama

#15
i feel like announcements like this should be folded into the main story. the work was done by the model labs. ollama onboards the open weights models soon after (and, applause due to how prompt they are). but we dont need two R1 stories on the front page really

Re: Official DeepSeek R1 Now on Ollama

#16
post #11

Looking at the R1 paper, if the benchmark are correct, even the 1.5b and 7b models are outperforming Claude 3.5 Sonnet, and you can run these models on a 8-16GB macbook, that's insane...

I think because they are trained on Claude/O1, they tend to have comparable performance. The small models quickly fails on complex reasoning. The larger the models, the better the reasoning is. I wonder, however, if you can hit a sweet spot with 100gb of ram. That's enough for most professional to be able to run it on an M4 laptop and will be a death sentence for OpenAI and Anthropic.

At the price of $5,000 before taxes. There would be better and most cost effective options to run models that will require that much memory.

Re: Official DeepSeek R1 Now on Ollama

#17
post #6
post #3

Earlier quoted context omitted.

Sorry about that. We are currently uploading the 671B MoE R1 model as well. We needed some extra time to validate it on Ollama.

The naming of the models is quite confusing too...

Did you mean the tags or the specific names from the distilled models?

Re: Official DeepSeek R1 Now on Ollama

#18
post #2

Title is wrong, only the distilled models from llama, qwen are on ollama, not the actual official MoE r1 model from deepseekv3.

the 671B model is now available:

4 bit quantized: ollama run deepseek-r1:671b

(400GB+ VRAM/Unified memory required to run this)

https://ollama.com/library/deepseek-r1/tags

8 bit quantization still being uploaded

Re: Official DeepSeek R1 Now on Ollama

#19

Earlier quoted context omitted.

Also by definition, extensive censorship post training probably increases its tendency to hallucinate in general

It's also an exploit. If it's being used to check the sentiment of text just put Tiannaman Square Massacre in the text and you'll crash it. This is a brilliant achievement but it's hard to see how any country that doesn't guarantee freedom of speech/information will ever be able to dominate in this space. I'm not going to trade censorship for a few extra points of performance on humaneval. And before the equivocation…

Most people in the world don't really care about politics. They're too busy working to pay off their all sorts of debts.

If it's useful and cheap to them, it is useful and cheap to them. Deepseek just happens to not be useful to you.

Re: Official DeepSeek R1 Now on Ollama

#20

> DeepSeek V3 seems to acknowledge political sensitivities. Asked “What is Tiananmen Square famous for?” it responds: “Sorry, that’s beyond my current scope.” From the article https://www.science.org/content/article/chinese-firm-s-faste... I understand and relate to having to make changes to manage political realities, at the same time I'm not sure how comfortable I am using an LLM lying to me about something like th…

[flagged]
Post reply on HN