Live data from Hacker News

Official DeepSeek R1 Now on Ollama

ollama.com

21–30 of 97 posts

Re: Official DeepSeek R1 Now on Ollama

#21
post #16
post #11

Earlier quoted context omitted.

I think because they are trained on Claude/O1, they tend to have comparable performance. The small models quickly fails on complex reasoning. The larger the models, the better the reasoning is. I wonder, however, if you can hit a sweet spot with 100gb of ram. That's enough for most professional to be able to run it on an M4 laptop and will be a death sentence for OpenAI and Anthropic.

At the price of $5,000 before taxes. There would be better and most cost effective options to run models that will require that much memory.

It is a laptop. The memory is also shared which means if you are looking for a non-gaming workload, you can use it. If you have laptop equivalents in the same memory range, feel free to share.

Re: Official DeepSeek R1 Now on Ollama

#22

Earlier quoted context omitted.

Also by definition, extensive censorship post training probably increases its tendency to hallucinate in general

It's also an exploit. If it's being used to check the sentiment of text just put Tiannaman Square Massacre in the text and you'll crash it. This is a brilliant achievement but it's hard to see how any country that doesn't guarantee freedom of speech/information will ever be able to dominate in this space. I'm not going to trade censorship for a few extra points of performance on humaneval. And before the equivocation…

Not until very recently, ChatGPT was responding to If Israel had a right to exist with "of course ..." and If Palestine had a right to exist with "It's complicated ..."

So I don't think our version is completely free of bias. I'm sure there are many other examples, I just wouldn't be able to point them out, considering the training data fed into ChatGPT was also fed into our human brains.

Re: Official DeepSeek R1 Now on Ollama

#23

Earlier quoted context omitted.

It's also an exploit. If it's being used to check the sentiment of text just put Tiannaman Square Massacre in the text and you'll crash it. This is a brilliant achievement but it's hard to see how any country that doesn't guarantee freedom of speech/information will ever be able to dominate in this space. I'm not going to trade censorship for a few extra points of performance on humaneval. And before the equivocation…

Most people in the world don't really care about politics. They're too busy working to pay off their all sorts of debts. If it's useful and cheap to them, it is useful and cheap to them. Deepseek just happens to not be useful to you.

You missed my point - post training censorship increases likelihood of hallucination in general

Re: Official DeepSeek R1 Now on Ollama

#24
post #15

i feel like announcements like this should be folded into the main story. the work was done by the model labs. ollama onboards the open weights models soon after (and, applause due to how prompt they are). but we dont need two R1 stories on the front page really

these are smaller qantized models that I can use on my 8 year old GPU, I can't even load the original deeppseek unqantized models

Re: Official DeepSeek R1 Now on Ollama

#25

> DeepSeek V3 seems to acknowledge political sensitivities. Asked “What is Tiananmen Square famous for?” it responds: “Sorry, that’s beyond my current scope.” From the article https://www.science.org/content/article/chinese-firm-s-faste... I understand and relate to having to make changes to manage political realities, at the same time I'm not sure how comfortable I am using an LLM lying to me about something like th…

That’s very likely coming from the API, not the model

Re: Official DeepSeek R1 Now on Ollama

#27

Earlier quoted context omitted.

Most people in the world don't really care about politics. They're too busy working to pay off their all sorts of debts. If it's useful and cheap to them, it is useful and cheap to them. Deepseek just happens to not be useful to you.

You missed my point - post training censorship increases likelihood of hallucination in general

It's just on Chinese politics, on a very select topics too.

Yeah I think Deepseek will be just fine.

Re: Official DeepSeek R1 Now on Ollama

#28
post #16
post #11

Earlier quoted context omitted.

I think because they are trained on Claude/O1, they tend to have comparable performance. The small models quickly fails on complex reasoning. The larger the models, the better the reasoning is. I wonder, however, if you can hit a sweet spot with 100gb of ram. That's enough for most professional to be able to run it on an M4 laptop and will be a death sentence for OpenAI and Anthropic.

At the price of $5,000 before taxes. There would be better and most cost effective options to run models that will require that much memory.

I see this comment all the time. But realistically if you want more than 1 token/s you’re going to need geforces, and that would cost quite a lot as well, for 100 GB.

Re: Official DeepSeek R1 Now on Ollama

#29
post #11

Looking at the R1 paper, if the benchmark are correct, even the 1.5b and 7b models are outperforming Claude 3.5 Sonnet, and you can run these models on a 8-16GB macbook, that's insane...

I think because they are trained on Claude/O1, they tend to have comparable performance. The small models quickly fails on complex reasoning. The larger the models, the better the reasoning is. I wonder, however, if you can hit a sweet spot with 100gb of ram. That's enough for most professional to be able to run it on an M4 laptop and will be a death sentence for OpenAI and Anthropic.

> I think because they are trained on Claude/O1, they tend to have comparable performance.

Why does having comparable performance indicate having been trained on a preexisting model's output?

I read a similar claim in relation to another model in the past, so I'm just curious how this works technically.

Post reply on HN