Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

851–860 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#851
post #613

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Funny, maybe OpenAI will achieve their initial stated goals of propelling AI research, spend investors money and be none profit. Functionally the same as their non-profit origins.

>Funny, maybe OpenAI will achieve their initial stated goals of propelling AI research, spend investors money and be none profit. Functionally the same as their non-profit origins.

Serves them right!!! This hopefully will give any non-profit pulling an OpenAI in going for-profit a second thought!!!! If you wanna go for-profit it is fine, just say it! Don't get the good will of community going and then do a bait and switch.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#852

Earlier quoted context omitted.

Most people I talked with don't grasp how big of an event this is. I consider is almost as similar to as what early version of linux did to OS ecosystem.

Precisely. This lets any of us have something that until the other day would have cost hundreds of millions of dollars. It's as if Linus had published linux 2.0, gcc, binutils, libc, etc. all on the same day .

people are doing all sort of experiments and reproducing the "emergence"(sorry it's not the right word) of backtracking; it's all so fun to watch.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#853
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

Actually it means we will potentially get 100x the economic value out of those datacenters. If we get a million digital PHD researchers for the investment then that’s a lot better than 10,000.

[deleted]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#854
post #707

Earlier quoted context omitted.

They just got 500 billion and they'll probably make that back in military contracts so this is unlikely (unfortunately)

that would be like 75%+ of the entire military budget

… in a year. Theirs is for 4 years.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#855
post #594

Earlier quoted context omitted.

I don't get it. I like DeepSeek, because I can turn on Search button. Turning on Deepthink R1 makes the results as bad as Perplexity. The results make me feel like they used parallel construction, and that the straightforward replies would have actually had some value. Claude Sonnet 3."6" may be limited in rare situations, but its personality really makes the responses outperform everything else when you're trying to…

Interesting thinking. Curious––what would you want to "edit" in the thought process if you had access to it? or would you just want/expect transparency and a feedback loop?

I ran the llama distill on my laptop and I edited both the thoughts and the reply. I used the fairly common approach of giving it a task, repeating the task 3 times with different input and adjusting the thoughts and reply for each repetition. So then I had a starting point with dialog going back and forth where the LLM had completed the task correctly 3 times. When I gave it a fourth task it did much better than if I had not primed it with three examples first.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#856
post #822

Earlier quoted context omitted.

Interesting thinking. Curious––what would you want to "edit" in the thought process if you had access to it? or would you just want/expect transparency and a feedback loop?

I personally would like to "fix" the thinking when it comes to asking these models for help on more complex and subjective problems. Things like design solutions. Since a lot of these types of solutions are belief based rather than fact based, it's important to be able to fine-tune those beliefs in the "middle" of the reasoning step and re-run or generate new output. Most people do this now through engineering longwi…

If you run one of the distill versions in something like LM Studio it’s very easy to edit. But the replies from those models isn’t half as good as the full R1, but still remarkably better then anything I’ve run locally before.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#857

Earlier quoted context omitted.

Probably shouldn't be firing their blood boys just yet ... According to Musk, SoftBank only has $10B available for this atm.

Elon says a lot of things.

While doing a lot of "gestures".

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#858

Earlier quoted context omitted.

IMO the deep think button works wonders.

Whenever I use it, it just seems to spin itself in circles for ages, spit out a half-assed summary and give up. Is it like the OpenAI models in that in needs to be prompted in extremely-specific ways to get it to not be garbage?

O1 doesn’t seem to need any particularly specific prompts. It seems to work just fine on just about anything I give it. It’s still not fantastic, but often times it comes up with things I either would have had to spend a lot of time to get right or just plainly things I didn’t know about myself.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#859

I have just tried ollama's r1-14b model on a statistics calculation I needed to do, and it is scary to see how in real time the model tries some approaches, backtracks, chooses alternative ones, checka them. It really reminds of human behaviour...

Please try QwQ 32B with the same question. In my experience it's even more "humane" while approaching a hard question.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#860
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

When I try to Sign Up with Email. I get.

>I'm sorry but your domain is currently not supported.

What kind domain email does deepseek accept?

Post reply on HN