Earlier quoted context omitted.
We will know soon enough if this replicates since Huggingface is working on replicating it. To know that this would work requires insanely deep technical knowledge about state of the art computing, and the top leadership of the PRC does not have that.
Researchers from TikTok claim they already replicated it https://x.com/sivil_taram/status/1883184784492666947?t=NzFZj...
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
471–480 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#472This might tempt me to get a graphics card and run local. What do I need minimum to run it?
It will run faster than you can read on a MacBook Pro with 192GB.
As for as fast as you can read, depends on the distilled size. I have a mac mini 64 GB Ram. The 32 GB models are quite slow. 14B and lower are very very fast.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#473Why not DeepLearn, what's it Seeking here ?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#474Earlier quoted context omitted.
The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.
It's also not a uniquely Chinese problem. You had American models generating ethnically diverse founding fathers when asked to draw them. China is doing America better than we are. Do we really think 300 million people, in a nation that's rapidly becoming anti science and for lack of a better term "pridefully stupid" can keep up. When compared to over a billion people who are making significant progress every day. Am…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#475Earlier quoted context omitted.
Those are not just-throw-money problems. Usually these tropes are limited to instagram comments. Surprised to see it here.
I know, it was simply to show the absurdity of committing $500B to marginally improving next token predictors.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#476Reddit's /r/chatgpt subreddit is currently heavily brigaded by bots/shills praising r1, I'd be very suspicious of any claims about it.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#477Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there
It's actually exactly 200 if you include the first author someone named DeepSeek-AI. For reference DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z.F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#478DeepSeek V3 came in the perfect time, precisely when Claude Sonnet turned into crap and barely allows me to complete something without me hitting some unexpected constraints. Idk, what their plans is and if their strategy is to undercut the competitors but for me, this is a huge benefit. I received 10$ free credits and have been using Deepseeks api a lot, yet, I have barely burned a single dollar, their pricing are t…
Can you tell me more about how Claude Sonnet went bad for you? I've been using the free version pretty happily, and felt I was about to upgrade to paid any day now (well, at least before the new DeepSeek).
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#479DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
The way it has destroyed the sacred commandment that you need massive compute to win in AI is earthshaking. Every tech company is spending tens of billions in AI compute every year. OpenAI starts charging 200/mo and trying to drum up 500 billion for compute. Nvidia is worth trillions on the basis it is the key to AI. How much of this is actually true?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#480Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.
Side note: I’ve read enough sci-fi to know that letting rich people live much longer than not rich is a recipe for a dystopian disaster. The world needs incompetent heirs to waste most of their inheritance, otherwise the civilization collapses to some kind of feudal nightmare.