Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)
I am not surprised if US Govt would mandate "Tiananmen-test" for LLMs in the future to have "clean LLM". Anyone working for federal govt or receiving federal money would only be allowed to use "clean LLM"
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
341–350 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#342DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#343Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#344Earlier quoted context omitted.
keyboard warrior strikes again lol. Most people would be thrilled to even be a small contributor in a tech initiative like this. call it what you want, your comment is just poor taste.
When Google did this with the recent Gemini paper, no one had any problem with calling it out as credential stuffing, but when Deepseek does it, it’s glorious unity and camaraderie.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#345Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years
More like three years. Even in the best case the retained value curve of GPUs is absolutely terrible. Most of these huge investments in GPUs are going to be massive losses.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#346DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.
LLaMA was huge, Byte Latent Transformer looks promising.. absolutely no idea were you got this idea from.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#347Earlier quoted context omitted.
Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.
DeepSeek was built on the foundations of public research, a major part of which is the Llama family of models. Prior to Llama open weights LLMs were considerably less performant; without Llama we might not have gotten Mistral, Qwen, or DeepSeek. This isn't meant to diminish DeepSeek's contributions, however: they've been doing great work on mixture of experts models and really pushing the community forward on that fr…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#348Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#349Earlier quoted context omitted.
The $500B is just an aspirational figure they hope to spend on data centers to run AI models, such as GPT-o1 and its successors, that have already been developed. If you want to compare the DeepSeek-R development costs to anything, you should be comparing it to what it cost OpenAI to develop GPT-o1 (not what they plan to spend to run it), but both numbers are somewhat irrelevant since they both build upon prior resea…
Thinking of the $500B as only an aspirational number is wrong. It’s true that the specific Stargate investment isn’t fully invested yet, but that’s hardly the only money being spent on AI development. The existing hyperscalers have already sunk ungodly amounts of money into literally hundreds of new data centers, millions of GPUs to fill them, chip manufacturing facilities, and even power plants with the impression t…
It's just data centers full of devices optimized for fast linear algebra, right? These are extremely repurposeable.