Earlier quoted context omitted.
I would think Meta - who open source their model - would be less freaked out than those others that do not.
The criticism seems to mostly be that Meta maintains very expensive cost structure and fat organisation in the AI. While Meta can afford to do this, if smaller orgs can produce better results it means Meta is paying a lot for nothing. Meta shareholders now need to ask the question how many non-productive people Meta is employing and is Zuck in the control of the cost.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
281–290 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#282Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)
Who cares? I ask O1 how to download a YouTube music playlist as a premium subscriber, and it tells me it can't help. Deepseek has no problem.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#283Earlier quoted context omitted.
Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…
The $500B is just an aspirational figure they hope to spend on data centers to run AI models, such as GPT-o1 and its successors, that have already been developed. If you want to compare the DeepSeek-R development costs to anything, you should be comparing it to what it cost OpenAI to develop GPT-o1 (not what they plan to spend to run it), but both numbers are somewhat irrelevant since they both build upon prior resea…
The existing hyperscalers have already sunk ungodly amounts of money into literally hundreds of new data centers, millions of GPUs to fill them, chip manufacturing facilities, and even power plants with the impression that, due to the amount of compute required to train and run these models, there would be demand for these things that would pay for that investment. Literally hundreds of billions of dollars spent already on hardware that’s already half (or fully) built, and isn’t easily repurposed.
If all of the expected demand on that stuff completely falls through because it turns out the same model training can be done on a fraction of the compute power, we could be looking at a massive bubble pop.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#284Earlier quoted context omitted.
From my casual read, right now everyone is on reputation tarnishing tirade, like spamming “Chinese stealing data! Definitely lying about everything! API can’t be this cheap!”. If that doesn’t go through well, I’m assuming lobbyism will start for import controls, which is very stupid. I have no idea how they can recover from it, if DeepSeek’s product is what they’re advertising.
So you're saying that this is the end of OpenAI? Somehow I doubt it.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#285Earlier quoted context omitted.
> I guess censorship doesnt have as bad a rep in china as it has here It's probably disliked, just people know not to talk about it so blatantly due to chilling effects from aforementioned censorship. disclaimer: ignorant American, no clue what i'm talking about.
on the topic of censorship, US LLMs' censorship is called alignment. llama or ChatGPT's refusal on how to make meth or nuclear bombs is the same as not answering questions abput Tiananmen tank man as far as the matrix math word prediction box is concerned.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#286Earlier quoted context omitted.
And (some people here are saying that)* if they are up-to-date is because they're cheating. The copium itt is astounding.
What’s the difference between what they do and what other ai firms do to openai in the us? What is cheating in a business context?
PS: I'm not criticizing them for it nor do I really care if they cheat as long as prices go down. I'm just observing and pointing out what other posters are saying. For me if China cheating means the GenAI bubble pops, I'm all for it. Plus no actor is really clean in this game, starting with OAI practically stealing all human content without asking for building their models.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#287Earlier quoted context omitted.
Not necessarily if you are pushing against a data wall. One could ask: after adjusting for DS efficiency gains how much more compute has OpenAI spent? Is their model correspondingly better? Or even DS could easily afford more than $6 million in compute but why didn't they just push the scaling?
right except that r1 is demoing the path of approach for moving beyond the data wall
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#288Earlier quoted context omitted.
In Communist theoretical texts the term "propaganda" is not negative and Communists are encouraged to produce propaganda to keep up morale in their own ranks and to produce propaganda that demoralize opponents. The recent wave of the average Chinese has a better quality of life than the average Westerner propaganda is an obvious example of propaganda aimed at opponents.
Is it propaganda if it's true?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#289I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…