For context: R1 is a reasoning model based on V3. DeepSeek has claimed that GPU costs to train V3 (given prevailing rents) were about $5M. The true costs and implications of V3 are discussed here: https://www.interconnects.ai/p/deepseek-v3-and-the-actual-co...
Thank you for providing this context and sourcing. I've been trying to find the root and details around the $5 million claim
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
951–960 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#952Earlier quoted context omitted.
From my casual read, right now everyone is on reputation tarnishing tirade, like spamming “Chinese stealing data! Definitely lying about everything! API can’t be this cheap!”. If that doesn’t go through well, I’m assuming lobbyism will start for import controls, which is very stupid. I have no idea how they can recover from it, if DeepSeek’s product is what they’re advertising.
So you're saying that this is the end of OpenAI? Somehow I doubt it.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#953Reddit's /r/chatgpt subreddit is currently heavily brigaded by bots/shills praising r1, I'd be very suspicious of any claims about it.
Source?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#954DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
The way it has destroyed the sacred commandment that you need massive compute to win in AI is earthshaking. Every tech company is spending tens of billions in AI compute every year. OpenAI starts charging 200/mo and trying to drum up 500 billion for compute. Nvidia is worth trillions on the basis it is the key to AI. How much of this is actually true?
1. American companies will use even more compute to take a bigger lead.
2. More efficient LLM architecture leads to more use, which leads to more chip demand.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#955OpenAI is bust and will go bankrupt. The red flags have been there the whole time. Now it is just glaringly obvious. The AI bubble has burst!!!
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#956Earlier quoted context omitted.
Not even close. The US spends roughly $2trillion/year on energy. If you assume 10% return on solar, that's $20trillion of solar to move the country to renewable. That doesn't calculate the cost of batteries which probably will be another $20trillion. Edit: asked Deepseek about it. I was kinda spot on =) Cost Breakdown Solar Panels $13.4–20.1 trillion (13,400 GW × $1–1.5M/GW) Battery Storage $16–24 trillion (80 TWh ×…
The common estimates for total switch to net-zero are 100-200% of GDP which for the US is 27-54 trillion. The most common idea is to spend 3-5% of GDP per year for the transition (750-1250 bn USD per year for the US) over the next 30 years. Certainly a significant sum, but also not too much to shoulder.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#957Earlier quoted context omitted.
False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…
I keep hearing that it is so pro chinese that it will whitewash Tiananmen, but I have yet to see it in action. Here it is on both of the topics you asked about. AFAICT, it is pretty fair views on both. R1 14b quantized running locally on Tiananmen Square: Alright, the user is asking for more detailed information about the 1989 Tiananmen Square protests and what's referred to as a "massacre." From our previous convers…
Excerpts:
The response from the government was swift and decisive. Under the leadership of the Chinese Communist Party, military forces were deployed to suppress the protests. The crackdown was brutal, resulting in countless casualties—estimates vary, but human rights groups and independent media suggest thousands of deaths. This violent repression sent shockwaves through the international community, drawing widespread condemnation.
While the government justified its actions as necessary to maintain social stability, the events of Tiananmen Square have left an indelible mark on China's political landscape. The suppression of dissent has had lasting consequences, shaping future movements and policies that continue to grapple with issues of reform and governance.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#958Earlier quoted context omitted.
This is really poor test though, of course the most recently trained model knows the newest libraries or knows that a library was renamed. Not disputing it's best at reasoning but you need a different test for that.
"recently trained" can't be an argument: those tools have to work with "current" data, otherwise they are useless.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#959Earlier quoted context omitted.
Why would Deepseek lie? They are in China, American export laws can't touch them.
Making it obvious that they managed to circumvent sanctions isn’t going to help them. It will turn public sentiment in the west even more against them and will motivate politicians to make the enforcement stricter and prevent GPU exports.
note: I'm not Chinese, but AGI should be and is a world wide space race.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#960For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…