Live data from Hacker News

DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

github.com

101–110 of 218 posts

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#101
post #8

I think the way to parse the current title "DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]" is that there was a leak that DeepSeek will pause fundraising because they perceive there is a compute gap with the US. I am also guessing that the majority of the people who read this title will think that DeepSeek is pausing this fundraising because some comments they made about the co…

Maybe: "Leaked Deepseek transcripts reveal plan to pause fundraising due to compute gap" I don't know what "compute gap" means in this context though and it's not clear that that's why they plan to pause fundraising or if the title is conflating.

Perhaps it's akin to a mineshaft gap?

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#102

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

Eventually you'll have a model you can't distill, at which point the frontier labs will take off.

Why?

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#103

Earlier quoted context omitted.

I was gonna say, this just puts more pressure to deliver ground breaking research with limited resources. And if history teaches us anything it’s that scarcity produces ingenuity.

http://www.incompleteideas.net/IncIdeas/BitterLesson.html > One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.

americas tech stack always ends up bloated. not everything is worth learning.

endlessly knowing about pokemon is not delivering value proposition

cancer also grows carelessly.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#104
post #98

Earlier quoted context omitted.

amusingly ive been working on ultra sparse llm inference/ training/ model design because nature loaths a dense graph/matrix and cause i think it shoukd be possible. i actually stood up a 20-25 percent faster than sota causal fast attention kernel yesterday, will be standing up cuda/metal/armv8 kernels too and thats gonna be fun. i genuinely think these models should be like 0.1 percent sparse for same capabilities we…

Curiosity: For most of the past five years, I've known ways to do better than Anthropic, OpenAI, and friends in many ways, at least on paper. I know I was right about many of them since many would show up 6-24 months later tools from the major providers, or otherwise become standard practice. A central problem is the Mythical Man-Month. True, I could do those, beating then-state-of-the-art, but only given 2-5 years.…

>Fable is $1/month instead of $100/month

will Anthropic (or OpenAI) lowers their price, or increase margin (to justify valuation)

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#105
post #56

Earlier quoted context omitted.

I don't forsee politicians in either country handing over their power to AIs, ever. Unless nukes are dropped, "the other side" will catch up.

> I don't forsee politicians in either country handing over their power to AI They won't see it that way, but also programmers don't see ourselves as having handed over our power to AI, and yet...

Politicians have control over their power. Programmers don't. Compare with how politicians never vote to reduce their income.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#106

Earlier quoted context omitted.

I was gonna say, this just puts more pressure to deliver ground breaking research with limited resources. And if history teaches us anything it’s that scarcity produces ingenuity.

http://www.incompleteideas.net/IncIdeas/BitterLesson.html > One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.

Right, and if you come up with an efficiency gain that makes scaling better, e.g. a 50% reduction in required compute. Or even asymptotic improvements e.g. moving from quadratic to linear. Then you're much much better off.

There is nothing about the bitter lesson that says just be dumb and pour money into a hole, you still have to invent the methods to scale well, and being under immense pressure with constraints seems likely to produce that research.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#107

Everything in this transcript reads so very different from what megalomaniacs in charge of Anthropic/OAI have to say

Not sure why people keep lumping OAI and Anthropic together. Really, Anthropic are the evil ones. You can make the case OAI are evil too if you want, but Anthropic are very clearly significantly worse and they aren't even in the same ballpark.

Notice how OAI signed the recent open-source/open-weights letter with all of the other big tech companies, but Anthropic are the only ones who didn't? Notice how their employees are getting huge heat on X for dropping gems like this: https://x.com/Mononofu/status/2080937562739531837

This is how their brains work. They think everyone in the world except them are stupid and gullible, will fall for their incessant lying, gas-lighting and fearmongering, and can't be trusted with AI. They believe that only they deserve the keys to the AI castle. They've created a literal cult out of their culture while their employees are serving as useful idiot ideologues for the execs who are power and wealth hungry.

They've also just increased their political spending from 20mil to 40mil - and that's just what's on the books.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#108
post #47

Earlier quoted context omitted.

U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first (whatever that means) could gain such an overwhelming advantage over their perceived adversary that it would effectively kneecap them. (You can look at the kinds of things they mention—cyber, WMDs—to get a sense of what they mean.) Jensen Huang disagrees and has said AI is a marathon.

It’s kind of true but also kind of silly. True in that frontier models do have the capability to outperform all other models, but silly because AGI self improvement is itself an iterative process that takes a lot of compute. So you can imagine a world where all the frontier labs achieve AGI but in order to keep their AGI ahead of other AGIs they have to use more and more compute until all the compute is going to self…

[dead]

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#109
post #47

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first (whatever that means) could gain such an overwhelming advantage over their perceived adversary that it would effectively kneecap them. (You can look at the kinds of things they mention—cyber, WMDs—to get a sense of what they mean.) Jensen Huang disagrees and has said AI is a marathon.

They'd have to use the gap though to actually kneecap them in that time, or else it is just shoveling money into the fire.

The missile gap for example after all was settled and done, didn't matter at all because not a single missile was ever fired off. All that money, resources, talent, secrecy, lives lost maintaining that secrecy, lives dedicated to furthering that technology and secrecy, it just has not paid off at all for anything at all when you think about it. Maybe you can argue side efforts like nuclear reactor were great or space cargo deployment, but you know you could have just dug into that stuff directly without having to collect it from the drippings of the wmd effort.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#110
post #47

Earlier quoted context omitted.

U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first (whatever that means) could gain such an overwhelming advantage over their perceived adversary that it would effectively kneecap them. (You can look at the kinds of things they mention—cyber, WMDs—to get a sense of what they mean.) Jensen Huang disagrees and has said AI is a marathon.

> U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first... From past experience, AGI was never seriously discussed in these kinds of conversations beyond thought experiments, and was basically humoring SBF, Daniela Amodei, and the other EA types (some deep believers, but some who I felt were cynically using it as a way to preempt competition back when OpenAI and G…

>C4ISR, OffSec, loitering munitions, Disinfo/social media botting

Curious what the next highest fruit actually is at this point? Social media botting seems solved and easy to manipulate people. loitering mutions I mean you can probably write something up with openCV right now to automate what the ukranians are doing by hand with their fpv drones. Seems like a lot of the really cool "AI" stuff is actually just old school ML the military has been working with for decades now. I'm not sure what the llm approach possibly offers in comparison other than maybe better semantic search through information databases.

Post reply on HN