Live data from Hacker News

Deepseek: The quiet giant leading China’s AI race

chinatalk.media

281–290 of 476 posts

Re: Deepseek: The quiet giant leading China’s AI race

#281
post #267

Earlier quoted context omitted.

Sorry, I edited my comment to ask this as the conclusion: > Why is a wet market a more convincing explanation than a lab whose explicit mission is studying coronaviruses, one which has been publicly called out for being uncooperative with the ensuing investigation? It's possible that the widespread belief in this explanation is a failure in science communication and there's a good reason for this, but it's not a fail…

I don't think the desire to see the conspiracy and not the boring, mundane explanation is a failure of communication, scientific thinking, or critical reasoning skills. I think it's simply the consequences of politically motivated reasoning. (At least, for people who have spent much time thinking about it.) > I've yet to see anyone effectively communicate why the intuitively more probable answer is the less probable…

> I just communicated why the market leak theory is both more intuitive, and more probable.

No, you didn't, you stated that it was.

The rest of your post makes sense as an explanation. Maybe lead with that next time instead of condescendingly telling people that they're politically motivated, stupid, or whatever else you meant to imply by calling it a conspiracy theory.

COVID-19, at least in the US, has been an enormous failure in science communication, and being condescending towards those who already feel alienated by the terrible communication isn't going to help.

Re: Deepseek: The quiet giant leading China’s AI race

#282
post #161

To this day, asking Deepseek "what model are you" typically gives the answer "I'm an AI language model called ChatGPT, created by OpenAI. Specifically, I'm based on the GPT-4 architecture, which is designed to understand and generate human-like text based on the input I receive. My training data includes a wide range of information up until October 2023, and I can assist with answering questions, generating text, and…

This is a common "gotcha" comment from people who don't understand LLMs very well. Occasionally if you ask Gemini it'll say this as well. It has everything to do with the fact that ChatGPT is the most talked about AI model rather than data being trained on it

When Gemini had its 'identity crisis,' mostly in 2023 (whatever they called it back then), it looked like a pre-training phenomenon, explainable by mentions in the available corpus. The consistency here, and exact and persistent match to how ChatGPT would answer identical queries, suggest training on Q&A pairs from ChatGPT transcripts, presumably at the post-training stage.

Re: Deepseek: The quiet giant leading China’s AI race

#283
post #88

Not personally surprised that a MoE model performs so well. I used Mixtral a lot for coding Rust, and it had qualities no other model had except GPT 3.5 and later Claude Sonet. The funny thing is Mixtral was based on Llama 2 which was not trained on code that much. DeepSeek v3: 671B parameters on total, and 37B activated sounds very good even though impossible to run locally. Question if some people happen to know: F…

DeepSeek v3 can run on CPU & RAM :

https://www.reddit.com/r/LocalLLaMA/comments/1hqidbs/deepsee...

Epyc Gen4 and 12 memory channels of DDR5 @4800 should give you 7 to 9 t/s.

Re: Deepseek: The quiet giant leading China’s AI race

#284
post #267

Earlier quoted context omitted.

I don't think the desire to see the conspiracy and not the boring, mundane explanation is a failure of communication, scientific thinking, or critical reasoning skills. I think it's simply the consequences of politically motivated reasoning. (At least, for people who have spent much time thinking about it.) > I've yet to see anyone effectively communicate why the intuitively more probable answer is the less probable…

> I just communicated why the market leak theory is both more intuitive, and more probable. No, you didn't, you stated that it was. The rest of your post makes sense as an explanation. Maybe lead with that next time instead of condescendingly telling people that they're politically motivated, stupid, or whatever else you meant to imply by calling it a conspiracy theory. COVID-19, at least in the US, has been an enorm…

The meat of the argument of the post was a restatement of the past three posts that I've made. I did lead with the argument, in somewhat less detail.

Re: Deepseek: The quiet giant leading China’s AI race

#285
post #284

Earlier quoted context omitted.

> I just communicated why the market leak theory is both more intuitive, and more probable. No, you didn't, you stated that it was. The rest of your post makes sense as an explanation. Maybe lead with that next time instead of condescendingly telling people that they're politically motivated, stupid, or whatever else you meant to imply by calling it a conspiracy theory. COVID-19, at least in the US, has been an enorm…

The meat of the argument of the post was a restatement of the past three posts that I've made. I did lead with the argument, in somewhat less detail.

No, you didn't, you led with this:

> In 2024, this isn't fact, it's just baseless conspiracy.

> All evidence has ended up pointing to bush meat contamination.

This isn't science communication, it's a condescending rebuke.

That said, I'm done here. Thanks for clarifying in the end, and happy new year!

Re: Deepseek: The quiet giant leading China’s AI race

#286
post #188

Earlier quoted context omitted.

> China creates a superbug via GOF research In 2024, this isn't fact, it's just baseless conspiracy. All evidence has ended up pointing to bush meat contamination.

Cleavage sites

https://news.ycombinator.com/item?id=42513063

Also Baltimore changed his mind.

Re: Deepseek: The quiet giant leading China’s AI race

#287

To this day, asking Deepseek "what model are you" typically gives the answer "I'm an AI language model called ChatGPT, created by OpenAI. Specifically, I'm based on the GPT-4 architecture, which is designed to understand and generate human-like text based on the input I receive. My training data includes a wide range of information up until October 2023, and I can assist with answering questions, generating text, and…

Counterpoint : your message is not synthetic data and will contribute to lots of LLMs saying the same. Many such cases ?

(It seems to me obvious that a fgrep would sanitize synthetic data obtained from competitors.)

Re: Deepseek: The quiet giant leading China’s AI race

#288
post #230

Earlier quoted context omitted.

The jump between 'data was withheld' and 'there was a coverup of an incredibly improbable thing happening' compared to 'very probable thing that was also supported by data happening' is colossal. Absence of data isn't a free pass that lets you fill in whatever blanks you want, to fit whatever improbable theory you want. Especially when a plausible, probable, data supported alternative exists. At the moment, given wha…

Sorry, I edited my comment to ask this as the conclusion: > Why is a wet market a more convincing explanation than a lab whose explicit mission is studying coronaviruses, one which has been publicly called out for being uncooperative with the ensuing investigation? It's possible that the widespread belief in this explanation is a failure in science communication and there's a good reason for this, but it's not a fail…

https://news.ycombinator.com/item?id=42513063 (See the astral codex ten link)

Re: Deepseek: The quiet giant leading China’s AI race

#289
post #6

I feel the GPU restrictions created an environment for Chinese Devs to be more innovative and do more with less. Kudos to the deepseek team!

Software expands to fill the available resources. If you want more efficient software, build it on less powerful hardware. AI training runs are no exception!

> Software expands to fill the available resources.

To make an analogy, that is why I think even with AI work expands to fill available resources (human+AI). I don't think jobless rate will be high, instead we will see demand expansion.

Re: Deepseek: The quiet giant leading China’s AI race

#290

I find that the gushing around deepseek is fascinating to watch. To me there are a few structural and fundamental reasons why deepseek can never outperform other models by a wide margin. On par maybe--as we reach the diminishing returns with our investment in the models, but not win by a wide margin. 1. The US trade war with china which will place deepseek compute availability at disadvantages, eventually, if we ever…

I don't think it's necessarily about DeepSeek, but about the wider competitive picture. There are two tacit assumptions being made about LLMs - that having a SOTA model is a substantial competitive advantage, and that the demand for compute will continue to grow rapidly. DeepSeek's phenomenal success in reducing training and inference cost points to the possibility of a very different future. If it's the case that SO…

> If DeepSeek don't have a competitive advantage, then no-one has a competitive advantage.

There is no moat. Smaller models are just a few months behind large proprietary ones. But the distribution of tasks might be increasingly solvable with smaller models, leaving little for the top models which are also more expensive.

Post reply on HN