Live data from Hacker News

DeepSeek-V3

github.com

11–20 of 44 posts

Re: DeepSeek-V3

#11
post #10
post #9

Earlier quoted context omitted.

well yes, locally, if you assume that someone's got about 300'000 dollars of hardware at hand... right? as you are not paying for Gemini, may I ask why, did you try it and find it inferior?

Gemini for coding does not work for me. It gets so many things wrong

You should try again. Gemini rates highest on coding at lmarena.

Re: DeepSeek-V3

#12
It still fails my private physics testing question half the time, where claude 3.5 sonnet and openai o1 (both web version) most of the time passes. So I'd say close to SOTA but not quite. However given deekseek already has the r1 lite preview, and they can achieve comparable performance for much less compute (assuming the API cost of close models roughly represent the inference cost), then it's not unreasonable to believe deepseek may be close to release very good test compute scaling model that is similar to o3 high effort.

Re: DeepSeek-V3

#13
post #8
post #3

You can run a model that can beat 4o which was released less than 6 months ago _locally_! I know this requires a ton of hardware but OpenAI will not be the leader in 2025 I can assume. Always bet on open source (or rather somewhat more open development strategies) The math and coding performance is what we really care about. I am paying for o1 Pro and also Sonnet, in my experience beside Sonnet being faster, it is al…

OpenAi toppled as LLM leader by an open source / open weight company? OpenAi has much more capital and compute than any of its competitors (especially deepseek); if that was to happen it would demonstrate that capital and compute doesn't matter as much as it is assumed ... (and it just might be the thing that pops the current ai bubble).

> OpenAi has much more capital and compute than any of its competitors

Isn't openai still losing money? I don't think they own any data centers.

Re: DeepSeek-V3

#14
In the introduction of the paper it says: "Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks." They have indeed a very strong infra team.

Re: DeepSeek-V3

#15

In the introduction of the paper it says: "Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks." They have indeed a very strong infra team.

[deleted]

Re: DeepSeek-V3

#16
Truly remarkable! Their approach to distributed inference is on an entirely new level. For the prefill stage, they utilized a deployment unit comprising 32 H800 GPUs, while the decoding stage scaled up to 320!! H800 GPUs per unit. Incorporates a multitude of sophisticated parallelization and communication overlap techniques, setting a standard that’s rarely seen in other setups.

[0] https://github.com/deepseek-ai/DeepSeek-V3/blob/main/DeepSee...

Re: DeepSeek-V3

#17
post #9
post #3

You can run a model that can beat 4o which was released less than 6 months ago _locally_! I know this requires a ton of hardware but OpenAI will not be the leader in 2025 I can assume. Always bet on open source (or rather somewhat more open development strategies) The math and coding performance is what we really care about. I am paying for o1 Pro and also Sonnet, in my experience beside Sonnet being faster, it is al…

well yes, locally, if you assume that someone's got about 300'000 dollars of hardware at hand... right? as you are not paying for Gemini, may I ask why, did you try it and find it inferior?

You actually can't pay for the latest models, they're only available as free with limits

Re: DeepSeek-V3

#18
> a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.

What kind of hardware do you need to run this?

Re: DeepSeek-V3

#19
post #9
post #3

You can run a model that can beat 4o which was released less than 6 months ago _locally_! I know this requires a ton of hardware but OpenAI will not be the leader in 2025 I can assume. Always bet on open source (or rather somewhat more open development strategies) The math and coding performance is what we really care about. I am paying for o1 Pro and also Sonnet, in my experience beside Sonnet being faster, it is al…

well yes, locally, if you assume that someone's got about 300'000 dollars of hardware at hand... right? as you are not paying for Gemini, may I ask why, did you try it and find it inferior?

I bought two (relatively) old datacenter GPUs with 48gb VRAM total for €200 that gets me 7 token/s for a 70b model.
Post reply on HN