Worth noting this is not only good on benchmarks, but significantly more efficient at inference https://x.com/_thomasip/status/1995489087386771851
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
211–220 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#212Earlier quoted context omitted.
If the Chinese model becomes better than competitors, these worries will suddenly disappear. Also, there are plenty startups and enterprises that are running fine-tuned versions of different OS models.
Yeah that’s not how Big Enterprise works… And most startups are just doing prompt engineering that will never go anywhere. The big companies will just throw a couple of developers at the feature and add it to their existing business.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#213Earlier quoted context omitted.
If the Chinese model becomes better than competitors, these worries will suddenly disappear. Also, there are plenty startups and enterprises that are running fine-tuned versions of different OS models.
No… Nobody I work for will touch these models. The fear is real that they have been poisoned or have some underlying bomb. Plus y’know, they’re produced by China, so they would never make it past a review board in most mega enterprises IME.
Companies just need to get to the “if” part first. That or they wash their hand by using a reseller that can use whatever it wants under the hood.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#214Earlier quoted context omitted.
Competitor != adversary. It is US warmongering ideology that tries to equate these concepts.
[flagged]
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#215Earlier quoted context omitted.
on what hypothetical grounds would you be more meaningfully able to sue the american maker of a self-hosted statistical language model that you select your own runtime sampling parameters for after random subtle security vulnerabilities came out the other side when you asked it for very secure code? put another way, how do you propose to tell this subtle nefarious chinese sabotage you baselessly imply to be commonpla…
[flagged]
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#216How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have their own proprietary DynamoDB but for the most people scale up MongoDB is more than suffice. Regardless brand or type of databases being used, you paid tons of money to Amazon anyway to be at scale.
Not everyone has the resource to host a SOTA AI model. On top of tangible data-intensive resources, they are other intangible considerations. Just think how many company or people host their own email server now although the resources needed are far less than hosting an AI/LLM model?
Google came up with the game changing transformer at its backyard and OpenAI temporarily stole the show with the well executed RLHF based system of ChatGPT. Now the paid users are swinging back to Google with its arguably more superior offering. Even Google now put AI summary as its top most search return results for free to all, higher than its paid advertisement clients.
[1]Google “We have no moat, and neither does OpenAI”:
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#217Earlier quoted context omitted.
How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?
If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…
What's cheap nowdays? I'm out of the loop. Does anything ever run on integrated AMD that is Ryzen AI that comes in framework motherboards? Is under 1k americans cheap?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#218Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#219Earlier quoted context omitted.
I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…
That might be the perspective of a US based company. But there is also Europe and basically it's a choice between Trump and China.