Who's afraid of Chinese models?
141–150 of 965 posts
Re: Who's afraid of Chinese models?
#142Earlier quoted context omitted.
mmm, the chinese models are also working on local GPUs at consumer grades. so theyre not just drainig cloud moats.
good luck running a 2.4T model on any local hardware. it’s not gonna happen. the arrow is to specialized hardware at least for the smartest models
Re: Who's afraid of Chinese models?
#143Earlier quoted context omitted.
>What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight s…
I'm arguing we can't trust retail prices because the marginal pricing isn't meaningfully connected to it anyway. But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price ( https://openrouter.ai/deepseek/deepseek-v4-flash ), so there's at least one example of a Chinese lab charging a predetermined price despite co…
their competitors are discounted at around 33%, so it's safe to say that's the margin, maybe less if their competitors have worse caching or quantization. Meanwhile claude code/codex resellers selling tokens for 90% off API price, presumably by reselling usage from fixed consumption plans, which gives an idea on how fat the american labs' margins are.
>And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
But composer is a closed model? If it's really that easy to get better coding performance, why haven't the chinese labs replicated it? And this is all assuming the performance boost is real and not from benchmaxxing. Moreover if you apply the "street price" discount I mentioned above, American labs look far more favorable.
Re: Who's afraid of Chinese models?
#144> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…
Felony contempt of business model.
Re: Who's afraid of Chinese models?
#145> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…
Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs. But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
Correct. It shouldn't be allowed to do that.
Re: Who's afraid of Chinese models?
#146Earlier quoted context omitted.
Forbidding distillation is like forbidding using a compiler to make another(perhaps better, more efficient) compiler.
Lots of software licenses have “non-compete” clauses that forbid you from using it to develop a competing product. Wouldn’t surprise me if there was a compiler or two out there with that restriction, most likely niche languages.
Re: Who's afraid of Chinese models?
#147The fact that Anthropic has a model like Mythos means that counterpart countries like Russia and China are not far behind, if they haven't already developed something similar or better.
lmao
Re: Who's afraid of Chinese models?
#148That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.
His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".
Re: Who's afraid of Chinese models?
#149According to openAI's own @deanwball: Even OpenAI isn't buying this distillation talk: https://xcancel.com/deanwball/status/2078133895766114412#m
Re: Who's afraid of Chinese models?
#150The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models…