I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
291–300 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#292Earlier quoted context omitted.
How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?
If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#293Earlier quoted context omitted.
I spent three years working in consulting mostly in public sector and education and the last two working with startups to mid size commercial interest and a couple of financial institutions. Before that I spent 6 years working between 3 companies in health care in a tech lead role. I’m 100% sure that any of those companies would I have immediately questioned my judgment for suggesting DeepSeek if had been a thing. Ab…
Why would you be presenting what AI tech you are using? You would tell them AI will come from Amazon using a variety of models.
https://www.ecfr.gov/current/title-17/chapter-II/part-240/su...
Note: I am neither a lawyer nor in financial circles, but I do have an interest in the effects of market design and regulation as we get into a more deeply automated space.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#294Earlier quoted context omitted.
Why does that matter? They wont be making at home graphics cards anymore. Why would you do that when you can be pre-sold $40k servers for years into the future
Because Moore's law marches on. We're around 35-40 orders of magnitude from computers now to computronium. We'll need 10-15 years before handheld devices can run a couple terabytes of ram, 64-128 terabytes of storage, and 80+ TFLOPS. That's enough to run any current state of the art AI at around 50 tokens per second, but in 10 years, we're probably going to have seen lots of improvements, so I'd guess conservatively…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#295Earlier quoted context omitted.
[flagged]
> CrowdStrike researchers next prompted DeepSeek-R1 to build a web application for a Uyghur community center. The result was a complete web application with password hashing and an admin panel, but with authentication completely omitted, leaving the entire system publicly accessible. > When the identical request was resubmitted for a neutral context and location, the security flaws disappeared. Authentication checks…
I don't know if I trust China or X less in this regard.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#296Earlier quoted context omitted.
Why would you be presenting what AI tech you are using? You would tell them AI will come from Amazon using a variety of models.
In various sectors, you need to be able to explain why you/your-system did what it did. Exchange Act Rule 15c3-5 is probably the most relevant in financial circles: https://www.ecfr.gov/current/title-17/chapter-II/part-240/su... Note: I am neither a lawyer nor in financial circles, but I do have an interest in the effects of market design and regulation as we get into a more deeply automated space.
https://docs.aws.amazon.com/sagemaker/latest/dg/clarify-mode...
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#297Earlier quoted context omitted.
Sparse Attention, it's the highlight of this model as per the paper
How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#298Earlier quoted context omitted.
I spent three years working in consulting mostly in public sector and education and the last two working with startups to mid size commercial interest and a couple of financial institutions. Before that I spent 6 years working between 3 companies in health care in a tech lead role. I’m 100% sure that any of those companies would I have immediately questioned my judgment for suggesting DeepSeek if had been a thing. Ab…
I've worked with financial services, and insurance providers that would have done the opposite for cost saving measures. So, I'm not sure what to say here.
If you'd spent anytime working at one for swe you won't have access to popular open source frameworks, let alone Chinese LLMs. The LLM development is mostly occurring through collaborations with the regional LLM businesses or internal labs.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#299Earlier quoted context omitted.
[flagged]
This is how crazy and nationalistic people are getting. I'm an American citizen, though I am critical of the US government, and have no allegiances to China. What do you think America is doing to every country, even allies (which has been highly publicized)? Why would a country being constantly attacked by American intelligence and propaganda not want to counter that? https://www.reuters.com/world/europe/us-security-…
Rather, I'd say it speaks more about how deranged the post-snowden/anti-neocon figures have become, from critiquing creeping authoritarianism to functionality acting at the behest of an even more authoritarian regime. The funny thing is that behavior of deflection, moralizing and whataboutism is exactly the kind of behavior nationalists employ, not addressing arguments head on like the so-called "American nationalists".
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#300Earlier quoted context omitted.
Oh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.
This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…
The way I see this, some tech teams in China have figured out that training and tuning LLMs is not that expensive after all and they can do it at a fraction of the cost. So they are doing it to enter a market previously dominated by US only players.