Earlier quoted context omitted.
Sir, this is a multi-billion dollar operation. There have to be some incentives.
You could argue that Linux is a trillion-dollar project, it doesn't shift the goalposts.
Qwen 3.8
451–460 of 793 posts
Re: Qwen 3.8
#452Deepseek 4 "final" version is imminent as well. Will probably be at Opus 4.8 level, and I find it pretty big deal because of Deepseek price...
Re: Qwen 3.8
#453Earlier quoted context omitted.
Yeah the Chinese totally have a really good history with being completely open and giving lol. The Chinese government totally has not been hacking into American and Western fortune 500 companies for the past few decades stealing R&D and tech to use for themselves. The Chinese also totally do not steal hundreds of billions of dollars of IP from America annually. Totally not something they would do!
It is hilarious to see people from arguable the most polarized political systems in the world believing the evil 1.5 billion people across the sea share one single mind, either a saint, or a devil.
Re: Qwen 3.8
#454Earlier quoted context omitted.
Yea the performance/price ratio for Deepseek is off the charts. I’ve been using V4 Flash a lot lately and it’s quite good.
I like that v4 flash is so fast! I run it on both FireWorks.ai in the US and bought some tokens directly from DeepSeek as an experiment. I only work on Open Source projects, so I don’t have to worry about my work being used to train models - I welcome AI’s being trained on my open content books and code (but not my conventionally published books: I am a party to the copyright suit against Anthropic).
Re: Qwen 3.8
#455> "compatible" instead of "comparable." It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.
if youre going to use ai for everything youre gonna start losing your edge as you focus less and less on what youre doing and this isnt me just talking out my ass, like... the front page here is peppered with study after study and blogpost after blogpost about how its overuse can come to the detriment of one's own abilities and skills.
coca cola had the ad with the magical truck that changed its design, shape and amount of tires it had and if nobody noticed that before releasing it then im not sure why anyone might think that the people peddling the LLMs would somehow be immune to this phenomenon
Re: Qwen 3.8
#456I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
These things take months to train. No chance this is a reaction to what just happened.
Re: Qwen 3.8
#457Re: Qwen 3.8
#458Earlier quoted context omitted.
Sir, this is a multi-billion dollar operation. There have to be some incentives.
You could argue that Linux is a trillion-dollar project, it doesn't shift the goalposts.
Re: Qwen 3.8
#459I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
I keep looking at the numbers. The power use numbers are not that problematic. Ordering a burrito on DoorDash uses more power than a few days of heavy AI use. The water argument applies to some locations, and is mostly a local governance problem... if the data centers are using too much water, it means they are not being charged enough for that water. Charge them more and they'll push toward closed loop cooling.
Yet the visceral pile-on here is so extreme, it feels fake.
One thing I've learned after 40 years on this planet is: propaganda works, and much of what a large fraction of people believe across the entire political spectrum (left, right, anything else) is there because someone paid to put it there. It's depressing but it's true, and it makes sense. Propaganda is an asymmetrical attack on human cognition and discourse, and in information security the attacker always has an easier job. Crafting viral bullshit is orders of magnitude easier than fact checking. On top of this, humans are busy and don't have time to fact check and logic check everything they read. As a result, much of what we believe is "sponsored content."
People get mad when you talk about this because everyone wants to believe they're too smart to fall for propaganda.
In any case, the US AI labs deserve to lose for their stupid "safety" regulatory capture monopolization push, which ended up blowing their own feet off and handing the lead to China.
Re: Qwen 3.8
#460Earlier quoted context omitted.
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
> It's hard to say what their motivation is. Feels pretty easy to me. They want to turn LLMs into a commodity, and watch the US AI labs crash and burn. There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens an…