> Oops! There was an issue connecting to Qwen3-Max.
> Content Security Warning: The input text data may contain inappropriate content.
431–440 of 450 posts
> Oops! There was an issue connecting to Qwen3-Max.
> Content Security Warning: The input text data may contain inappropriate content.
It just occured to me that it underperforms Opus 4.5 on benchmarks when search is not enabled, but outperforms it when it is - is it possible the the Chinese internet has better quality content available? My problem with deep research tends to be that what it does is it searches the internet, and most of the stuff it turns up is the half baked garbage that gets repeated on every topic.
> is it possible the the Chinese internet has better quality content available? That’s a huge leap of logic. The simpler explanation is that it has better searching functionality and performance. The models are multi-lingual and can parse results from global websites just fine.
I think existence of Wikipedia is a red herring, there's no historical inevitability that people will band together to curate a high-quality encyclopedia on every imaginable topic.
There might be similar, even broader/better efforts on the Chinese internet we (I) know nothing about.
It also might be that Chinese search engines are better than Google at finding high quality data.
But I reiterate - these search based LLMs kinda suck in the West, because Google kinda sucks. Every use of deep research usually ended up with the model citing the same crap articles and data you could find on Google manually, but whereas I could tell the data was no good, AI took it at face value.
One thing I’m becoming curious about with these models are the token counts to achieve these results - things like “better reasoning” and “more tool usage” aren’t “model improvements” in what I think would be understood as the colloquial sense, they’re techniques for using the model more to better steer the model, and are closer to “spend more to get more” than “get more for less.” They’re still valuable, but they op…
i'm no expert, and i actually asked google gemini a similar question yesterday - "how much more energy is consumed by running every query through Gemini AI versus traditional search?" turns out that the AI result is actually on par, if not more efficient (power wise) than traditional search. I think it said its the equivalent power of watching 5 seconds of TV per search. I also asked perplexity to give a report of th…
Can't wait for the benchmark at artificial analysis. Qwen team doesn't seem to have updated the information about this new model yet https://chat.qwen.ai/settings/model . I tried getting an api key from alibabacloud, but the amount of steps from creating an account made me stop, it was too much. It should be this difficult. Incredible work anyways!
https://boutell.dev/misc/qwen3-max-pelican.svg
I used Simon Willison's usual prompt.
It thought for over 2 minutes (free account). The commentary was even more glowing than the image.
It has a certain charm.
Earlier quoted context omitted.
Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.
It would be trivial to detect such gaming, tho. That's the beauty of the test, and that's why they're probably not doing it. If a model draws "perfect" (whatever that means) pelicans on a bike, you start testing for owls riding a lawnmower, or crows riding a unicycle, or x _verb_ on y ...
Earlier quoted context omitted.
What would a good coding model to run on an M3 Pro (18GB) to get Codex like workflow and quality? Essentially, I am running out quick when using Codex-High on VSCode on the $20 ChatGPT plan and looking for cheaper / free alternatives (even if a little slower, but same quality). Any pointers?
18gb RAM it is a bit tight with 32gb RAM: qwen3-coder and glm 4.7 flash are both impressive 30b parameter models not on the level of gpt 5.2 codex but small enough to run locally (w/ 32gb RAM 4bit quantized) and quite capable but it is just a matter of time I think until we get quite capable coding models that will be able to run with less RAM
Current test version runs in 8GB @ 60tks. Lmk if you want to join our early tester group!
Earlier quoted context omitted.
The version backed by photographic and video evidence, I imagine. I haven't looked it up personally. What are the different versions, and which would you expect to see in the results?
It's all about framing. Photos and videos will be recontextualized to show that President Trump, saviour of America, did all he could to encourage peaceful protest on January 6th. Other versions will be created to show that he's not the savior of America, and that he was actually instigating violence on that day.
There were other indications, to be sure, but that was certainly one of them.
Earlier quoted context omitted.
Because corporate aint gonna approve the vendor that lets their dumbass employees generate child porn when the other vendors do not.
This still doesn’t make any sense. If they are worried a dumbass employee would generate CP using the tool, wouldn’t that same employee also download it from the web? Or use a web based tool to generate it? Any employee that is going to use a corporate AI tool to generate CP is going to use other corporate tools to do worse things. There is no point in worrying about it.
Company A - First to the market, Reasonable Cost, Most well known name. Very easy integration.
Company B - Well regarded tools, Higher cost, Better performance and reviews from team. More difficult to integrate
Company C - Reasonably Priced, Performance is reasonable, Has a connection to an extremely controversial individual, Currently being lambasted for being an CP/Revenge Porn generator
Ok, now pretend you're talking to a guy who signs your paycheques. Which one are you NOT gonna pick?
Earlier quoted context omitted.
There’s a domestic AI price war in China, plus pricing in mainland China benefits from lower cost structures and very substantial government support e.g., local compute power vouchers and subsidies designed to make AI infrastructure cheaper for domestic businesses and widespread adoption. https://www.notebookcheck.net/China-expands-AI-subsidies-wit...
All of this is true and credit assignment is hard, but the brutal competition between Chinese firms, especially in manufacturing, differentiates them from and advances them over economies in the west. It makes investment hard as profits are competed away, which is blasphemy in Thiel's worldview, but is excellent for consumers both local and global.