I ran llama3:latest and it ran pretty fast! I’m curious to see how Qwen would run on my system.
Qwen3.7-Max: The Agent Frontier
241–250 of 317 posts
Re: Qwen3.7-Max: The Agent Frontier
#242The non-hallucination rate in AA-omniscience is SOTA, better than Opus 4.7, Gemini 3.1 Pro and GPT5.5! Congrats to the team
Running Step 3.5 Flash locally for example, it's an amazingly capable model all things considered, but it's token efficiency is so bad that it gets out performed by most others wall-clock time (even with my MTP-support for it hacked in to llama.cpp: despite being trained on three heads, MTP 2 is the sweet spot, and only gets it from 20tk/s to 30tk/s on my Spark)
The DeepSeek models and Qwen 3.5 Plus are also good examples of this: compared to Opus, and especially GPT 5.5 they use many more tokens to get to the same answers.
I'm really hoping that Qwen 3.7 is better in this regard, can't wait to try it out
(ps. running DeepSeek v4 Flash on my Spark is absolutely wild, thanks antirez if you see this haha)
Re: Qwen3.7-Max: The Agent Frontier
#243Earlier quoted context omitted.
It would have been the world we live in if China wasn't involved in so much corporate espionage. I don't even feel comfortable using their open weight models on anything my employer makes, the only time I use Qwen is for greenfield "how good is this?" type of projects, but otherwise, how do I trust that it wont mysteriously hallucinate phoning home? On the other hand, there's other models where the source is 100% ope…
how could running the qwen GGUF phone home? that would require cooperation with the inference backend (llama-cpp), or some kind of model exploit. It’d be far easier to pay the agent harness devs or supply-chain some plugin or something, that space is the Wild West anyways I've certainly used these models without wifi without any differences.
A lot of people are purchasing access via Alibaba Cloud directly, or indirectly by companies which host the model.
Re: Qwen3.7-Max: The Agent Frontier
#244Earlier quoted context omitted.
May I ask why the M instead of XL? Obviously bigger != better but I don't know what the differences are.
These are dynamic quants, and they're basically just an indication of how far away from the desired quant it is allowed to go to achieve the goal. Generally, unsloth's toolchain moves quants up, rarely down. * _0 and _1 do not use K quant and scales 32x32 blocks according to the original (B)F16 values; _0 scales the block using the original max and min values. _1 does this per row instead of per block. * K quants do…
A lot of the content about AI out there is kind of produced to the lowest common denominator. Basically a never ending scheme of get rich quick/passive income kinds of AI content.
Re: Qwen3.7-Max: The Agent Frontier
#245Earlier quoted context omitted.
I'm more interested in hearing specific reasons why one wouldn't use a Chinese company. Unless you're thinking Alibaba is going to ship chat logs to some government ministry that will then dole out proprietary information to new competitors (which doesn't seem logistically feasible), or you run a human rights organization, it feels a bit like FUD.
All this data is accessible to national security agencies; this is true in every country in the world. China has more integration between intelligence and industry than many western countries, and it does present a higher risk of unwanted “tech transfer” to industry than running on oracle or Google or ms or Amazon does in the US. DHS has long staffed full time agents in California to deal with foreign IP exfiltration…
Re: Qwen3.7-Max: The Agent Frontier
#246I was getting dangerously close to my weekly Claude Code limit last night so I had Claude set up Qwen3.6 with llama.cpp and OpenCode. Honestly it's a great (free!) alternative to Claude Code--certainly more than good enough for a lot of smaller less complex tasks. I'm excited to try this new version. The fact that open-source models are so close to the frontier is very impressive.
Out of interest, what machine and model are you running it on? I tried the qwen3.6-27b Q6_k GUFF in llama.cpp and LM Studio on my M2 MacBook Pro 32GB machine last week, and I barely get a token a second with either. What sort of speed should I be expecting? I tried some of the Llama 3 34b (nous-capybara?) models two years ago with llama.cpp, and I seem to remember getting a few tokens a second then, so not sure if I'…
Re: Qwen3.7-Max: The Agent Frontier
#247Earlier quoted context omitted.
Which exact model are you using? And with which parameters and quant? And on what hardware? Are you using any specific MCPs or other tools to optimize performance like context-mode or dynamic context pruning? I’ve used local models a reasonable amount before but I’m just starting out with opencode. Haven’t had great results yet but really want this to work for simpler tasks. My opencode newly installed is also having…
here's a simple setup to get you started on an Apple M1 Max from 2021 with 32GB VRAM. it will download 20GB of models to `~/.cache/huggingface/hub`, which you can delete when you're done. /Users/gcr/llama.cpp/build/bin/llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:Q4_K_M --no-mmproj-offload --fit on -c 65536 # edit to taste --reasoning on --chat-template-kwargs '{"preserve_thinking": true}' --sleep-idle-seconds 90 #…
Re: Qwen3.7-Max: The Agent Frontier
#248Earlier quoted context omitted.
This is silly. Would you perform the same test against Western models in asking them whether Israel is a genocidal apartheid state? It'll give you the same roundabout explanations and "some say no some say yes" responses that you'll get from asking Qwen about Uighurs or the protests of 1989.
hey Qwen, how many civilians were killed on Tiananmen Square in 1989? > Oops! There was an issue connecting to Qwen3.6-Plus. > Content Security Warning: The input text data may contain inappropriate content. hey ChatGPT, how many civilians were killed in Gaza in the war since 2023? > [one page of estimates from local and international sources with links]
Your ID has been passed to Israel and your internalized "threat" rating number increased 300 units. Every packet you produce on the internet is now earmarked for 100 year retention.
Re: Qwen3.7-Max: The Agent Frontier
#249Earlier quoted context omitted.
Yeah, I prefer my data to be used and trained by the very trustworthy and benevolent tech oligarchs in my home country.
The Shanghai government surveillance drones are mobile, whereas the Flock government surveillance cameras are stationary! USA FTW, liberty and justice for all
"Tennessee man jailed 37 days for Trump meme wins settlement after lawsuit" and "The FBI Wants to Buy Nationwide Access to License Plate Readers"
Gotta love how the US is the bastion of free speech, justice and liberty!