While I agree with some of the gist of the article, 2 remarks: 1. Unfortunatly in my tests the open models do not (yet?) rival, at least Claude Opus, for software development/engineering and adjacent tasks. 2. Enjoy while it lasts. I'll be genuinly amazed these open models will not be declared 'illegal' under some security pretense by the end of the year. And I say 'pretense' because the primary driver will be regula…
There is minimal downside to switching to open models
141–150 of 351 posts
Re: There is minimal downside to switching to open models
#142Earlier quoted context omitted.
I don't think we will. The open model labs are too resource constrained to approach Fable or even Opus on the general case and I don't see that changing within a year. Right now, due to profound shortfalls in both data and hardware compared to the US labs, the OSS models are IMO basically technology demonstrators that in practise are even more jagged than the US labs' efforts. The high points of the jaggedness are cl…
deepseek-v4-pro, probably the representative cheap opensouce LLM, was released in 2026.4 One year before, what OAI had in hand was gpt-4.1 and gpt-o3. I think it is not very controversial to say that deepseek is stronger than them, at most you can point to some post-training problems, basically the instability you mentioned. Also I am not sure if it is because the people who are best at using AI -- the people making…
But the question was about whether the Chinese labs will have fable-equivalence in 1 year. I am by no means some kind of insider, but knowing the vaguest outlines of what went into Mythos, they just can't do it. The compute is not there. The Chinese engineers are incredible, but they're not literal magicians.
Of course there could be something incredible to come out of left field and overturn the apple cart yet again, but that's speculation. It would be awesome, sure! But I wouldn't bet too heavily on it.
And FWIW - again, no disrespect at all to the Chinese engineers but I don't rate GLM5.2 as being even close to opus 4.6. It can hit a few benchmarks, sure, that's the top edge of the "jag". But filling in the rest of the capabilities - again, it takes compute and data the OSS labs just don't have, that anyone knows about at least.
Re: There is minimal downside to switching to open models
#143Earlier quoted context omitted.
They would quantize the model. That'd make it cheaper to run, and have slightly worse output but it would still generate outputs with a similar feel, derived from a compressed version of the same knowledge base etc. They wouldn't even need to do this uniformly, quantized versions of the model could be routed only a subset of the requests. They could do this to nerf the old model, or more likely just to give themselve…
I have had the same experiences you've had with 4.6 and it was ever since they brought out 4.7. It's fairly obvious they're doing something like you've said here.
Re: There is minimal downside to switching to open models
#144I guess this will happen soon. There are two catalysts needed for this to happen: 1. Evals that can quickly tell you how much downside there is to switching 2. Something like OpenRouter that can help you run those evals quickly Now #2 is starting to become popular, and I think we'll soon see more people adopting a model-agnostic approach. Of course, there will still be high-intelligence use cases where nothing comes…
Whether you're using SDK or harness based agents, having evals means you're able to modify any part of your agent and still know what satisfies your "good enough".
It's great for designing products that are easy to change as well.
Re: There is minimal downside to switching to open models
#145Earlier quoted context omitted.
There's at least the possibility that they intentionally degrade the models as time passes. We can't really verify that we're getting what we're paying for all of the time. All the more reason to invest in local inference.
What if the new model is exactly as good as the last model on launch day but better than the last model was on the new model's launch day because it was degraded? Every single time?
Re: There is minimal downside to switching to open models
#146I’ve been wanting to get better acquainted with local inference but I don’t have the hardware, which has made me think about something I haven’t seen discussed, which is local collaboratives. The economics makes it seem like a group of people joining together to run good hardware and an open model might make sense, but I haven’t seen anything like this mentioned. Have I been missing it? I think it would be pretty nea…
The reason you don't see more of this is because everyone does the math, realizes it's not a good deal, and then gives up on the idea. There's a post at the top of /r/localllama about this exact math right now: https://www.reddit.com/r/LocalLLaMA/comments/1ubrcwj/tokenom... TL;DR: Running GLM 5.2 is going to cost about $20K minimum, and that's going to be painfully slow compared to the cloud hosted versions. Even the…
The appeal to me is that we can run that, but we can also run smaller models on your laptop _and it’s functional!_ I can run DeepSeek v4 flash and a qwen 3.6 on my laptop! Thats crazy good.
Re: There is minimal downside to switching to open models
#147Earlier quoted context omitted.
> With all the issues in the US and generally wrong direction, I can’t remember them ever arresting people for mean tweets in the way that Germany and the UK have. Then you haven't been paying attention. The constitution prevents citizens from being convicted, but that doesn't stop arrests or being turned away at the border (even for permanent residents who've lived in the US for decades), and US citizens don't seem…
>and US citizens don't seem to care I think maybe you haven't been paying attention. Most of us do care. Trump's approval rating is pretty low at 36%, and his disapproval rating is high. Just because he's still causing chaos doesn't mean the majority of us don't care about it. There's just no legal way to remove him, and his cronies simply won't do it - there's not enough votes in congress or he would have been gone…
By contrast, Biden at the same point in his term was hovering around 39%, for the heinous crime of... rebuilding the US economy? Including some woke riders in his infrastructure bill?
At this point, a fair assessment of US citizens is that on average, they seem to consider that being a right-wing autocrat wannabe, threatening to invade allied countries "as a negotiating tactic", being a climate change denier, starting a humiliating failed war, trying to blackmail the press into compliance, etc, are about 3% worse than being a cringe center-left bureaucrat.
"US citizens don't seem to care" is an apt hyperbole.
Re: There is minimal downside to switching to open models
#148> Open models are served via various means, some by the companies that released them and some by third parties like OpenRouter. Unfortunately, both of these routes are dodgier in terms of privacy and data sharing, and I would not feel the same comfort sending API calls containing client or confidential data to them. That's why I'm using eurouter.ai with the following routing rule for all my requests: { "model": "glm-…
I had a look at eurouter.ai and it seems like an extremely bad offer. - The prices are ridiculous (15 % markup for free account). - They have a rate limit of 1000 requests per month, unless you pay 40€ per month for ... what exactly is their value proposition? - They have a single provider (TensorX) for DeepSeek-V4-Pro, with a cache read cost that is over 100 times higher than DeepSeek ($0.44 vs $0.003625). Notably,…
Re: There is minimal downside to switching to open models
#149I think it's interesting that people write off open weight models because they're "a few months behind" proprietary models. I know LLMs move at the speed of light (especially these past few quarters), but if Opus and GPT "a few months ago" were really like open weight models, then there's really no reason to not switch, especially for those who were using these models a few months ago. Your codebase didn't change, so…
Re: There is minimal downside to switching to open models
#150Earlier quoted context omitted.
People talk about this a lot. What I have never seen is a discussion of methods they might employ to degrade the models. Let’s say I’m a bad faith LLM operator, and I want to degrade my model so the next release looks better and people want to switch to the more expensive one. How would I do that?
Weight quantization, n-expert capping, routing to smaller model, context window truncation, aggressive sampling constraints, lossy speculative decoding and probably more.