Earlier quoted context omitted.
All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…
It takes 2 minutes to fix Opus 5 https://code.claude.com/docs/en/output-styles
Qwen3.8 Max now ranked as the best overall model by agentic index
341–350 of 364 posts
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#342Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…
Hint: The new models are really good at burning tokens. I've had to use it a bit for work, and it's been remarkable watching the degradation in performance with the default suggested current models (Opus 5 as a prime example) vs the models that got them huge attention a year ago (Opus 4.6) If you give 4.6 a spec, or existing code to implement a feature in, it will ask some pointed questions if there's something uncle…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#343Earlier quoted context omitted.
Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.
This reply must have cost dozens of dollars.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#344Earlier quoted context omitted.
Copyright law only considers illegal ownership of a work, so the crime - or tort - was making/acquiring copies without permission or payment. Training from copies has been ruled fair use because it's "transformative" and not simply "derivative." This is obviously debatable, but that's where the debate is at the moment.
> but that's where the debate is at the moment. Because of the rulings of a couple of judges. Is that actually what the majority of people think? > Copyright law only considers illegal ownership of a work That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing. Similarly, copyright has something to say if I read a legal copy of harry potter and then c…
There's a good reason for the law not to be based on what the majority thinks.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#345Earlier quoted context omitted.
> They didn’t accept guilt for incorporating all of human output into their model without consent. Because that use case is actually permitted by law.
I mean... that's one interpretation of the law, sure. The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage. So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.
It's not an either-or.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#346Earlier quoted context omitted.
They have something like 80% gross margins, are at a $100B/yr ARR, and are growing at 10x per year... If that keeps up, they're going to be doing more revenue than Google in a year ($400B ARR, 20% per year growth)
Where did that $100B figure come from? I thought they were at ~10B at the end of 2025, so they're either not at 100B yet, or they're growing way faster than 10x / year.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#347Earlier quoted context omitted.
I'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.
Why isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#348Earlier quoted context omitted.
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
I think it's more likely to be the effect of synchronization of launches, and the fact that models that do not challenge SOTA in some way do not get launched (think Gemini Pro delays), launched quietly or do not get any attention.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#349Earlier quoted context omitted.
> but that's where the debate is at the moment. Because of the rulings of a couple of judges. Is that actually what the majority of people think? > Copyright law only considers illegal ownership of a work That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing. Similarly, copyright has something to say if I read a legal copy of harry potter and then c…
> Because of the rulings of a couple of judges. Is that actually what the majority of people think? There's a good reason for the law not to be based on what the majority thinks.
Sure i don’t think the majority get to dictate things like who has rights or who the law applies to. That doesn’t apply here tho.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#350Earlier quoted context omitted.
I mean... that's one interpretation of the law, sure. The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage. So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.
> So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework. It's not an either-or.