Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

341–350 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#341

Earlier quoted context omitted.

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

It takes 2 minutes to fix Opus 5 https://code.claude.com/docs/en/output-styles

What output style have you found to actually fix Opus 5's grating prose then? I find it leans hard into its preferred grammatical structures and rote sayings no matter what I include in the output style.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#342

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

Hint: The new models are really good at burning tokens. I've had to use it a bit for work, and it's been remarkable watching the degradation in performance with the default suggested current models (Opus 5 as a prime example) vs the models that got them huge attention a year ago (Opus 4.6) If you give 4.6 a spec, or existing code to implement a feature in, it will ask some pointed questions if there's something uncle…

The entire ecosystem of CC is designed to facilitate burning tokens. You have to ask the LLM to write a script for the app to tell you which folder you're working in and which branch you're on. There are commands that just diagnose your Claude Code setup and try to "optimize" it. Adding skills or plugins bloats the context window. Developing plans means that you work through questions before you get to it in the code, but that matters way more for human programmers than LLMs, so it's probably just a waste of tokens.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#343

Earlier quoted context omitted.

Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.

This reply must have cost dozens of dollars.

Imagine how much it must have cost to have him read your comment! I hope he sends you an invoice.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#344

Earlier quoted context omitted.

Copyright law only considers illegal ownership of a work, so the crime - or tort - was making/acquiring copies without permission or payment. Training from copies has been ruled fair use because it's "transformative" and not simply "derivative." This is obviously debatable, but that's where the debate is at the moment.

> but that's where the debate is at the moment. Because of the rulings of a couple of judges. Is that actually what the majority of people think? > Copyright law only considers illegal ownership of a work That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing. Similarly, copyright has something to say if I read a legal copy of harry potter and then c…

> Because of the rulings of a couple of judges. Is that actually what the majority of people think?

There's a good reason for the law not to be based on what the majority thinks.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#345

Earlier quoted context omitted.

> They didn’t accept guilt for incorporating all of human output into their model without consent. Because that use case is actually permitted by law.

I mean... that's one interpretation of the law, sure. The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage. So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.

> So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.

It's not an either-or.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#346
post #102

Earlier quoted context omitted.

They have something like 80% gross margins, are at a $100B/yr ARR, and are growing at 10x per year... If that keeps up, they're going to be doing more revenue than Google in a year ($400B ARR, 20% per year growth)

Where did that $100B figure come from? I thought they were at ~10B at the end of 2025, so they're either not at 100B yet, or they're growing way faster than 10x / year.

Good question, I heard it on a podcast, but going back to the transcript, looks like that's their forecast, not that they've hit it, they estimated a current $70B, but they've been revising their forecasts up, so yeah, it's probably >10x. Latest solid number they reported was $47B in May.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#347

Earlier quoted context omitted.

I'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.

Why isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?

Because the companies did it underhandedly without prior consent? Imagine you walked into a hardware store to grab some lumber, didn't pay for it, and the store had to call the cops to swing around your house? That's hardly the typical shopping experience, now is it?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#348
post #326

Earlier quoted context omitted.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

I think it's more likely to be the effect of synchronization of launches, and the fact that models that do not challenge SOTA in some way do not get launched (think Gemini Pro delays), launched quietly or do not get any attention.

[deleted]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#349

Earlier quoted context omitted.

> but that's where the debate is at the moment. Because of the rulings of a couple of judges. Is that actually what the majority of people think? > Copyright law only considers illegal ownership of a work That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing. Similarly, copyright has something to say if I read a legal copy of harry potter and then c…

> Because of the rulings of a couple of judges. Is that actually what the majority of people think? There's a good reason for the law not to be based on what the majority thinks.

If you’re gonna equate copyright law being tied to a majority sense of morality, with some kind of mob rule, ya lost me.

Sure i don’t think the majority get to dictate things like who has rights or who the law applies to. That doesn’t apply here tho.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#350

Earlier quoted context omitted.

I mean... that's one interpretation of the law, sure. The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage. So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.

> So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework. It's not an either-or.

Kinda is.
Post reply on HN