Live data from Hacker News

$500 GPU outperforms Claude Sonnet on coding benchmarks

github.com

221–230 of 311 posts

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#221
post #215

Earlier quoted context omitted.

Generating big chunks code is all I do, all day. I don't write code by hand any more, neither at work, nor for side projects. I work mostly in Rust and TypeScript at a developer tools company.

[flagged]

Why? Because writing code is the only measure of quality when producing tools? What about Unit and Integration Tests, UX research, and Performance tests.

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#222

Earlier quoted context omitted.

In the US/Western Europe? Because for devs especially in the former, $200 is pocket change, especially for a core productivity tool. And the rent would be in the $1200 to $3000 easily. Same for houses. Maybe not in NY or SF, but in most of the US there's no shortage of house spaces and redundant rooms.

I've seen those comments about $200/month and empty rooms here, so I suppose they mainly come from the US, yes. So yes, you describe a situation that I feel like a lot of people here don't understand is not the norm. I compared the subscription with my rent precisely because it's easier to compare: with your numbers it would be like paying from $600 up to $1500 / month. Pretty hard to justify.

> Because for devs especially

Are you not a dev? If not, what would you use a coding tool for? They still require handholding for anything largeish. Still much cheaper than outsource.

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#223
post #215

Earlier quoted context omitted.

Generating big chunks code is all I do, all day. I don't write code by hand any more, neither at work, nor for side projects. I work mostly in Rust and TypeScript at a developer tools company.

[flagged]

We have the quietest on-call rotation of any company I've ever worked at.

We have a high standard for code review, static verification, and tests.

The fact that the code isn't hand-rolled artisanal code, and is generated by AI now, has so far turned out to have no impact on product quality or bugs reported.

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#224
post #10

It's a race to the bottom. DeepSeek beats all others (single-shot), and it is ~50% cheaper than the cost of local electricity only. > DeepSeek V3.2 Reasoning 86.2% ~$0.002 API, single-shot > ATLAS V3 (pass@1-v(k=3)) 74.6% ~$0.004 Local electricity only, best-of-3 + repair pipeline

I will "suffer" through .004 of electricity if I can run it on my own computer

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#225
post #148

Earlier quoted context omitted.

My usage is in the $60 tier, but that doesn't exist so I have to cough up $100. And then get all shaky if I don't use up my weekly quota.

Do you mostly just hit the session limits? If so I know it's not ideal but you could wait an hour or two for that to reset. Not sure if that would work for you but just a suggestion

I get to 80% when on a single session and cap out a hour off the rest if I’m working on two.

But I like to have that forced hour to stop, it’s moment to take a breath.

It depends on the kind of work though, some things are more token intensive.

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#226

Earlier quoted context omitted.

I hope so too, but I think it's wishful thinking. Be prepared for the mother of all financial bailouts from the world governments to make sure that doesn't happen

I can understand why banks got bailed out by the US gov in 2008, but why would a government feel the need to bail out AI labs? I hope you are not going to say, "to avoid a global recession or depression caused by the popping of the AI bubble". That would be unnecessary and harmful (in its second-order effects), and governments do have advisors who are competent enough in economics to advise against such a move.

The U.S. has an admin right now that has made it clear the only important metric for country health is the stock market, which is single-handedly propped up by AI right now.

That's why huge concessions nobody asked for were made to the AI industry in the Big Beautiful Bill.

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#227
post #14

Earlier quoted context omitted.

> cheaper than the cost of local electricity only. Can you explain what that means?

I think they mean that the DeepSeek API charges are less than it would cost for the electricity to run a local model. Local model enthusiasts often assume that running locally is more energy efficient than running in a data center, but fail to take the economies of scale into account.

Is it economies of scale, or is it unpaid externalities?

Re: $500 GPU outperforms Claude Sonnet on coding benchmarks

#230

Earlier quoted context omitted.

Maybe you haven't tried any other AI product with an actual preexisting project. Or blindly trust every BS Claude feeds you. I haven't had any such frustrations with Codex Claude is specially annoying because of their submarining and people thinking it's the best

I use both, read what I need to read and fix small issues myself. Both Agents are pure magic and none of their issues warrant a tantrum on a public forum.

I posted a more detailed report in case you can't see it in your thread view: https://news.ycombinator.com/item?id=47541369

and other comments further back in my history

> none of their issues warrant a tantrum on a public forum

I don't get frustrated if a problem is genuinely difficult to solve and the product creator is trying their best,

I get frustrated when a problem has been solved by other similar products but a specific creator or provider refuses to follow suit and fix their shit.

Claude's Electron app vs. Codex's native app is one such example right off the first impression of both products.

Post reply on HN