Show HN: State of the Art of Coding Models, According to Hacker News Commenters
21–30 of 96 posts
Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#22Interpreting these metrics is quite interesting. One thing for sure is that while Claude is currently taking the #1 spot in mentions, it carries a lot of negative sentiment due to API pricing policies and frequent server downtime. On the other hand, the runner-up, GPT-5.5, actually seems to have more positive feedback. Personally, my experience with Codex wasn't as good as with Claude Code (Codex freezes on Windows m…
They are cheaper! All signals point to them staying cheaper because they are built more sustainably. Also, some of the latest entries can run on 1 GPU! Literally available at your desktop where there can be no service interruptions. Not even network latency. People are one and few shotting little games for 0 dollars because they bought a GPU to play video games this year. To me that's an unbeatable value. Once the tooling catches up and a few more model releases, it could change everything completely.
Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#23Interpreting these metrics is quite interesting. One thing for sure is that while Claude is currently taking the #1 spot in mentions, it carries a lot of negative sentiment due to API pricing policies and frequent server downtime. On the other hand, the runner-up, GPT-5.5, actually seems to have more positive feedback. Personally, my experience with Codex wasn't as good as with Claude Code (Codex freezes on Windows m…
Of course, when I tried it on something else it rewrote every line in the file for no good reason, applied changes directly when I told it just to plan, etc.
So maybe it has one strength.
Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#24Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#25Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#26Before harnesses, I'd fix the methodology/claims. A saner methodology would be to see comments that compare two models, say 'gpt5.5>opus4.7' and infer context ('ctx:frontend', for example). For your current methodology, 'opus 4.6 was very smart, opus4.7 is a disappointing upgrade to 4.6' would make normal aspect-based sentiment analysis consider 4.6 is smarter than 4.6. But considering you have <300 mentions total, p…
The context would be really nice to have, but reading the comments myself, it often just isn't very clear what exactly users are building or which programming language they are using.
I think analyzing more comments is promising. If you get enough data, you can generalize across use cases and get more meaningful ratings. The obvious lever is including more posts, although it might hit diminishing returns. I'll play around with it.
For the context, I want to try giving Gemini a "scratch pad", where it can note down strengths and weaknesses per model that it finds in the comments. Something like "some users say that model x is good for writing tests". Then on each run, I let it update the scratch pad and publish the results as more of a qualitative analysis.
For the wording, I'd like to keep a certain amount of click bait, sorry ;)
Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#27Interesting to see the positive sentiment around kimi2.6 qwen3.6 and deepseek relative to the negative. I hope the trend of people appreciating open models continue. They aren't namesakes yet, but it's a higher percentage then I thought it would be. Especially on HN where we are all talking about businesses. I am upset because now anthropic, openai, meta, etc will continue their smear campaigns here. But I am also ha…
What I want is more fully open models where everything is shared. Data, training algorithms, weights. That way we can figure out if we should trust it.
Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#28Surely "Claude Opus 4.7" and "Claude Opus Latest" should be the same, right?
I thought I'd keep these as a rating for model families rather than specific models. But tbh it's probably better to remove them, too confusing.
Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#29How noisy is the sentiment classification? Feels like that could skew results a lot
Re: Show HN: State of the Art of Coding Models, According to Hacker News Commenters
#30Interesting to see the positive sentiment around kimi2.6 qwen3.6 and deepseek relative to the negative. I hope the trend of people appreciating open models continue. They aren't namesakes yet, but it's a higher percentage then I thought it would be. Especially on HN where we are all talking about businesses. I am upset because now anthropic, openai, meta, etc will continue their smear campaigns here. But I am also ha…
Is it just “smear campaigns”? Don’t get me wrong - I don’t want big tech or big AI monopolies and appreciate the open weight models. But it’s also true that Chinese companies are basically stealing through distillation and also that they censor to align to CCP rules. They’re problematic in a different way. What I want is more fully open models where everything is shared. Data, training algorithms, weights. That way w…
I think it's also unfair to say their success is solely due to stealing data. They are contributing a lot of advances to the literature about what they are doing. The proof is in the results we have 27b models you can vibe code with. Not 1t+
It's murky sure. But there are smear campaigns about how people can't trust China too. There's some truth to that too but we can't trust the US either so local models are an interesting way for China to offer us some level of sovereignty.