Viewing profile — benjiro29
benjiro29
HN member- Joined
- Sun, Jun 14, 2026, 10:53 AM UTC
- HN karma
- 528
- Public activity
- 82 items
- HN profile
- View on Hacker News ↗
About benjiro29
No profile information was provided.
Recent public activity
-
comment
Comment #49217733
HBM requires stacking the chips. So they need to shave the layers, glue, stack more, shave again. They also require a substrate what is even more wafers. The issue is that a error …
-
comment
Comment #49209414
Ironically, we are also moving to more capable / faster models that use less power. DeepSeek V4 Flash 0731 is so extreme capable and comparability to a lot of models cheap to run. …
-
comment
Comment #49202341
Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. What cost the most in API. Input, Cached Input, or Output. There you have your answ…
-
comment
Comment #49190183
A very interesting cost analyze of using AI and without by a software engineering. Interesting quote by OP on reddit: > So I wouldn't say "AI saved me $200k". What actually happene…
- story
-
comment
Comment #49189721
Because every company may have different needs that are not fulfilled by standards software. We have seen the large number of companies whose goal is to make custom software for ot…
-
comment
Comment #49188740
> Claude code does some of this by handing off the "explore" agent work to haiku. That is not handing off to a specialized model, its just handing off to a lighter and interior mod…
-
comment
Comment #49145586
Repost from above: https://en.wikipedia.org/wiki/Film_industry#Largest_markets_ ... ("Largest markets by box office revenue") United States and Canada $8,870,000,000 2025[175] Euro…
-
comment
Comment #49145570
Whoever put those trailer together really knew how to sell it as a "do not bother to see movie". And frankly, Jason Momoa fatigue may also be a issue. Jason momoa playing Jason mom…
- comment
-
comment
Comment #49121594
https://www.tbench.ai/leaderboard/terminal-bench/2.1 > 78.4 The real score is always the official benchmark. We need to see later if DS4 flash 0731 is going to maintain the score b…
-
comment
Comment #49121564
Probably the same. When the same base model is trained, the weight do not tend to change a lot. GLM 5.0 > 5.1 > 5.2 are the same base model, that just kept being trained. Weights h…
-
comment
Comment #49121550
I think it does not matter. Most people who use these types of models are into the IT world, and will know/be informed very fast that there is a difference. People will likely also…
-
comment
Comment #49121506
MiMo, the overlooked sidekick to the hero. Will be interesting to see what Xiaomi bring to the table. These massive jumps in cheap models, is really great times!
-
comment
Comment #49121460
But on Terminal bench, its * DS4 Flash: 82.7 * GPT 5.6 Luna: 75.7 For reference, that puts it on the third spot behind GPT 5.5 and Fable 5. For some reason GPT 5.6 Sol is not showi…
-
comment
Comment #49103383
Seems like it might be more advantageous to just adjust reasoning effort to retain cache. Some agents when you alter the reasoning level, partially or complete wipe the cache. Neve…
-
comment
Comment #49088502
> You Could Have Come Up with ... Creating or combining to have something new, that does not already exist is actually freaking hard! The moment its presented and people go "o, tha…
-
comment
Comment #49068712
Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its a…
-
comment
Comment #49068644
I really do not understand how software developer think anymore. Using a LLM to translate a project in a short time, is by itself incredible. Just like one-shot whatever office clo…
-
comment
Comment #49042078
Same answer i gave to somebody else up here... If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are…
-
comment
Comment #49042018
Then your comparing to a level of GPT 5.6 High, what is 50% cheaper then Opus Medium for the same intelligence / score. You see the issue, if you try to scale effort down, you also…
-
comment
Comment #49041208
> At half the price and less likely to auto-downgrade, it sounds like a reasonable claim Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase fo…
-
comment
Comment #49041158
Also the cost per task. https://www.vals.ai/benchmarks/vals_index !!! Vals !!! Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. …
-
comment
Comment #49021810
Not how it works ... Anything public viewable is still subject to a TOS. And the whole copyright or whatever still applies. If your argument was valid, we can scrape news websites …
-
comment
Comment #49021709
> It was only 15 days. That, and usage was also half. Given the fact that we seen not any evidence beyond people "see, it one shot X game looks similar to Fable" (a lot of one shot…