Viewing profile — KaoruAoiShiho
KaoruAoiShiho
HN member- Joined
- Sat, Oct 01, 2011, 1:49 AM UTC
- HN karma
- 3,254
- Public activity
- 1,695 items
- HN profile
- View on Hacker News ↗
About KaoruAoiShiho
Recent public activity
-
comment
Comment #49047099
Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
-
comment
Comment #48874716
You're probably just responding to the headline but this person is an AI bull and isn't claiming it's a big deal, she's going into it and explaining it.
-
comment
Comment #48834756
Click through to the link, the answer is no it uses the latest gpt models now.
-
comment
Comment #48748657
I think the parent's point is that if you are genuinely open to losing, the arguments can be productive because you can learn something instead... So stopping arguments is just ano…
-
comment
Comment #48712475
And before you know it, you invented some openrouter provider from first principles...
-
comment
Comment #48712466
Are you sure fireworks is unquant? It's not listing precision on openrouter like everyone else.
-
comment
Comment #48594196
Terrible zero value article, I am extremely surprised it is upvoted. That being said Artificial Analysis just came out with a brand new benchmark where it scored between opus 4.8 a…
-
comment
Comment #48570245
This is really held back by one bench (omniscience accuracy) where it's really very far behind otherwise i think it's got at least a couple of points higher.
-
comment
Comment #48560401
Fable largely fixed the annoying chatterness so sucks that it's gone now.
-
comment
Comment #48200804
Then that's just a video game might as well as play a video game why limit yourself to still confined to the rules of chess.
-
comment
Comment #48192392
Kimi 2.5 has the best long context. For raw coding benchmark scores you can just post train on top of it with more specialized data. 2.5 is kinda old, 2.6 is the current release wh…
-
comment
Comment #48188625
Do people like "gacha"? I thought people played games for the game experience, story, etc, and the gacha is just the monetization mechanic. It's like making a big deal out of payin…
-
comment
Comment #47961610
No I think the best agent with hundreds of millions in ARR should be worth more than the 15th best model company with tiny revenue. ur the joke.
-
comment
Comment #47926171
Manus is saved, 2 billion is such an undervaluation considering much worse companies like minimax is valued at 30 billion.
-
comment
Comment #47885333
SOTA MRCR (or would've been a few hours earlier... beaten by 5.5), I've long thought of this as the most important non-agentic benchmark, so this is especially impressive. Beats Op…
-
comment
Comment #47836912
Huh, that's not a thing?
-
comment
Comment #47795284
Might be sticking with 4.6 it's only been 20 minutes of using 4.7 and there are annoyances I didn't face with 4.6 what the heck. Huge downgrade on MRCR too.... 256K: - Opus 4.6: 91…
-
comment
Comment #47743831
Talking nonsense.
-
comment
Comment #47742128
Well they're not public yet so you'll have to put up with rumors. But the numbers are available for companies like DeepSeek say they have an 80% profit margin, so it stands to reas…
-
comment
Comment #47739935
After googling https://www.reddit.com/r/singularity/comments/1psesym/openai...
-
comment
Comment #47719290
TLDR: Writer hasn't heard of agents.
-
comment
Comment #47710103
I feel like netflix is definitely very cheap, with OpenClaw or whatever your favorite agent is, it's trivial to subscribe to watch one show and then have it cancel immediately.
-
comment
Comment #47679366
Blog post is new but the model is about 2 weeks in public.
-
comment
Comment #47679355
The non-awesome context window is the sad part, but I think a better harness can deal with this.
-
comment
Comment #47665480
Can you paste the relevant section in your soul please?