Viewing profile — stared
stared
HN member- Joined
- Sun, Sep 02, 2012, 3:25 PM UTC
- HN karma
- 13,516
- Public activity
- 3,220 items
- HN profile
- View on Hacker News ↗
About stared
Now: benchmarking AI at: https://quesma.com/blog/
Previously: co-founder & CTO at https://quantumflytrap.com/
Recent public activity
- story
-
comment
Comment #49190303
Good to know! Is it that it wasn't accepted yet, or are there issues with how it was run?
-
comment
Comment #49190228
It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010 . I am curious - how does it fare for other benchmarks, or everyday p…
-
comment
Comment #49158654
I am curious how Claude Opus 5 fares - similar, better, or (my guess) worse than Fable 5.
- story
-
comment
Comment #49146604
We updated our entry with DeepSeek-V4-Flash-0731, which rocks when it comes to intelligence per $. More on the benchmark: https://quesma.com/blog/baba-is-bench/
- story
-
comment
Comment #49133425
"A Visualization Language for the AI Era", as a tagline, sounds weird. I had the best success with popular and versatile packages like matplotlib and ggplot2 - even 1.5 ago (vide h…
-
comment
Comment #49132512
Apps for one, but when I think there is even a mild potential others may use it, I share. * Doom launcher and Megawad Steam-like library for macOS: https://github.com/stared/rusted…
-
comment
Comment #49128383
Call me biased, but if Grok 4.5 is above GPT-5.6 Sol, I don’t trust this benchmark.
-
comment
Comment #49125595
For an update on Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash, see: https://quesma.com/blog/baba-kimi-k3-opus-5/ This post was directly mentioned in an OpenAI release on the ARC…
- story
- story
- story
-
comment
Comment #49095128
Now the friction is close to none, at least for techies. Set a GitHub repo, use pnpm + Astro, and your favourite AI to migrate data there, brew some coffee, and before you finish i…
- story
- story
-
comment
Comment #49045399
It's called frog boiling. We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age. If you don't believe me, create something comp…
-
comment
Comment #49045336
Also top on the freshly released Frontier-Bench, by a large margin: https://www.frontierbench.ai/
- story
-
comment
Comment #48977093
On another note - who needs WordPress in the age of Astro and LLMs? I mean, it's not a taunt, but a serious question - do people keep WordPress because it used to be the easiest so…
-
comment
Comment #48957470
See the "caveats" section. It was also our initial assumption, but to our surprise there were no signs of models "knowing" the solutions (we investigated all trajectories). Compare…
-
comment
Comment #48957244
It is on the way! Sadly, this Kimi K3 API rate limits are devastating, at least on the OpenRouter. So far it solved all The Intro levels (as expected) - I am much more curious how …
-
comment
Comment #48957215
I am almost certain it is star or spark. But again, if someone is looking for assholes, they will find them.
-
comment
Comment #48957197
In a similar vein, some time ago I got curious why the Grafana logo looks like the Zerg emblem see https://www.reddit.com/r/grafana/comments/1o79zxy/grafana_lo... .