Live data from Hacker News

Viewing profile — stared

stared

HN member
Joined
Sun, Sep 02, 2012, 3:25 PM UTC
HN karma
13,516
Public activity
3,220 items

About stared

A curious being, doctor of sorcery. Posts, projects & resume at: https://p.migdal.pl/

Now: benchmarking AI at: https://quesma.com/blog/

Previously: co-founder & CTO at https://quantumflytrap.com/

Recent public activity

  1. story
  2. comment
    Comment #49190303

    Good to know! Is it that it wasn't accepted yet, or are there issues with how it was run?

  3. comment
    Comment #49190228

    It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010 . I am curious - how does it fare for other benchmarks, or everyday p…

  4. comment
    Comment #49158654

    I am curious how Claude Opus 5 fares - similar, better, or (my guess) worse than Fable 5.

  5. story
  6. comment
    Comment #49146604

    We updated our entry with DeepSeek-V4-Flash-0731, which rocks when it comes to intelligence per $. More on the benchmark: https://quesma.com/blog/baba-is-bench/

  7. story
  8. comment
    Comment #49133425

    "A Visualization Language for the AI Era", as a tagline, sounds weird. I had the best success with popular and versatile packages like matplotlib and ggplot2 - even 1.5 ago (vide h…

  9. comment
    Comment #49132512

    Apps for one, but when I think there is even a mild potential others may use it, I share. * Doom launcher and Megawad Steam-like library for macOS: https://github.com/stared/rusted…

  10. comment
    Comment #49128383

    Call me biased, but if Grok 4.5 is above GPT-5.6 Sol, I don’t trust this benchmark.

  11. comment
    Comment #49125595

    For an update on Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash, see: https://quesma.com/blog/baba-kimi-k3-opus-5/ This post was directly mentioned in an OpenAI release on the ARC…

  12. story
  13. story
  14. story
  15. comment
    Comment #49095128

    Now the friction is close to none, at least for techies. Set a GitHub repo, use pnpm + Astro, and your favourite AI to migrate data there, brew some coffee, and before you finish i…

  16. story
  17. story
  18. comment
    Comment #49045399

    It's called frog boiling. We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age. If you don't believe me, create something comp…

  19. comment
    Comment #49045336

    Also top on the freshly released Frontier-Bench, by a large margin: https://www.frontierbench.ai/

  20. story
  21. comment
    Comment #48977093

    On another note - who needs WordPress in the age of Astro and LLMs? I mean, it's not a taunt, but a serious question - do people keep WordPress because it used to be the easiest so…

  22. comment
    Comment #48957470

    See the "caveats" section. It was also our initial assumption, but to our surprise there were no signs of models "knowing" the solutions (we investigated all trajectories). Compare…

  23. comment
    Comment #48957244

    It is on the way! Sadly, this Kimi K3 API rate limits are devastating, at least on the OpenRouter. So far it solved all The Intro levels (as expected) - I am much more curious how …

  24. comment
    Comment #48957215

    I am almost certain it is star or spark. But again, if someone is looking for assholes, they will find them.

  25. comment
    Comment #48957197

    In a similar vein, some time ago I got curious why the Grafana logo looks like the Zerg emblem see https://www.reddit.com/r/grafana/comments/1o79zxy/grafana_lo... .