Live data from Hacker News

Viewing profile — meander_water

meander_water

HN member
Joined
Sun, Feb 09, 2025, 3:11 AM UTC
HN karma
1,285
Public activity
246 items

About meander_water

I write software and I write words about software:

https://vivis.dev

https://findsubstack.com

https://pythonkoans.substack.com

Recent public activity

  1. comment
  2. story
  3. comment
    Comment #49149219

    This is what I was looking for, thanks!

  4. comment
    Comment #49148772

    Sure, but then you would just use an image generation or multimodal model to generate that image. I don't think you'd want a weird looking svg.

  5. comment
    Comment #49148641

    Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.

  6. comment
    Comment #49047697

    Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than …

  7. comment
    Comment #49046304

    The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because …

  8. story
  9. comment
    Comment #49028976

    1. The benchmark is run with a python script - https://github.com/sunblaze-ucb/exploitgym using an agent harness. I suspect they used codex. So the model has access to the environm…

  10. comment
    Comment #49028256

    Seems similar to Openrouter Fusion - https://openrouter.ai/docs/guides/routing/routers/fusion-rou...

  11. story
  12. story
  13. comment
    Comment #48725766

    I thought all model providers are doing this under the hood anyway in their UI? They certainly seem to when A/B testing different models, and Fable routes to Opus 4.8 when guardrai…

  14. comment
    Comment #48668729

    GPTZero is much better at handling humanized outputs. Also has a similar false positive rate to Pangram.

  15. comment
    Comment #48660300

    > However, it’s your job to go down the rabbit hole, learn the 100%, and sprinkle in your 3%. I would say that there is a big difference between stealing without acknowledgement, a…

  16. comment
    Comment #48628364

    > I don't think you should waste time reviewing every single line of code in here and just use AI to review it! > What you bring is the knowledge that the author nor the LLM doesn'…

  17. comment
    Comment #48627163

    Thanks, I didn't mean to be brusque, but I have seen a lot of these vibe tests lately that come to grand conclusions like "X model is better than Y" from the result of a single pro…

  18. comment
    Comment #48627015

    > So we ran it head-to-head against Claude Opus 4.8: same one-shot prompt, build a 3D platformer in raw WebGL from scratch Running a single one-shot prompt is not a benchmark, not …

  19. story
  20. comment
    Comment #48485287

    Really curious to understand why I'm being downvoted. I don't think it's a particularly spicy take - Just choose the right tool for the job.

  21. comment
    Comment #48483411

    As someone who has built both react based frontends and html based ones (with htmx), there is a law of diminishing returns at play. To start off, writing a basic crud website with …

  22. comment
    Comment #48469073

    All the model releases we've seen this year have only made incremental improvements in benchmarks. This feels like the first release that feels like a significant step up in terms …

  23. story
  24. comment
  25. comment
    Comment #48261932

    Not the first study, and they all largely report the same results: https://www.nature.com/articles/s41562-025-02259-6 https://www.theguardian.com/money/2019/feb/19/four-day-week-..…