Viewing profile — meander_water
meander_water
HN member- Joined
- Sun, Feb 09, 2025, 3:11 AM UTC
- HN karma
- 1,285
- Public activity
- 246 items
- HN profile
- View on Hacker News ↗
About meander_water
https://vivis.dev
https://findsubstack.com
https://pythonkoans.substack.com
Recent public activity
- comment
- story
-
comment
Comment #49149219
This is what I was looking for, thanks!
-
comment
Comment #49148772
Sure, but then you would just use an image generation or multimodal model to generate that image. I don't think you'd want a weird looking svg.
-
comment
Comment #49148641
Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.
-
comment
Comment #49047697
Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than …
-
comment
Comment #49046304
The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because …
- story
-
comment
Comment #49028976
1. The benchmark is run with a python script - https://github.com/sunblaze-ucb/exploitgym using an agent harness. I suspect they used codex. So the model has access to the environm…
-
comment
Comment #49028256
Seems similar to Openrouter Fusion - https://openrouter.ai/docs/guides/routing/routers/fusion-rou...
- story
- story
-
comment
Comment #48725766
I thought all model providers are doing this under the hood anyway in their UI? They certainly seem to when A/B testing different models, and Fable routes to Opus 4.8 when guardrai…
-
comment
Comment #48668729
GPTZero is much better at handling humanized outputs. Also has a similar false positive rate to Pangram.
-
comment
Comment #48660300
> However, it’s your job to go down the rabbit hole, learn the 100%, and sprinkle in your 3%. I would say that there is a big difference between stealing without acknowledgement, a…
-
comment
Comment #48628364
> I don't think you should waste time reviewing every single line of code in here and just use AI to review it! > What you bring is the knowledge that the author nor the LLM doesn'…
-
comment
Comment #48627163
Thanks, I didn't mean to be brusque, but I have seen a lot of these vibe tests lately that come to grand conclusions like "X model is better than Y" from the result of a single pro…
-
comment
Comment #48627015
> So we ran it head-to-head against Claude Opus 4.8: same one-shot prompt, build a 3D platformer in raw WebGL from scratch Running a single one-shot prompt is not a benchmark, not …
- story
-
comment
Comment #48485287
Really curious to understand why I'm being downvoted. I don't think it's a particularly spicy take - Just choose the right tool for the job.
-
comment
Comment #48483411
As someone who has built both react based frontends and html based ones (with htmx), there is a law of diminishing returns at play. To start off, writing a basic crud website with …
-
comment
Comment #48469073
All the model releases we've seen this year have only made incremental improvements in benchmarks. This feels like the first release that feels like a significant step up in terms …
- story
- comment
-
comment
Comment #48261932
Not the first study, and they all largely report the same results: https://www.nature.com/articles/s41562-025-02259-6 https://www.theguardian.com/money/2019/feb/19/four-day-week-..…