Viewing profile — sam-paech
sam-paech
HN member- Joined
- Thu, Apr 10, 2025, 8:11 AM UTC
- HN karma
- 5
- Public activity
- 9 items
- HN profile
- View on Hacker News ↗
About sam-paech
No profile information was provided.
Recent public activity
-
comment
Comment #45689214
Those higher level kinds of mode collapse are hard to quantify in an automated way. To fix that, you would need interventions upstream, at pre & post training. This approach is tar…
-
comment
Comment #43643882
All the judge outputs (including rubric) and model outputs are in the samples reports. Sorry you don't like the displayed metrics. I find them very useful / revealing of the things…
-
comment
Comment #43643598
None of those factors go into the scoring fwiw. They are just informational. The scoring is done to a rubric, like a teacher would grade an essay, on various criteria for good & ba…
-
comment
Comment #43643583
Different benchmark, those are for the short form creative writing leaderboard here: https://eqbench.com/creative_writing.html
-
comment
Comment #43643134
Personally what I find interesting is getting insight into the trajectory of model abilities over time. Over the time I've been running these benchmarks, the writing has gone from …
-
comment
Comment #43642649
Oops, should be: https://eqbench.com/creative_writing.html Sample outputs: https://eqbench.com/results/creative-writing-v3/gemini-2.5-p...
-
comment
Comment #43642005
Hey, I made this! Cool to see it show up on hackernews.
-
comment
Comment #43641988
The old version of the creative writing eval had several "in the style of" prompts actually! But I got tired of reading bad Hemingway impersonations so I cut them out of the new ve…
-
comment
Comment #43641948
Not internal consistency exactly, but there are criteria checking how well the chapter plan was followed (which is all the way up at the top of the context window). This is done pe…