Live data from Hacker News

Viewing profile — sam-paech

sam-paech

HN member
Joined
Thu, Apr 10, 2025, 8:11 AM UTC
HN karma
5
Public activity
9 items

About sam-paech

No profile information was provided.

Recent public activity

  1. comment
    Comment #45689214

    Those higher level kinds of mode collapse are hard to quantify in an automated way. To fix that, you would need interventions upstream, at pre & post training. This approach is tar…

  2. comment
    Comment #43643882

    All the judge outputs (including rubric) and model outputs are in the samples reports. Sorry you don't like the displayed metrics. I find them very useful / revealing of the things…

  3. comment
    Comment #43643598

    None of those factors go into the scoring fwiw. They are just informational. The scoring is done to a rubric, like a teacher would grade an essay, on various criteria for good & ba…

  4. comment
    Comment #43643583

    Different benchmark, those are for the short form creative writing leaderboard here: https://eqbench.com/creative_writing.html

  5. comment
    Comment #43643134

    Personally what I find interesting is getting insight into the trajectory of model abilities over time. Over the time I've been running these benchmarks, the writing has gone from …

  6. comment
    Comment #43642649

    Oops, should be: https://eqbench.com/creative_writing.html Sample outputs: https://eqbench.com/results/creative-writing-v3/gemini-2.5-p...

  7. comment
    Comment #43642005

    Hey, I made this! Cool to see it show up on hackernews.

  8. comment
    Comment #43641988

    The old version of the creative writing eval had several "in the style of" prompts actually! But I got tired of reading bad Hemingway impersonations so I cut them out of the new ve…

  9. comment
    Comment #43641948

    Not internal consistency exactly, but there are criteria checking how well the chapter plan was followed (which is all the way up at the top of the context window). This is done pe…