Live data from Hacker News

Viewing profile — jauws

jauws

HN member
Joined
Mon, Sep 11, 2023, 6:04 PM UTC
HN karma
17
Public activity
20 items

About jauws

No profile information was provided.

Recent public activity

  1. comment
    Comment #46880093

    Definitely on the to-do list! Right now, there's smth called fork (inspired by Github fork), where it lets you remix the story with a given input. It might be cool for you to mess …

  2. comment
    Comment #46880074

    Thanks Josh! I tried GEPA previously back when it was still 1-shot generation. It actually ended up working really well for some models and horrible for others, so I decided to scr…

  3. comment
    Comment #46877855

    If you look at similar live benchmarks like LMArena or Design Arena, there's an extremely large number of unique annotators, with a low number of annotations per person - which is …

  4. comment
    Comment #46877661

    I hope it's clear that the stories aren't being generated one-shot. I'm sure there are flaws that I haven't perfectly accounted for in the agent-loop, but because we randomize the …

  5. comment
    Comment #46877376

    Realistically, I don't think anyone will be spending hours here instead of reading real fiction anytime soon (I personally wouldn't). There's just so much nuanced complexity when i…

  6. comment
    Comment #46876837

    Happy to engage if you have concrete criticisms.

  7. comment
    Comment #46876833

    If you have specific objections, I’m open to hearing them.

  8. comment
    Comment #46876822

    Would love to chat! Here's my email: team@narrator.sh

  9. comment
    Comment #46876782

    Thanks for the feedback. What would you need to see to change your mind?

  10. comment
    Comment #46876758

    Thanks for the feedback - looking at the rest of the comments, I definitely agree it seems to be a common theme. Will do better to fix those issues so there's less noise.

  11. comment
    Comment #46876695

    I think there's interesting work to be built on this data beyond just generating and sorting slop. I didn't build this because I enjoy having people read bad fiction. I built it be…

  12. comment
    Comment #46876451

    There's 151 models there right now (with all the latest Anthropic models), it's all randomized, it's just that there aren't enough annotations for the anthropic models to be elicit…

  13. comment
    Comment #46875561

    Thanks for letting me know - the UI issues are definitely on me (fixing asap). Feel free to generate a story or two - right now there's not enough annotations to make "top-rated" a…

  14. comment
    Comment #46875419

    Ah shoot - thanks for letting me know. I'm still a noob on frontend so still learning as I go.

  15. story
    Show HN: I built "AI Wattpad" to eval LLMs on fiction

    I've been a webfiction reader for years (too many hours on Royal Road), and I kept running into the same question: which LLMs actually write fiction that people want to keep readin…

  16. comment
    Comment #44914098

    Thanks! Anecdotally, I'd tend to say that Claude 3.7 tends to improve the most, but it seems like (via the leaderboard), some people really prefer Grok-3 lol.

  17. comment
    Comment #44909184

    Thanks for the comment! Do you mind linking the site - would love to check it out! That's a very fair point about the technical error aspect. Though with all the confounding variab…

  18. comment
    Comment #44909090

    This is an amazing suggestion! Will definitely try to figure out a way to incorporate this into the leaderboard without making it a constant each time. I'm currently using OpenRout…

  19. comment
    Comment #44909054

    Thanks Johnny! I totally agree with you, really appreciate you for checking out my project!

  20. story
    Show HN: Evaluating LLMs on creative writing via reader usage, not benchmarks

    Hey HN! I'd love to get some people to mess around with a little side project I built to teach myself DSPy! I've been a big fan of reading fiction + webnovels for a while now, and …