Viewing profile — jauws
jauws
HN member- Joined
- Mon, Sep 11, 2023, 6:04 PM UTC
- HN karma
- 17
- Public activity
- 20 items
- HN profile
- View on Hacker News ↗
About jauws
No profile information was provided.
Recent public activity
-
comment
Comment #46880093
Definitely on the to-do list! Right now, there's smth called fork (inspired by Github fork), where it lets you remix the story with a given input. It might be cool for you to mess …
-
comment
Comment #46880074
Thanks Josh! I tried GEPA previously back when it was still 1-shot generation. It actually ended up working really well for some models and horrible for others, so I decided to scr…
-
comment
Comment #46877855
If you look at similar live benchmarks like LMArena or Design Arena, there's an extremely large number of unique annotators, with a low number of annotations per person - which is …
-
comment
Comment #46877661
I hope it's clear that the stories aren't being generated one-shot. I'm sure there are flaws that I haven't perfectly accounted for in the agent-loop, but because we randomize the …
-
comment
Comment #46877376
Realistically, I don't think anyone will be spending hours here instead of reading real fiction anytime soon (I personally wouldn't). There's just so much nuanced complexity when i…
-
comment
Comment #46876837
Happy to engage if you have concrete criticisms.
-
comment
Comment #46876833
If you have specific objections, I’m open to hearing them.
-
comment
Comment #46876822
Would love to chat! Here's my email: team@narrator.sh
-
comment
Comment #46876782
Thanks for the feedback. What would you need to see to change your mind?
-
comment
Comment #46876758
Thanks for the feedback - looking at the rest of the comments, I definitely agree it seems to be a common theme. Will do better to fix those issues so there's less noise.
-
comment
Comment #46876695
I think there's interesting work to be built on this data beyond just generating and sorting slop. I didn't build this because I enjoy having people read bad fiction. I built it be…
-
comment
Comment #46876451
There's 151 models there right now (with all the latest Anthropic models), it's all randomized, it's just that there aren't enough annotations for the anthropic models to be elicit…
-
comment
Comment #46875561
Thanks for letting me know - the UI issues are definitely on me (fixing asap). Feel free to generate a story or two - right now there's not enough annotations to make "top-rated" a…
-
comment
Comment #46875419
Ah shoot - thanks for letting me know. I'm still a noob on frontend so still learning as I go.
-
story
Show HN: I built "AI Wattpad" to eval LLMs on fiction
I've been a webfiction reader for years (too many hours on Royal Road), and I kept running into the same question: which LLMs actually write fiction that people want to keep readin…
-
comment
Comment #44914098
Thanks! Anecdotally, I'd tend to say that Claude 3.7 tends to improve the most, but it seems like (via the leaderboard), some people really prefer Grok-3 lol.
-
comment
Comment #44909184
Thanks for the comment! Do you mind linking the site - would love to check it out! That's a very fair point about the technical error aspect. Though with all the confounding variab…
-
comment
Comment #44909090
This is an amazing suggestion! Will definitely try to figure out a way to incorporate this into the leaderboard without making it a constant each time. I'm currently using OpenRout…
-
comment
Comment #44909054
Thanks Johnny! I totally agree with you, really appreciate you for checking out my project!
-
story
Show HN: Evaluating LLMs on creative writing via reader usage, not benchmarks
Hey HN! I'd love to get some people to mess around with a little side project I built to teach myself DSPy! I've been a big fan of reading fiction + webnovels for a while now, and …