Did you skip Anthropic models? I honestly can't take this seriously if you're not looking at all the leading providers but you did look at some obscure ones.
Show HN: I built "AI Wattpad" to eval LLMs on fiction
11–20 of 44 posts
Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#12Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#13A very cool idea in theory and something very hard to pull off, but I think in order to get the data you need on how readable each story is you'll need to work on presentation and recommendation so those don't distract from what you're actually testing.
Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#14Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#15> The surge of AI, large language models, and generated art begs fascinating questions. The industry’s progress so far is enough to force us to explore what art is and why we make it. Brandon Sanderson explores the rise of AI art, the importance of the artistic process, and why he rebels against this new technological and artistic frontier. What It Means To Be Human | Art in the AI Era https://www.youtube.com/watch?v…
Do watch the video as it makes a compelling argument against this exact kind of thing. From a product design perspective, you're asking people to read a bunch of slop and organize it into slop piles. What's the point of that? Honestly it seems like a huge waste of everyone's time.
More broadly, crowdsourced data where human inputs are fundamentally diverse lets us study problems that static benchmarks can't touch. The recent "Artificial Hivemind" paper (Jiang et al., NeurIPS 2025 Best Paper) showed that LLMs exhibit striking mode collapse on open-ended tasks, both within models and across model families, and that current reward models are poorly calibrated to diverse human preferences. Fiction at scale is exactly the kind of data you need to diagnose and measure this. You can see where models converge on the same tropes, whether "creative" behavior actually persists or collapses into the same patterns, and how novelty degrades over time. That signal matters well beyond fiction, including domains like scientific research where convergence versus originality really matters.
Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#16Do you have a contact email?
Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#17Hard to find the signal in the noise and know what stories I should even read to get a sense of baseline quality; partially because that's just a hard problem inherent to floods of any content, but also because the recommendation system seems to lack enough data (and also might be weighting the wrong things, e.g. the rank #1 story is also the lowest-rated...). A very cool idea in theory and something very hard to pul…
Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#18[flagged]
Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#19Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction
#20I have a lot of engagement data on LLMs from running a creative writing oriented consumer AI app and spending s lot of time on quality improvements and post training Do you have a contact email?