Really like the shift from synthetic benchmarks to actual reader engagement — feels way more aligned with what “good writing” actually means. Curious if you’ve noticed certain models consistently improving more with feedback than others.
Show HN: Evaluating LLMs on creative writing via reader usage, not benchmarks
11–13 of 13 posts
Re: Show HN: Evaluating LLMs on creative writing via reader usage, not benchmarks
#12I run a site that does something similar, but on a more granular level (prompts at the page level rather than the chapter) I think right now we're at the point where novelcrafter is an excellent proxy for the best models for readers, because LLMs are still mostly losing engagement due to technical errors as opposed to subjective ones: That's repetition problems, moralizing/soft-censorship, grammatical quirks, missing…
Thanks for the comment! Do you mind linking the site - would love to check it out! That's a very fair point about the technical error aspect. Though with all the confounding variables (author skill differences, model selection based on price/speed, etc.) I'd say it's probably the most mature signal we have right now, but still far from ideal. Really interested in what you've been working on for the past year! Are you…
Most of it has been fine-tuning (SFT/DPO/GRPO), but also a lot of prompting and adding steps between the user's prompt and the output