Live data from Hacker News

Show HN: I built "AI Wattpad" to eval LLMs on fiction

narrator.sh

21–30 of 44 posts

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#23
post #18
post #14

[flagged]

Thanks for the feedback. What would you need to see to change your mind?

There's more quality fiction out there than you or I will ever have time to read. I don't see a purpose in flooding the world with more mediocre to unreadable fiction.

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#24
post #18
post #14

[flagged]

Thanks for the feedback. What would you need to see to change your mind?

I am not going to argue this on the basis of LLM's suck at fiction, because even if it's true, it's not really that relevant. The problem is that what LLM's are good at is producing mediocre fiction particular to the tastes of the individual reading at. What people will keep reading is fiction that an LLM is writing because they personally asked it to write it.

I don't want to read fiction generated from someone else's ideas. I want to read LLM fiction generated from my weird quirks and personal taste.

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#25
post #23
post #18

Earlier quoted context omitted.

Thanks for the feedback. What would you need to see to change your mind?

There's more quality fiction out there than you or I will ever have time to read. I don't see a purpose in flooding the world with more mediocre to unreadable fiction.

Realistically, I don't think anyone will be spending hours here instead of reading real fiction anytime soon (I personally wouldn't). There's just so much nuanced complexity when it comes to creative writing as a domain (long-form outputs, creativity, etc.) that coming up with better annotation methods has massive applications in other research, like in scientific discovery. "AI Wattpad" just happens to be a convenient form factor for crowdsourcing from an HCI perspective. I hope you give it a chance.

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#26
post #22

[flagged]

Happy to engage if you have concrete criticisms.

have you read any of the generated stories? if you can honestly tell me this is not complete drivel (even worse, wildly generic and poorly written) then i will consider giving real feedback but i would find that hard to believe.

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#27
post #22

Earlier quoted context omitted.

Happy to engage if you have concrete criticisms.

have you read any of the generated stories? if you can honestly tell me this is not complete drivel (even worse, wildly generic and poorly written) then i will consider giving real feedback but i would find that hard to believe.

I hope it's clear that the stories aren't being generated one-shot. I'm sure there are flaws that I haven't perfectly accounted for in the agent-loop, but because we randomize the models for each of the brainstorming -> writing -> memory parts, bad intermediate outputs will affect the final output as well. That's why unless we have above average models across all 3 stages, it might be worse than what you're used to. It's a trade-off to get more granular results. Hope you can give it a chance.

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#28
post #25
post #23

Earlier quoted context omitted.

There's more quality fiction out there than you or I will ever have time to read. I don't see a purpose in flooding the world with more mediocre to unreadable fiction.

Realistically, I don't think anyone will be spending hours here instead of reading real fiction anytime soon (I personally wouldn't). There's just so much nuanced complexity when it comes to creative writing as a domain (long-form outputs, creativity, etc.) that coming up with better annotation methods has massive applications in other research, like in scientific discovery. "AI Wattpad" just happens to be a convenie…

OK, so you already recognize these stories aren't something that people are going to spend time sorting through. How could you possibly then get any usable preference data out of this?

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#29
post #28
post #25

Earlier quoted context omitted.

Realistically, I don't think anyone will be spending hours here instead of reading real fiction anytime soon (I personally wouldn't). There's just so much nuanced complexity when it comes to creative writing as a domain (long-form outputs, creativity, etc.) that coming up with better annotation methods has massive applications in other research, like in scientific discovery. "AI Wattpad" just happens to be a convenie…

OK, so you already recognize these stories aren't something that people are going to spend time sorting through. How could you possibly then get any usable preference data out of this?

If you look at similar live benchmarks like LMArena or Design Arena, there's an extremely large number of unique annotators, with a low number of annotations per person - which is normal. However, since this platform is designed to generate fiction catered to individual interests, my hypothesis is that it'll be an added boost of novelty that will help aggregate enough usable data over time.

Re: Show HN: I built "AI Wattpad" to eval LLMs on fiction

#30
post #29
post #28

Earlier quoted context omitted.

OK, so you already recognize these stories aren't something that people are going to spend time sorting through. How could you possibly then get any usable preference data out of this?

If you look at similar live benchmarks like LMArena or Design Arena, there's an extremely large number of unique annotators, with a low number of annotations per person - which is normal. However, since this platform is designed to generate fiction catered to individual interests, my hypothesis is that it'll be an added boost of novelty that will help aggregate enough usable data over time.

I tried reading the two top rated stories. They're both unreadable gunk. Why would I (or anyone else) go for another? Why would I tell anyone I know to spend their time reading this?
Post reply on HN