Live data from Hacker News

Show HN: Humor Arena – Which frontier model is funniest?

laugh.so

1–10 of 11 posts

Show HN: Humor Arena – Which frontier model is funniest?

#1
What if you could measure humor?

Well we've trained a model on our own dataset of ~50k human ratings to detect what jokes people find funniest. We know it's part objective, part subjective component. Subjective is out of our depth for now haha

The main results: Fable 5 is funniest - beating the average model 67% of the time, with GPT 4o last at 17%.

Other findings: - The models never refused to try, even with dark prompts - Thinking longer has a slight benefit - Absurdness correlates negatively with joke quality

Some methodology notes: - We benchmarked our model against the human majority and it agreed 72% of the time in a blind sample test. - We had 51 US adults rate the jokes, each blind to the models, with joke order randomized, and quality checked for attention and speed. - To rate some yourself visit https://pair.laugh.so

The full benchmark here:

https://laugh.so/benchmark

Am taking requests if there's more research you want to see! Cheers

Show HN: Humor Arena – Which frontier model is funniest?
laugh.so

Re: Show HN: Humor Arena – Which frontier model is funniest?

#4
measuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.

Re: Show HN: Humor Arena – Which frontier model is funniest?

#5
post #4

measuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.

Nice a good way to 10x inference costs... worth seeing though ahaha

Re: Show HN: Humor Arena – Which frontier model is funniest?

#7

Can you also measure how often the LLM response makes people laugh? Sometimes the responses that aren't attempting a joke are the funniest, and I'd be more interested in stats of which LLM succeed in that metric.

Interesting point - thoughts on how to do this? Honestly most models are not-to-kinda funny so I'd be surprised if there were many lol moments. Oral delivery is something I think is v interesting though
Post reply on HN