Settings

Theme

Show HN: Humor Arena – Which frontier model is funniest?

laugh.so

10 points by killiandunne1 · 10 comments · 1 min read

Reader

What if you could measure humor?

Well we've trained a model on our own dataset of ~50k human ratings to detect what jokes people find funniest. We know it's part objective, part subjective component. Subjective is out of our depth for now haha

The main results: Fable 5 is funniest - beating the average model 67% of the time, with GPT 4o last at 17%.

Other findings: - The models never refused to try, even with dark prompts - Thinking longer has a slight benefit - Absurdness correlates negatively with joke quality

Some methodology notes: - We benchmarked our model against the human majority and it agreed 72% of the time in a blind sample test. - We had 51 US adults rate the jokes, each blind to the models, with joke order randomized, and quality checked for attention and speed. - To rate some yourself visit https://pair.laugh.so

The full benchmark here:

https://laugh.so/benchmark

Am taking requests if there's more research you want to see! Cheers

3 threads
dlcarrier

Can you also measure how often the LLM response makes people laugh? Sometimes the responses that aren't attempting a joke are the funniest, and I'd be more interested in stats of which LLM succeed in that metric.

  • killiandunne1OP

    Interesting point - thoughts on how to do this? Honestly most models are not-to-kinda funny so I'd be surprised if there were many lol moments. Oral delivery is something I think is v interesting though

    • dlcarrier

      You'll probably still get plenty of smirks, smiles, and silent chuckles. OpenCV or a similar machine vision system should be able to pick those out from a webcam aimed at a reviewer.

maxzhdev

I can’t imagine how to correctly assess the ability to humor, because this is a very subjective and relative phenomenon. The work ahead is serious

rafaepta

measuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection