Which AI has the best sense of humor?

9 min read Original article ↗

Findings

September 4: GPT-6 improves on GPT-5.6 Sol as a judge

GPT-6 Astra scored 69% across the four human-voted tests, compared with 62% for GPT-5.6 Sol. Each test contributes equally to that average.

The gain is 7.0 percentage points. In a comparison scored as a correct pick, an incorrect pick, or half credit when the two presentation orders disagree, that is about seven more points per hundred than GPT-5.6 Sol on this mix of tests.

The improvement is clearest on clean and dark jokes. Paired 95% bootstrap intervals exclude zero on the mixed, clean, and dark tests. On frontier-written jokes the gain is 1.5 points, with an interval from −4.2 to +6.9, so that test does not establish an improvement. These are individual intervals without adjustment for multiple comparisons.

Our judge still has the highest overall point estimate at 74%, 5.1 points above Astra. On clean jokes Astra scores 74% against our judge’s 73%, with a paired interval that includes zero. The remaining advantage is concentrated in other tests, especially mixed register. Several tests come from the same apps and reader population used to train our judge, which limits how broadly we can apply its lead.

Astra completed 1,833 pairs in both orders across six tests, with no missing verdicts. The four headline tests contain 1,339 pairs with a directional human majority. The curator test and the separate 100-pair clean test are excluded from the headline mean. An overlap check found no exact pairs or normalized jokes shared between these six tests and our judge’s training set.

The generation result uses our fine-tuned judge to score Astra’s writing. This test instead measures how Astra’s own choices agree with recorded human votes.

Which model agrees with people most?

For every pair of jokes, did the judge pick the same one as the majority of human raters?

Among the frontier models, GPT-6 Astra agrees with people the most

Overall averages four human-voted tests; the editor’s curated picks are shown separately. Shading marks each column’s best score. Hover for uncertainty ranges.

How wide are the gaps between the best and worst judges?

We took our main test and worked out the uncertainty range for each judge.

On the mixed test, our judge scores 78.7%; GPT-6 Astra scores 68.6%

Decisive only includes choices unchanged when joke order swaps. Undecided counts order-dependent choices; bars show 95% uncertainty intervals.

95% interval, marker at the point estimate

Does taste change on dark jokes?

We compared how well each frontier model matches human taste on clean jokes and on dark jokes.

GPT-6 leads the general models on dark jokes, but scores lower than on clean jokes

Each line connects a judge’s agreement with people on clean and dark jokes. Hover for exact values.

50%60%70%80%Clean jokesDark jokesLL judge 1 75%GPT-6 Astra 68%Kimi K2.6 50%

LL judge 1Frontier models

The clean-to-dark change varies by model

Change is dark minus clean, in percentage points, largest gain to largest loss.

Can a judge rank comedians?

Picking the winner of one pair and ranking whole systems are different jobs. Judges shown in this chart scored 5,028 pairs from the SemEval-2026 MWAHAHA competition, where about a hundred annotators ranked 31 joke-writing systems.

Good at pairs and good at ranking are different skills

Horizontal: similarity to the human ranking of 31 systems. Vertical: agreement on individual joke pairs. Hover to identify each judge.

55%60%65%70%75%0.650.700.750.800.85How closely its ranking of the 31 systems matches the human ranking (1 = identical)Agreement with human majority (4 human-voted tests)LL judge 1Gemini 3.1 ProClaude Opus 5Kimi K2.6

LL judge 1Frontier models

GPT-6 Astra has not run the separate ranking test, so it has no dot here.

Does a smarter model judge humor better?

We compared the models with an available LMArena text score against their agreement with people here. GPT-6 Astra is absent from this earlier comparison.

Smarter models tend to be better judges of humor (correlation +0.42 on a scale from -1 to +1)

Horizontal: LMArena text score on August 21, 2026. Vertical: average agreement with people on humor. Hover to identify each judge.

55%60%65%14201440146014801500LMArena text score (higher is smarter)Overall agreement with peopleGemini 3.7 FlashClaude Fable 5Claude Opus 5Mistral Medium 3.5Kimi K2.6

Our humor-only judge has no LMArena score, so it is excluded.

What do refusals and ties tell us?

Whether a judge answers at all, and whether it changes its mind when the jokes swap sides.

Claude Fable 5 refuses 5% of pairs. Nobody else refuses, and its successor Fable 5.1 answers everything

Refused is the share of judgments declined and excluded from scoring. Undecided is the share of pairs where swapping joke order changed the pick.

Other findings

  • More thinking does not help. In separate runs we tried the higher reasoning settings on Gemini 3.7 Flash, Gemini 3.1 Pro, GPT-5.5 and both Kimi models. None improved and several got worse, so the frontier runs used the low-reasoning requests and provider fallbacks listed in the methodology.
  • Writing and judging give different rankings. Fable 5 leads the Humor Arena, while GPT-6 Astra has the highest agreement among the general judges. Astra places third as a writer by point estimate.
  • Intelligence helps, moderately. Across the 15 frontier models, the LMArena text score and agreement with people correlate at +0.42 on a scale from -1 to +1 (rank correlation +0.30): a real but loose link, and with 15 models the uncertainty is wide. The top LMArena model, Claude Fable 5, sits mid-table here, so a high general score is no guarantee.
  • Undecided pairs are common. Most judges change their pick on 15 to 30% of pairs when the jokes swap sides. Mistral Medium 3.5 flips on more than 40%.

Methodology

The roughly 80% human reference measures agreement between two halves of this panel. It is not a fixed ceiling or the same measure as one person’s agreement.

Each pair is shown in both orders. A stable choice matching the human majority earns 1 point; a stable choice against it earns 0; an order disagreement earns 0.5. Refused pairs are excluded for that model. We average within each test, then give each of the four tests equal weight. “121k verdicts” counts model judgments across the evaluation runs, not human votes.

The judge panel

The frontier models tested so far, earlier releases for comparison, and our own judge. The most recent addition is GPT-6 Astra on September 4, 2026.

LL judge 1Laugh Labs

GPT-6 AstraOpenAI

Gemini 3.7 FlashGoogle

GPT-5.5OpenAI

Grok 4.5xAI

Gemini 3.1 ProGoogle

DeepSeek V4 ProDeepSeek

GLM-5.3Z.ai

Kimi K3Moonshot AI

Claude Fable 5.1Anthropic

Claude Fable 5Anthropic

Grok 4.6xAI

GPT-5.6 SolOpenAI

Claude Sonnet 5Anthropic

Gemini 3.8 FlashGoogle

TMInklingThinking Machines

Claude Opus 5Anthropic

Qwen3.8 MaxAlibaba

Mistral Medium 3.5Mistral

Kimi K2.6Moonshot AI

Three of the four headline tests draw on the same apps and largely the same reader population used to train our judge; on the separate test with new raters, its lead has overlapping uncertainty intervals.

LL judge 1, our own model

  • What it is. A Kimi K2.6 model fine-tuned to pick the funniest jokes (adapter kimi-k2p6-judge-haprolific), trained on 1,822 head-to-head votes from readers of our apps. It answers without explanation or reasoning.
  • What it was checked against. A panel of 51 raters who never trained it: on the pairs where it commits to an answer, it agrees with that panel more often than an extra human rater does (72% against 52%).
  • Where it stops. It declines to pick on about a quarter of pairs, and it is poor at ranking whole systems, so we route that job to a frontier model.
  • None of these pairs were in its training. Every test on this page was checked for overlap against everything it was trained on.

The tests

Each test is a list of joke pairs written for the same brief, with a human answer key. Sizes: mixed 239 pairs, clean 786, dark 147, frontier jokes 280 (167 with a human side), curated 281, ranking 5,028.

How the test was run

  • LLM and human judges never saw the author, model name, or vote count.
  • Every pair was shown twice with the jokes swapped, so a judge that follows screen position counts as undecided.
  • A judge scores 1 for picking the human side, 0 for the other side, 0.5 when its two answers disagree.
  • The frontier models used one shared prompt and the low-reasoning requests and provider fallbacks listed below; our fine-tuned judge used its terse A/B protocol. The initial grid ran on August 25; later entrants used the same frozen tests. GPT-6 Astra ran through the direct OpenAI API on September 4.

Models, settings and judging materials

These are the requested settings recovered from the runner. Where the API rejected a setting, the runner tried the listed alternatives; it did not save which fallback succeeded. We cannot recover those accepted settings from the saved verdicts.

Exact model IDs and requested settings

Gemini requests used temperature 0; other frontier requests left temperature unset. Fable 5.1 has saved predictions but is missing from the checked-in model registry, so its exact request configuration is marked unknown rather than inferred.

Read the frontier judging prompt
You are an expert comedy judge. Decide which of two jokes is funnier for the brief. End your reply with exactly 'FUNNIER: A' or 'FUNNIER: B' on the final line.

Brief: {brief}

Joke A: {joke_a}

Joke B: {joke_b}

Which joke is funnier?

Our fine-tuned judge used a separate terse A/B prompt, included in the download.

Caveats

We built the judge that wins most of these tables, so here is where the deck is stacked and where it is not.

  • Home advantage. The mixed, clean and dark tests come from our own apps and largely the same readers whose votes trained our judge, and our judge answers in the format it was trained on while the frontier models got one untuned prompt. The frontier-jokes test is the fair one: 280 pairs of jokes written by the Humor Arena models, voted on by 51 raters who never trained our judge. Our judge scores 67.7% there and leads the best frontier models by 3 to 6 points, with overlapping intervals. Treat the frontier numbers as a floor.
  • Ties and refusals. An undecided pair counts as half right for every judge, and a refused pair is dropped rather than scored wrong. On the pairs our judge commits to, its agreement is higher still, so the tie rule does not manufacture its lead.
  • Humans disagree too. On the closest calls two people agree with each other only about 54% of the time, and two halves of the same panel agree on about 80% of pairs overall. This measures disagreement in this sample. It does not impose a fixed ceiling on every judge or test; use the reported intervals when comparing models.
  • Not about joke writing. This page says nothing about which model writes the funniest jokes. That is the Humor Arena.
Per-test scores, intervals and eligible pairs