Is a hot dog a sandwich? 12 AI models answer, every week — HOTDOG BENCHMARK

HOTDOG BENCHMARK

3 min read Original article ↗

Every week, the largest AI models are asked the question:

One word answer.

  1. Claude Opus 5Anthropic

    No.reasoning

    2.7 sReasoned for 2.7 s (100% of the call) on 93 tokens, then answered · 3 of 3 runs agreed

  2. Claude Sonnet 5Anthropic

    Yes.reasoning

    1.5 sReasoned for 1.4 s (96% of the call) on 20 tokens, then answered · 2 of 3 runs agreed

  3. Claude Haiku 4.5Anthropic

    Yes.

    607 msNo reasoning, answered straight away · 3 of 3 runs agreed

  4. GPT-5.6 SolOpenAI

    Yes.

    2.4 sNo reasoning, answered straight away · 3 of 3 runs agreed

  5. GPT-5.5OpenAI

    Yesreasoning

    1.9 sReasoned for 1.7 s (89% of the call) on 29 tokens, then answered · 3 of 3 runs agreed

  6. GPT-5.4 miniOpenAI

    No

    1.3 sNo reasoning, answered straight away · 2 of 3 runs agreed

  7. Grok 4.6xAI

    Noreasoning

    10.4 sReasoned for 10.3 s (100% of the call) on 530 tokens, then answered · 2 of 3 runs agreed

  8. Grok 4.3xAI

    Noreasoning

    4.6 sReasoned for 4.5 s (99% of the call) on 484 tokens, then answered · 2 of 3 runs agreed

  9. Grok 4.20 (non-reasoning)xAI

    Yes.

    464 msNo reasoning, answered straight away · 3 of 3 runs agreed

  10. Mistral Medium 3.5Mistral AI

    no answer

    rate limit

  11. Mistral Small 4Mistral AI

    no answer

    rate limit

  12. DeepSeek V4 ProDeepSeek

    No.reasoning

    3.3 sReasoned for 3.2 s (99% of the call) on 137 tokens, then answered · 3 of 3 runs agreed

5 said yes5 said no

Recorded week 37, 2026. Real durations, verbatim words. Teal is the wait before the first word, hatched where the model spent it reasoning; the rest is answering.Read the report →

Same question, different minds

They do not agree with each other.

Thinking alike: Claude Opus 5 & GPT-5.4 mini · Claude Sonnet 5 & Claude Haiku 4.5 · Claude Sonnet 5 & GPT-5.6 Sol · and 11 more pairs

Read the 6 reports →One straight-faced analyst report per question: standings, the certainty quadrant, every verbatim answer under every framing, and a PDF for each.

Tell them the answer

Some of them believe you.

Share of questions where a model changed its answer once a system prompt stated the answer as fact. Holding firm and following instructions are both defensible; the methodology grades neither.

  1. GPT-5.6 Sol50%6 of 12
  2. GPT-5.550%6 of 12
  3. GPT-5.4 mini50%6 of 12
  4. Grok 4.20 (non-reasoning)50%6 of 12
  5. Claude Sonnet 533%4 of 12
  6. Claude Haiku 4.525%3 of 12
  7. Claude Opus 517%2 of 12
  8. DeepSeek V4 Pro17%2 of 12
  9. Grok 4.68%1 of 12
  10. Grok 4.38%1 of 12

Submit your own question

Ask the models something.

Send it in. An accepted question appears here under Up next, credited to you if you want, then joins an edition and gets its own report. Every question is asked the same way, so it ends with One word answer.; we add that if you leave it off.

Where it goes:

Open source

Point it at your own question.

One repo, MIT-licensed: adapters for every provider, the framings, the site. Clone it, swap the question, add whatever keys you have, and you get the same cross-model, cross-framing analysis for cents. Pull requests welcome.

GitHubSelf-hostingAdd a modelContributing

Have a question the models should get? Send it in.

git clone https://github.com/en-dash-consulting/hotdogbenchmark.git
cd hotdogbenchmark && npm install
npm run bench -- run --mock --out tmp/mock-run.json
npm run dev

Week 37, 2026 · published September 7, 2026 · 2 editions so far