HIPPO/EVAL — The benchmark was benchmarked

1 min read Original article ↗

A small change to our very serious benchmark

We retired the pelican.
Long live the hippo.

01 / THE PROBLEM

A benchmark everyone has seen is a benchmark a model might have seen, too. Pelican-on-a-bike is charming, compact, and surprisingly revealing—but familiarity can masquerade as capability.

So we changed the nouns, kept the difficulty, and made the test weird again.

“Same coordination problem. New animal. Less benchmark leakage.”— our entire methodology, basically

02 / SIDE-BY-SIDE

Pelican vs. hippo

Reserved for outputs from the latest local open-model runs. Same renderer, same scoring rubric, fresh prompt.

ModelOld promptNew prompt

Qwenlocal · candidate 01next run

SVG output placeholderpelican + bicycle

awaiting run

SVG output placeholderhippo + pogo stick

awaiting run

DeepSeeklocal · candidate 02

SVG output placeholderpelican + bicycle

awaiting run

SVG output placeholderhippo + pogo stick

awaiting run

Llamalocal · candidate 03

SVG output placeholderpelican + bicycle

awaiting run

SVG output placeholderhippo + pogo stick

awaiting run

Mistrallocal · candidate 04

SVG output placeholderpelican + bicycle

awaiting run

SVG output placeholderhippo + pogo stick

awaiting run

03 / WHAT STAYS CONSTANT

New mascot. Same test.

We still score the things that make SVG generation useful—not whether the animal looks cute in a screenshot.

  1. 01Valid, editable SVG
  2. 02Prompt adherence
  3. 03Object relationships
  4. 04Visual coherence

THE NEW CONTROL PROMPT

v2.0

“Create an SVG of a hippo riding a pogo stick.”

Copy it. Try it. Send us the weird ones.