#1Robert Fagles · Penguin
69% pairwise
Sing to me of the man, Muse, the man of twists and turns driven time and again off course, once he had plundered the hallowed heights of Troy.
Results through July 21, 2026 at 6:30 p.m. ET
We hid the translators’ names and asked readers to choose between eight English versions. As of July 26, 1,538 readers have completed this test making over 18,133 pairwise comparisons. The initial snapshot covers the first 685 completed tests and 8,043 head-to-heads.
Cumulative counted results at each cutoff: July 26 at 9:30 p.m. ET and July 21 at 6:30 p.m. ET.
Robert Fagles led the initial analysis of 674 readers. His Penguin translation ranked first by raw matches, in the regularized pairwise model, and in the scene-adjusted model. It also finished first in all 50,000 whole-ballot bootstrap samples.
The July 26 snapshot included 1,538 readers and 18,133 head-to-heads. Applying the same under-90-second rule left 1,524 readers and 17,973 head-to-heads. Fagles was still first. All eight pairwise positions were unchanged.
Pairwise strength estimates how often a translation would beat an average opponent in this field. It accounts for which opponents each translation faced, so it is not a reader percentage. The 95% ranges are shown by default.
The result is limited to eight translations, eight short scenes, and a self-selected, mostly American social-media sample. Greek fidelity and the experience of reading a complete translation were outside the test.
After publication, we heard from three scholars and one of the translators in the test:
“My compliments to you and your colleagues for your brilliant work.”
— Gregory Nagy, Harvard
“An ingenious and timely online poll that gives real nuance to the ballooning interest in all things Odyssey today.”
— Michael Witmore, former director of the Folger Shakespeare Library
“I appreciate that you’re using technology not to short-circuit the reading experience but to expand our conversations about it—and how we encounter the rhetorical texture of great books. It’s a shame that Pope’s technical virtuosity has fallen on hard times.”
— Pasquale Toscano, assistant professor of English at Vassar College
The test presented text first, while Homeric epic began as oral performance. On July 24, 2026, we spoke with Dr. Stanley Lombardo, professor emeritus of Classics at the University of Kansas. His Odyssey translation was one of the eight tested. We asked what the text-first format missed:
“Readers should hear the verse and the nuance that’s created by rhythm. It was an oral performance from the beginning and it was oral for generations before it was written down. That’s how I would go about it and respect it and honor it and try to transmit it.”
— Stanley Lombardo, professor emeritus of Classics at the University of Kansas
We considered 17 translations and chose the eight with the highest estimated named search attention. The estimates cover trailing 12-month U.S. searches for each exact full-name-plus-Odyssey query. They gave us a reproducible cutoff for the test, but say nothing about sales, readership, classroom use, or literary importance.
Linear scale · trailing 12-month average
1Emily Wilson9,617
2Robert Fagles2,212
3Robert Fitzgerald704
4Richmond Lattimore419
5Daniel Mendelsohn355
6Samuel Butler303
7Alexander Pope226
8Stanley Lombardo121
9T. E. Lawrence70
10Stephen Mitchell68
11E. V. Rieu59
12W. H. D. Rouse55
13Peter Green49
14George Chapman35
15A. T. Murray15
16Barry Powell< floor
17Anthony Verity< floor
4.3× Wilson over Fagles80× Wilson over Lombardo
Mobile96.7%
Desktop2.7%
Tablet0.7%
United States92.8%
International7.2%
Male50.3%
Female48.3%
Unknown1.4%
Most readers were in the United States, on mobile phones, and reached through social media. The largest source was our Instagram post, with additional traffic from TikTok, Facebook, YouTube, and readers sharing the test.
The campaign reached people who already cared about Homer or Christopher Nolan’s Odyssey. That is the study’s clearest selection bias.
Device, geography, and source figures come from campaign and visitor analytics. Gender comes from paid campaign-delivery data and was not collected on the ballot. These figures describe the audience reached by the campaign; they cannot be mapped precisely to either the 685 counted completions or the 674 analyzed ballots. We did not retain raw IP addresses or direct identifiers; one-way IP hashes were used to detect retakes.
Each lane is scaled to its own peak so the timing remains visible.
Named search attention did not predict the blind ranking very well. Emily Wilson had the most search attention but did not lead the test.
Eight tested translations · US Google search attention
Moving right means more named search attention; moving up means stronger pairwise strength. The lines at 1,000 average monthly searches and 50% pairwise strength are arbitrary visual guides. Moving the guides changes the quadrant labels, not the points or underlying data.
Underrated
FitzgeraldLattimoreMendelsohnLombardo
Popularity uses DataForSEO's trailing 12-month estimate of average monthly US Google searches for each exact full-name-plus-Odyssey query. It measures named search attention. It does not measure sales, readership, or literary quality.
Emily Wilson argues that readers may mistake archaic language for authority or authenticity:
“Mild stylistic archaism is often accepted without question in translations of ancient texts and can be presented as if it were a mark of authenticity. But of course, the English of the nineteenth or early twentieth century is no closer to Homeric Greek than the language of today. The use of a noncolloquial or archaizing linguistic register can blind readers to the real, inevitable, and vast gap between the Greek original and any modern translation.”
The chart compares publication date with blind preference. Newer translations look stronger across the full 300-year span, but much of that pattern comes from Pope. Remove him and it weakens. Among translations published since 1960, it nearly disappears. Readers were not simply choosing the oldest-sounding English.
With Pope, the line slopes upward across all eight translations (correlation 0.71). Among translations published since 1960, it becomes -0.16. The pattern mostly separates older verse styles from the modern field. Within the modern group, newer does not consistently mean stronger.
Google AI mentions and blind preference had a correlation of 0.22 across the eight translations. Wilson appeared most often in DataForSEO’s observed corpus but ranked sixth in pairwise strength.
The ChatGPT correlation was 0.32. It surfaced some of the leaders, but rarely mentioned Mendelsohn, who ranked third in the blind test.
We excluded 11 timed completions under 90 seconds from the initial analysis. Restoring them did not change the winner.
DiscardedSkimSlowMedian 4:01
The 90-second cutoff removes 11 completions. Finishing that quickly required reading 24 excerpts and making 12 decisions at several hundred words per minute, with no time left to choose. The skim group took 1:30 to 2:14. The slow group took 2:15 or more.
The all-reader view contains the full analyzed cohort of 674 ballots. Fagles leads overall.
Fagles still leads when all 11 completions under 90 seconds are restored. The pairwise order is unchanged across all 685 readers. Raw match shares barely move; Fitzgerald and Wilson tie at 79 top matches.
Pairwise strength estimates how often a translation would beat an average opponent in this field, accounting for which opponents it faced.
We include this view so readers can audit the 90-second cutoff. The rest of the report applies that rule consistently, but the cutoff did not create the winner.
We found no coordinated bot or voting campaign. The site counted one ballot per anonymous browser and blocked retakes and later submissions from browsers that had already cast a counted ballot.
To check shared-network activity, we converted each IP address into a private one-way hash. This let us compare matching addresses without keeping the raw address in ballot records. Retakes never entered the totals.
We used GPT-5.6 Sol to build a fixed pool of 25 scenes from Samuel Butler’s public-domain Project Gutenberg translation. We ranked those scenes by how often they appeared in nine study guides, then asked the model to choose eight that covered different parts and themes of the poem.
The model never saw the other seven translations or any reader results while choosing scenes. We then compared its panel with all 1,081,575 possible eight-scene subsets from the same pool. Only 5,059 subsets, or 0.47%, matched or beat it on both recognition and semantic breadth. The benchmark checks our selection against the candidate pool. The scenes were not randomly sampled.
Our panel compared with every eight-scene combination from the same 25 candidates
Middle 50% of panelsChosen panel
98.4th percentile
Indexed quotation recurrence
90.2nd percentile
Google-result visibility
85.0th percentile
Study-guide inclusion
Fagles leads overall, but other translators win individual scenes. Pick one to compare its first- and eighth-place passages.
1,197 blind choices
#1Robert Fagles · Penguin
69% pairwise
Sing to me of the man, Muse, the man of twists and turns driven time and again off course, once he had plundered the hallowed heights of Troy.
#8Alexander Pope · OUP
26% pairwise
The man for wisdom’s various arts renown’d, Long exercised in woes, O Muse! resound; Who, when his arms had wrought the destined fall Of sacred Troy, and razed her heaven-built wall.
This model gives equal weight to scenes, opponents, and recruitment waves and controls for first/second position. The top three stay in place. Fitzgerald and Lombardo swap fourth and fifth.
Our original analysis required 20% separation and found no distinct reader camps. At a lower 15% threshold, one split appears. Its separation is 16.4%, and it returns in 97.6% of resamples.
674 ballots · candidate solutions from two to five groups
0 qualifying solutions
A reader type had to clear both gates. Some splits repeated reliably, but every split fell short on separation.
Gate 20%
What this means: the aggregate winner does not appear to conceal a large, coherent minority camp. It does not mean every reader wants the same thing.
Fagles ranks first on both sides of the split. The largest difference is at the bottom: one side favors Wilson and rejects Pope, while the other reverses them. Lattimore barely moves.
15% separation gate
Wilson side
56%
377 readers
Pope side
44%
297 readers
0%25%50%75%
Pope
Wilson
Butler
Lombardo
Fitzgerald
Fagles
Mendelsohn
Lattimore
We are not classicists and do not read Ancient Greek. The sources below helped us understand why “most accurate” is not one question. A translation may preserve wording, poetic effect, or ambiguity, but rarely all three at once.
Homer calls Odysseus πολύτροπον (polytropon), or “many-turning.” It can mean traveled, clever, adaptable, evasive, or deceptive.
Emily Wilson argues that Fagles adds drama and cumulative motion to Homer’s compact opening. Her comparison is useful, but it also makes the case for her own method.
Literal facts are not the whole poem. In Book 18, Homer uses πετάννυμι (petannymi), “to open or spread out,” when desire wakes in Penelope’s suitors.
Corinne Pache uses this example to show how a less literal sentence can keep the Greek image. Meter constrains the translator’s available choices: line length, rhythm, syntax, and imagery compete for limited space. Wilson uses short pentameter. Mendelsohn uses a longer six-beat line. Kim Montpelier explains Mendelsohn’s method. Gregory Nagy explains why he values several translations.
Every translation picks a meaning. πολύμητις (polymetis) means “rich in metis.” That can mean intelligence, planning, cunning, trickery, or deception. Wilson’s “lord of lies” picks one side. Pache thinks it picks too hard.
Social language has the same problem. “Maid” can hide enslavement. “Sluts” or “whores” can add a judgment Homer did not make. Wilson writes about inherited bias, and Pache reviews her choices. Pasquale Toscano argues that a new translation should reveal new possibilities, not try to replace its predecessors.
“My goal is not to rank them, nor to criticize my esteemed predecessors and colleagues.” - Emily Wilson
Wilson’s essay compares how particular translations echo and depart from the Greek. We cannot judge the Greek ourselves, so we have linked the sources we used.
We displayed the excerpts without their original verse lineation. The test therefore captures diction, syntax, imagery, and short stretches of cadence better than meter, enjambment, couplet structure, or the experience of a complete work.
Flexible six-beat verseAccentual verseIambic pentameterProse
The top three translations use flexible six-beat verse, but length alone does not explain the ranking. Pope’s excerpts were nearly as long as the leaders’ and ranked eighth. Strict iambic pentameter and prose did not lead.
As of July 22, 2026, Scrivium had no financial or editorial relationship with any tested translator or publisher. We do benefit when people visit our interactive mythology course. That is our commercial interest in the subject.
We hid translator names, dates, and publishers. Readers saw up to 12 comparisons. On 668 of the 674 analyzed ballots, the opening eight formed a balanced cycle in which every available translator appeared twice. Before each of the final four comparisons, the test refit a regularized Bradley–Terry model using that reader’s earlier head-to-heads. It first covered translators with the least evidence. If skips had split the comparison network, it favored pairs that joined those fragments. The remaining score rewarded close contenders, uncertainty, and pairs the reader had not already seen. Scene assignment followed the phase-specific rules detailed below. Left and right positions were independently randomized for every comparison. Skips changed neither translator’s preference score, although the pair and scene still counted as having been shown. Detailed implementation facts:
Initialization: Every translator began at ability 0, equivalent to a neutral 50% predicted win probability.
Updating scores: After each completed head-to-head, the browser refit the model from all previous completed head-to-heads. It ran 80 gradient iterations with L2 regularization of 0.7, shrinking sparse estimates toward zero. Skips were omitted from the ability fit.
Eligible pairs: The starting universe was every unordered pair among the translators present in that ballot—28 pairs for eight translators.
Coverage and connectivity filters: If completed head-to-heads left disconnected groups, only pairs joining two groups were considered. It then retained pairs containing the largest possible number of translators tied for the fewest decisions.
Repeated pairs: Previously shown pairs were excluded whenever an unseen pair survived those filters. Repeats were permitted only when necessary and were penalized according to 1 / (1 + previous meetings). A skipped pair still counted as previously shown.
Closeness: For abilities aᵢ and aⱼ, closeness was 1 / (1 + |aᵢ − aⱼ|). This was only one part of the composite score, alongside contender strength, uncertainty, connectivity, and novelty.
Final selection: The algorithm kept the six highest-scoring eligible pairs and made a weighted selection among them. Higher scores were more likely, but the highest score was not automatically chosen.
Scene assignment: The method changed prospectively. Analyzed ballots 1–6 used the earlier four-scene cycle. Ballots 7–447 balanced scene exposure while also considering editorial strength, lexical contrast, and pair freshness. For analyzed ballots 448–674, the selector removed the immediately previous scene when alternatives existed, retained the least-shown scenes, sorted them by ID, and selected one with a ballot-specific deterministic hash. Only this final phase was fully blind to candidate identity, wording, and the reader’s selections.
Unavailable material: Pair selection only saw translators in the session’s candidate set; scene assignment only saw passages present in that session. Nothing unavailable was imputed or backfilled.
Left/right order: The logical pair was independently flipped with a 50/50 cryptographic random draw before every comparison.
Randomness and seeds: There was no fixed study-wide assignment seed. The server cryptographically shuffled translator and passage order and generated a random ballot UUID. That UUID became a per-ballot seed for reproducible adaptive-pair and scene assignment. Left/right flips remained separately cryptographically random and unseeded.
A JSON validation bug changed which options were available early in collection. Daniel Mendelsohn was missing for the first 6 readers and entered with reader 7. Harmonious household and The Lotus-Eaters were missing for the first 457 readers. Of those, 447 remain in the analysis. The scenes entered with reader 458, or analyzed ballot 448.
We did not rewrite or remove ballots because of the bug. The report separately excludes 11 ballots completed in under 90 seconds, including 10 collected while the two scenes were missing. Opening rounds balanced the options available at the time. The scene-adjusted model accounts for translator-by-scene and recruitment-wave effects.
The published ranking fits a regularized Bradley–Terry model:
P(i ≻ j) = σ(βᵢ − βⱼ) = 1 / (1 + e−(βᵢ−βⱼ))
β is a translation’s estimated ability. We report pairwise strength as its average probability of beating each of the other seven translations. A unit L2 penalty pulls weakly observed estimates toward the field average instead of allowing extreme scores from sparse matchups.
The compensated scene model adds partially pooled scene effects, recruitment wave effects, and a first-position term:
logit P(i ≻ j) = (αᵢ − αⱼ) + (sₖᵢ − sₖⱼ) + (wₜᵢ − wₜⱼ) + δ
k identifies the scene, t the recruitment wave, and δ the first-option effect. Partial pooling keeps small scene samples from overreacting. In five-fold grouped cross-validation, mean held-out log loss improved from 0.669 to 0.633; the final sampler had R̂ = 1.00 and no divergences.
The 95% ranges come from 50,000 whole-ballot bootstrap resamples, preserving the dependence among one reader’s choices. The initial analyzed charts use 674 ballots and 7,919 head-to-heads. Retakes and later writes from the same anonymous browser never entered the counted totals.
50,000 / 50,000
whole-ballot reruns ranked Fagles first
5.5% lower
held-out prediction error after scene adjustment
R̂ 1.00 · 0
maximum R̂ and sampler divergences; minimum effective sample size 1,230
Below are the frozen snapshot, analysis settings, downloads, and reconciliation checks.
Aggregate downloads
These are the rankings and model diagnostics used in the report. The files contain aggregate results only. They do not include ballot IDs, IP addresses, or individual ballot records.
The original report and July 26 update use separate frozen sources.
July 21 Report
Cutoff: July 21, 2026 at 6:30 p.m. ET
July 26 Update
Cutoff: July 26, 2026 at 9:30 p.m. ET
Shared analysis settings
Regularized and scene-adjusted Bradley-Terry models; 50,000 whole-ballot bootstrap runs.
For interviews, data and methodology questions, publication-ready charts, or factual corrections, contact Kyle Cureau at kyle@scrivium.com or text Kyle at 415-390-6475 and include your outlet and deadline.