“Storms make trees take deeper roots”: The impact of Dolly Parton’s death on Wikipedia

· Diff ·

5 min read Original article ↗
Dolly Parton accepts Woodrow Wilson Award in 2008.
Dolly Parton accepts Woodrow Wilson Award in 2008.

At differing times on 25 August, the world heard news they didn’t quite believe. Singer, songwriter, literacy campaigner, and overall icon Dolly Parton had died. Many of us did what we do when we want to be sure: we went to Wikipedia to verify the news. Every day, that reflex pulls hundreds of millions of readers to the site. Almost always, they arrive at Wikipedia articles that load, no matter how sudden the surge in interest. For about forty minutes on 25 August, the page did not load for some of our readers.

At around 18:18 UTC, four of the seven database servers in the Foundation’s Ashburn datacenter that hold cached versions of parsed Wikipedia pages became unavailable. Readers routed through Ashburn (much of Europe, Eastern North America and South America) saw errors and long timeouts. At its peak, MediaWiki was returning around 10,000 error responses per second.

Dolly Parton, 79, with her decades of work, was too much for any one page to hold. Editors have been trying anyway. Her English-language article had been edited more than 6,000 times before her death; volunteers made hundreds more in the hours after.

Wikipedia’s infrastructure has sustained similar events before, and mostly quite well. When Michael Jackson died in 2009, the sudden concentration of readers on a single freshly edited page overwhelmed the servers and their caching mechanisms. Engineers named the pattern the Michael Jackson effect. In its wake, they built PoolCounter, a lock manager that stops many servers from doing the same rendering work at once. When Prince died in 2016, PoolCounter did what it was designed to do; readers saw a rare error message rather than an outage. In the years since — through papal elections, Queen Elizabeth II’s death, and other moments that briefly focused global attention on a single article — the mitigations held.

So what changed with Dolly?

Wikipedia uses an extension called FlaggedRevs — known on the English Wikipedia as Pending Changes — to let communities review certain edits before they appear to readers. It uses PoolCounter’s coordination in a way that was not fully aligned with how PoolCounter was designed: under enough load, that misalignment can transform the protection into the very database storm it was built to prevent. The specific interaction had gone undetected for eleven years. More than 4,000 English Wikipedia articles use Pending Changes, some of them drawing hundreds of thousands of readers a month. But being widely read was not enough to trigger the flaw. It needed a particular set of conditions to line up at once, and until 25 August, they never had. 

On 25 August, both the English and German articles on Dolly Parton had FlaggedRevs active. Both were being furiously edited. Their cache keys landed on the same servers. The interaction that had never surfaced had now done so.

What happened next is the part of the story that tends to get less attention. Alarms fired at 18:20 UTC. The SRE team knew who to page. The right engineers within the MediaWiki team were quickly identified, and the root cause was found. Within six days, the incident report was written and made public. The system worked because the people around it had built processes for exactly this kind of moment.

The report itself is candid about a longstanding question. FlaggedRevs is a complex extension that has lacked a dedicated maintainer for years. A code stewardship review was opened in 2018; in 2022, active work on it was paused pending clearer organisational priorities. A community wishlist filed in August 2024 asked for regular support of the extension — a small but committed group of contributors flagging the gap before it caused an incident. Within hours of the Dolly outage, the Foundation’s Service Ownership working group opened its process for FlaggedRevs. The Moderator Tools team, which works on the tools volunteer patrollers and administrators use to keep content reliable, has since taken ownership of the extension, and will be responsible for maintenance, refactoring and bug triage where needed, going forward. 

FlaggedRevs is not the only piece of software whose ownership has been unclear. For the past few years, the Foundation has been working on a question that sounds simple and isn’t: who is responsible for what? The answer is taking shape as the Service Catalog, an inventory of the software components that keep Wikimedia projects running, each with defined ownership roles and responsibilities. The catalog runs to hundreds of entries and was published last week. Its length is part of the point: a codebase built up over more than two decades, by staff and volunteers alike, does not map itself. 

From a reliability perspective, the Service Catalog is a major milestone for our Engineering organization. It will allow us to resolve production incidents faster, as well as ensure all components are sufficiently covered and maintained.

In the case of FlaggedRevs, this appears to be the best outcome we could get out of the incident: a sustainable owner with expertise when we need it most. 

Dolly used to say storms make trees take deeper roots. Wikipedia roots are no different.

Can you help us translate this article?

In order for this article to reach as many people as possible we would like your help. Can you translate this article to get the message out?

Start translation