Wikipedia:Village pump (WMF) - Wikipedia

265 min read Original article ↗

Discussion page for matters concerning the Wikimedia Foundation

The Wikimedia Foundation (WMF) section of the village pump is a community-managed page. Editors or Wikimedia Foundation staff may post and discuss information, proposals, feedback requests, or other matters of significance to both the community and the Foundation. It is intended to aid communication, understanding, and coordination between the community and the foundation, though Wikimedia Foundation currently does not consider this page to be a communication venue.

Threads may be automatically archived after 14 days of inactivity.

Behaviour on this page: This page is for engaging with and discussing the Wikimedia Foundation. Editors commenting here are required to act with appropriate decorum. While grievances, complaints, or criticism of the foundation are frequently posted here, you are expected to present them without being rude or hostile. Comments that are uncivil may be removed without warning. Personal attacks against other users, including employees of the Wikimedia Foundation, will be met with sanctions.

« Archives, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17

Note from Editing Team: Please see #Context from Editing Team for details on how this experimental feedback-requesting feature is intended to work.

When Simple Summaries rolled out, the overwhelming response from the community was that we do not want AI features. One year later, we are now getting more AI features, e.g., AI-generated edit suggestions. This has been in the works for about 2 months and is due to be announced any time now, so you heard it here first. This is going to be long, sorry, there is a lot of ground to cover.

First: These appear to be the suggestions, because based on how hard that URL was to dig up I assume the announcement wasn't going to link to them in one place. (It's unclear which of these lists if any is the actual list used in production, but there do not seem to be major quality differences based on spot-checking several of them, the newer lists are not obviously better and the larger lists are not obviously spottier). There are three broad problems here, besides the fact that adding new AI features is the opposite of what people have asked for:

  1. Besides the most obvious low-hanging fruit (typos), the suggestions do not contain any concrete fixes. I am pretty sure this is due to not wanting to create a vector for adding AI-generated text (from the study: We don’t have any plans nor desires to use models to create edits directly), and I agree with that. The problem is that, based on the kinds of suggested edits we've seen in the past (e.g., Newcomer Tasks), people will resolve this ambiguity by using AI anyway.
  2. To that point, it is unclear who this is for. A lot of the suggestions are obvious low-hanging fruit that competent copy editors would have noticed on their own anyway. People who are not fluent in English will not be able to do much with a suggestion like "Rephrase the sentence to present the information in a neutral tone and qualify the superlative with a time reference." Newcomers who might have been overwhelmed will probably be even more overwhelmed because of the lack of direction.
  3. Several of the suggestions are just bad. Sorry, but there's no kinder way to put that: they are bad suggestions that encourage bad edits. This is just a spot check and is incomplete, but I've identified a few common categories of bad suggestions (drawn from multiple lists at the link):

Political/geopolitical nightmares: I strongly suspect the geographical name suggestions are going to or have run into some geopolitical snarl at some point. But the NPOV suggestions already have, and demonstrate the reason why using AI to "correct" non-neutral point of tone is the opposite of "low-risk" (as described in the writeup) and generally a bad idea.

  • FreedomWorks: This article about a conservative organization contain(ed?) the sentence "During the 2020 election campaign, FreedomWorks pushed false and misleading claims about mail-in-voting, targeting ad campaigns on swing states with high concentrations of minority voters." It is cited to a Washington Post article describing clearly false and misleading claims. The AI's suggestion, however, is The original wording presents FreedomWorks' actions as definitively false and misleading, which is non‑neutral. It should be rephrased to a neutral description of the disputed nature of the claims.
  • Yishuv: The original article contains the following sentence: "The League of Nations codified support for the eventual 'establishment in Palestine of a national home for the Jewish people' into the foundational document of the British Mandate in Palestine, thereby facilitating what detractors later regarded not as aliyah, or the flight of refugees from Nazi and Fascist atrocities, but as the Zionist colonization of Palestine." Obviously that's not perfect, but this suggestion seems to be a misinterpretation: The passage uses loaded language that frames the British Mandate support for a Jewish national home as ""Zionist colonization"", reflecting a partisan perspective. It should be rephrased to present the differing interpretations without endorsing one.
  • The LLM is oversensitive to even the most factual descriptions of political views or affiliations. Example: DeAndrea G. Benjamin, re: the sentence "During her confirmation hearing, Republican senators questioned her decisions granting bond and early release of defendants": The phrase ""Republican senators"" introduces a partisan label that is unnecessary for a factual description of the hearing. The senators who raised these questions are factually members of the Republican Party, all three sources frame it as such, and the partisan breakdown is an inherent fact of the situation.

Suggestions to introduce factual errors:

  • Frost Children, regarding a sentence about the album called Smile! :D: The sentence contains an emoticon, which is non‑neutral and informal. If someone followed this suggestion they would turn a correct statement into a wrong one.
  • 1797 in Denmark: The word "pyusician" is a misspelling; it should be corrected to "musician". The actual correct spelling here would be "physician".

Parsing errors producing nonsense: This is the same problem Simple Summaries had. The text parsing breaks in many ways, and the LLM generates suggestions based on the broken version.

  • The LLM has trouble with wikitables, and will often directly reference a JSON snippet it received, resulting in bizarre suggestions like Remove the nonsensical JSON list and replace it with a brief, readable statement or omit it entirely. (Markazi Jamiat Ahle Hadith)
  • The tool seems to assume that excerpts are one full sentence long, even when they're not. This results in several suggestions to "break up" sentences that already are. This happens a lot, but an illustrative example is Santa Fe Place (the "Stores" paragraph), where the actual prose problem is the opposite as Overly long sentence with many clauses; needs to be broken up for simplicity): the sentences are choppy and some could be combined. (also there's an obvious comma splice the AI fails to mention)
  • Sometimes titles, sidebars, etc. get interpreted as article text: 2019–20 Philadelphia 76ers season: The lead repeats ""NBA professional basketball team season"" twice, creating redundancy. Obviously, it does not; the culprit is that the sentence the LLM interpreted was "NBA professional basketball team season NBA professional basketball team season The 2019–20 Philadelphia 76ers season was the 71st season of the franchise in the National Basketball Association (NBA)." (See also Boadicea Haranguing the Britons, where it does this for the template)
  • This also happens with templates, as in Chen Lijun (actress): Redundant repetition of the subject’s occupation; the phrase “Chinese” appears twice. The first "Chinese" comes from the "lang-zh" template.
  • Cross-wiki links get it too, as in Switzerland in the Eurovision Song Contest 1966: Extraneous ""[it]"" markup after Mascia Cantoni; it is an artifact of extraction and should be deleted.
  • So do blockquotes, as in Nacht und Nebel: The passage contains a run‑on sentence and an incomplete citation phrase ""According to historian Wolfgang Sofsky:"" that leaves the reader expecting a quotation.
  • So do stub templates: The stub template line includes an unnecessary ""vte"" fragment and could be phrased more cleanly. (the "vte" shows up a lot) or Remove the redundant stub messages that appear as ordinary text at the end of the article.
  • Some suggestions are already fixed. For instance, Bomb-making instructions on the Internet was (very obviously) vandalized in Special:Diff/1358132635, and the vandalism got reverted by ClueBot basically immediately. The suggestion nevertheless refers to the vandalized version (Remove the non‑encyclopedic, opinionated rant). I don't know how the LLM got hold of a revision that existed for only a few seconds.

Suggestions to conceal the symptoms of a larger problem:

  • As seen above in the bomb-making example, the LLM doesn't seem to know about vandalism and will never describe it as such, regardless of how obvious it is. This seems likely if not certain to encourage someone to "fix" the tone of vandalism without addressing the actual claim, which is how we get years-long hoaxes.
  • One thing that happens fairly frequently is that the suggestion feature will flag an article that, in context, is clearly AI-generated. (Examples: Jeremy Coller, Vimbuza). It never picks up on this and suggests minor tweaks to wording that would just put a band-aid on the issue. (The csv is actually fairly useful for this, but only in full searchable list form.)

LLM-specific tics: Obviously the suggestion text itself are AISIGNS overload but what I mean here is that LLM edit suggestions/summaries have some consistent quirks. I haven't done an in-depth look for them but two known ones do show up frequently:

  • For some reason LLMs have a fixation on "superlatives" being inherently non-neutral, when sometimes they are just true. I don't know where this comes from -- WP:NPOV doesn't mention anything about superlatives -- but it shows up all the time in LLM-generated revision suggestions, and also all the time here. For instance, in List of Hercules: The Legendary Journeys and Xena: Warrior Princess characters: The description of Hercules is too lengthy and includes biased phrasing such as ""strongest man in the world"" [...] Or on Trinidad, California, a sentence starting "On December 31, 1914, the largest recorded ocean wave ever to hit the United States West Coast" (in a paragraph with 5 citations) is criticized with The sentence makes an unqualified superlative claim about the wave’s size, which could be seen as non‑neutral. Adding an attribution phrase mitigates this.
  • LLMs will over-justify anything. The following reads like a parody but it is actually a suggestion from these: The word 'ithe' is a typo; it should be 'the'. Fixing this error improves the sentence's correctness.

I don't really know what to say at this point. Based on the convoluted phabricator issue snarl it seems that there was some human review done at some point, but nevertheless it took me only about ~1-2 hours to find the above issues, and that was only a spot check. This seems like a reasonable amount of human QA to expect before pushing an editorial feature to prod. It also seems reasonable to expect a full audit of potentially controversial subject matter (by "full audit" here I mean even just the results of CTRL-F "Israel", "Palestine", etc.), and ideally a review by people who are familiar with AI-generated revision suggestions and the general areas in which they get things wrong. (None of this deviates much from the thousands of similar justifications I've seen in AI-generated edit summaries.)

I also think the whole premise is just flawed. LLM and machine learning tools are not only going to have high rates of false positives, but false positives that take time to evaluate. (For instance, there are various LLM tools to scan articles/edit summaries for possible issues, but the point is that they generate lists to be manually reviewed later.) They do not scale to a scenario like automated edit suggestions, where the assumption is that the suggestions are pre-vetted and can be evaluated quickly. Gnomingstuff (talk) 17:54, 15 August 2026 (UTC)Reply

@Gnomingstuff I get your frustration, and agree with some of your points, but I still think this is worth a try. I will often ask the LLM-bots to proofread my articles. Some of the suggestions are obviously good (mostly low-level stuff like spelling, repeated words, etc). That's the kind of stuff that once you've read your own writing 100 times, you read right past and don't notice, so I find it an invaluable service. The higher-level suggestions (tone, phrasing, flow) I'm much more likely to reject, but I accept them often enough that it's worth doing. But that's really no different from when I'm working with a human reviewer. I'll often push back and say "Nah, I think the way I've got it now is fine".
So I think the trick here is to figure out how to educate people that these really are just suggestions and they need to apply their human judgement about whether to accept them or not. You are correct that for new editors, that may be problematic. Still, I think this is something worth trying as long as we monitor how well it's working out and be willing to pull the plug if it turns out to not be useful.
LLMs are just the most recent technology step between scribes writing on clay tablets and where we are today. We can dig in our heels and chant "LLMs bad, down with AI, all power to the humans!" Or we can experiment with them (inevitably with some failures) and learn how to take the best advantage of them to improve our product. I vote for the latter. RoySmith (talk) 18:17, 15 August 2026 (UTC)Reply
Please don't ping me to a discussion that I started less than an hour ago and am clearly aware of.
I don't think that We can dig in our heels and chant "LLMs bad, down with AI, all power to the humans!" is a fair assessment of something that I spent actual time looking into. Gnomingstuff (talk) 18:30, 15 August 2026 (UTC)Reply
The question is whether this will help new editors to develop good judgement. LittlePuppers (talk) 04:55, 16 August 2026 (UTC)Reply
Please just throw me in a ditch. Polygnotus (talk) 18:48, 15 August 2026 (UTC)Reply
Per WP:TALK and the header of this talk page, please avoid fact-free rants and aggressive exclamations, and try to contribute actual arguments instead. Regards, HaeB (talk) 07:01, 16 August 2026 (UTC)Reply
Fixing this error improves the comment's correctness.  — Hex • talk 15:46, 16 August 2026 (UTC)Reply
The above exclamation was not nearly aggressive enough. Let me rephrase for clarity: This is one of the worst features I have ever seen proposed for anything, ever. Enabling a feature like this is literally insane. –jacobolus (t) 20:10, 29 August 2026 (UTC)Reply
At the very least this specific feature will be community configurable via MediaWiki:Editcheck-config.json/Special:EditChecks once it goes live so we can turn it off. But yeah, my first reaction is the same as Polygnotus'. * Pppery * (alt) in solidarity 18:58, 15 August 2026 (UTC)Reply
Kill this with fire, and fire whoever wanted to impose this upon us. Dishraceful and going against clearly expressed community sentiment. The Wmf should not produce any tools that make content suggestions ever, this is not what they zxist for. Fram (talk) 20:08, 15 August 2026 (UTC)Reply
Kill it with fire. The only good thing about this is that it appears we have the ability to turn it off. Tazerdadog (talk) 20:54, 15 August 2026 (UTC)Reply
I fear the suggested edit feature has shown that new editors have a bad tendency to follow suggestions blindly. This isn't the fault of new editors, but poor explanations of what is being suggested and that they are only suggestions.
Looking at the example above make me think this will only make the situation worse, especially as the LLM shows that it doesn't understand policy (a common problem for ever LLM). -- LCU ActivelyDisinterested «@» °∆t° 20:59, 15 August 2026 (UTC)Reply
@ActivelyDisinterested, @Kowal2701 and Gnomingstuff, and others in this thread, you are looking at a feature in it's pre-pre-pre-alpha stage. What the ticket tells you is that the Editing team is preparing to deploy a very very early experimental version of the feature to experienced editors who have opted into enabling a suggestion mode beta and append a specific parameter to the URL (i.e. basically nobody will get this feature unless they specifically click a link and have a very specific beta preference enabled). The code is being enabled so that it can be demoed to folks, used to perform rudimentary qualitative A/B tests and gain very preliminary feedback from Wikipedians at conferences (which occurs before wider consultations with the community). This is nowhere close to being deployed anytime soon without significant bug fixes and the call-to-action in the thread, "is due to be announced any time now" is just patently false. For what it's worth, I personally haven't made my mind up about this specific feature, but I'm willing (and would strongly urge other folks) to provide the team with the ability to spend some more time atleast trying to iterate and experiment on the feature to see if some variation of it could be made useful to some Wikipedians. Sohom (talk) 22:50, 15 August 2026 (UTC)Reply
I realize that this is an experimental version of the feature, but based on the actual content that exists, this isn't the "show to editors and assume they like it" stage, it's the "internal minimum-viable-product demo" stage -- and a MVP you'd need to very carefully babysit to make sure something like The current content is a series of JSON objects that do not convey readable information to the reader. doesn't pop up onscreen. Arguably it's not even that, but the stage of "go back to square one and rethink because the premise is inherently flawed."
I think that "due to be announced any time now" is a fair interpretation of there being an August ticket called "Announce availability of 'experimental' suggestions" with the description The announcement we publish ought to equip volunteers with the info. they need to answer the following questions.... Like... it's due to be announced. That's... what the ticket... says.....
  • "Assume" is not my wording, it's directly from the ticket: In T428311 and T431376, we – staff, in collaboration with experienced volunteers across a range of Wikipedias – will have assumedly determined the initial batch of LLM-generated MoS suggestions to be reliable.
Gnomingstuff (talk) 06:46, 16 August 2026 (UTC)Reply
Let's nip this in the bud please before it becomes a fait accompli (if it hasn't already). Incredible that there's been no community consultation about this AFAICT Kowal2701 (talk, contribs) 22:35, 15 August 2026 (UTC)Reply
I'm somewhat interested in what an llm could dig up in a widespread analysis of MOS:GEO, but unfortunately the suggestions for MOS:GEO are not really about MOS:GEO but are normal typos and (misunderstood) context suggestions. This may be in pre-alpha, but it is frustrating to read "we’ve been successful in developing bespoke, one-off models that surface specific kinds of editing suggestions in a reliable way. For example, we use the Add-a-Link model to suggest relevant inline links between articles" when there have been deep flaws to the add-a-link model that have been unaddressed since its implementation. It is also known that the revert metric used is flawed, so it is disappointing to see it still being referred to. (I recently provided an example to WMF devs of a tone check edit making the article more promotional, but I don't know if that's an edge case or a more widespread issue like add-a-link has.) The "tools that show promise" user story is also quite cheeky; I can't decide where that lies on the amusement to annoying scale, it could be seen as endearing.On the current suggestions, there is a mix of "valid and useful" and quite wrong. The valid and useful ones I saw were mostly typo suggestions. The MOS:GEO ones that went beyond that were sometimes nonsensical. The NPOV ones I checked I would avoid suggesting. The metric being used to assess readiness, in this case "The size of this vetted set of suggestions has given the team the confidence", needs to be relooked at. A consideration that seems to be lacking from all these suggestion ideas is that putting any of these into a formal structure gives them an imprimatur of authority. That's tricky to work around, but the "accelerate the speed with which we can surface meaningful signals" language suggests it isn't a strong consideration. "low-risk edit suggestions (as defined above)" is another metric that needs to be reassessed, if you're trying to touch upon NPOV you have left low-risk behind. If it is true that "it can take more than a year to produce a single type of suggestion", it does seem like far too much effort given the quality of the results. I hope the experimental team will have another think about the fundamental assumptions here and the metrics used for assessment, to help shape future development. CMD (talk) 23:42, 15 August 2026 (UTC)Reply
The idea itself is interesting, and I wouldn't be against experimenting with AI as a way to surface article quality issues (e.g., what EditCheck is doing). This seems to go much further, with the model presenting specific, actionable suggestions, although it stops short of ready-to-post edits.Is this still experimental? Absolutely. As @Sohom Datta points out, this is about to be deployed as "experimental suggestions" open for further feedback, to a testing audience clearly distinct from its target audience. I doubt the kind of editor knowledgeable about beta features and actively seeking out to test this will be misled by the AI's suggestions.I will concede that there is an ambiguity in the way the ticket presents the matter, which is not ideal: the purpose of these edit suggestions is to both edit more effectively with tools that show promise and contribute to making those suggestions more reliable by using and evaluating them in real editing contexts. This, while technically true, should be clarified to shift the emphasis towards the latter.Now, what gives? Of course, no one wants the current version to be shown to newcomers, given the major flaws pointed by Gnomingstuff and others above. What we can do, however, is twofold. Now, discuss whether this feature could be developed to provide constructive help in theory, not considering its current lack of readiness. This is an open question. And later, once the feature is sufficiently mature and ready for a rollout to newcomers, discuss whether that current state is worth rolling out, whether it needs further development, or if the project should be cut short. Chaotic Enby (in solidarity · talk · contribs) 23:42, 15 August 2026 (UTC)Reply
But we've done this so many times where features seem to never get cut short after a certain point in development regardless of feedback, it's only the rare occasion when enough people kick and scream Kowal2701 (talk, contribs) 00:29, 16 August 2026 (UTC)Reply
Agree, and this is why the sunk cost fallacy has to be considered. In fact, I can see the opposite outcome from this initial discussion, namely the developers being left with the impression that all the issues the community has are with the current state of the project, and that investing more will make it worthwhile and gain community acceptance. This is in fact far from obvious, and why we should, I believe, center the discussion on the viability of the project as a whole. Chaotic Enby (in solidarity · talk · contribs) 00:40, 16 August 2026 (UTC)Reply
It's just very difficult to have a constructive discussion/approach when most are worried/anxious they're going to end up ignored and powerless Kowal2701 (talk, contribs) 01:14, 16 August 2026 (UTC)Reply
Pointing down to Peter's comment below, but basically: these specific suggestions were done in a way to try to avoid sunk costs. The development effort on Editing's side has been quite low (gerrit:1299651 + gerrit:1320209) and has mostly been about setting up a framework for a suggestion-type that can ask an API to hand us fairly arbitrary suggestions on an article and then for us to log feedback from a user about whether they seem valid. The broad idea was to make them available to people who knew how to toggle past multiple layers of "are you sure? this is a beta / experimental", and gather data about which specific instances were considered helpful and which weren't. (Screenshot below as well, but we're also not showing the content-specific suggestions you can see in the raw data file; we just show the static_description field for each one, so you just get the "this might violate MOS:GEO" level of prompting about it...) DLynch (WMF) (talk) 13:30, 16 August 2026 (UTC)Reply
I think the first wave of responses show useful feedback on what instances are helpful, what aren't, and what need extensive tuning to filter out tricky cases and focus on the most useful / confident / low-risk suggestions. Glad that NPOV is being dropped; that's extremely complex and contextual.
A good recurring issue that won't show up in spot checks of individual suggestions, is that articles overall deserve article-level checks before spending time fixing small details. @Gnomingstuff put this well above. I would be interested to see a version that allows article-level suggestions (e.g., for tags that might apply to entire articles or sections), especially for new or single-author articles. – SJ + 22:53, 18 August 2026 (UTC)Reply
I am curious what you mean by whether this feature could be developed to provide constructive help in theory. My experience as an engineer is that these sorts of discussions are not productive. You simply make things, you learn a lot along the way, some of the stuff you scrap, other stuff ends up being very useful, but not towards the goal you originally set out to achieve, and on occasion you actually end up producing what you originally set out to do. Discussions about what sort of engineering efforts would theoretically produce good products are simply a waste of time. Czarking0 (talk) 15:52, 16 August 2026 (UTC)Reply
Hey all -- I'm Marshall Miller, director of product at WMF (this project is with the teams that I work with). I'm commenting to let you know that we see this and that members of the Editing team will be able to comment with more detail, background, and clarifications this week.
But yes, let me first say that this is the very earliest stage of testing/trying/experimenting with this idea, and just for experienced editors. As we have done for all the edit check and suggestion features so far, we will only advance this feature farther if communities are supportive, if the suggestions are reliable, and if the data shows that they make a positive difference for the wiki. And all these checks and suggestions are configurable by communities at Special:EditChecks.
About why we're pursuing this: we have seen good success and community support with edit checks and edit suggestions, and the idea of suggestion mode. So far, these have run off of either simple logic ("this blob of text was pasted from ChatGPT") or small machine learning models ("this sentence uses peacock words"). What they all essentially do is point out to human editors when they are violating a wiki's policies in some way -- and they have been shown to reduce revert rates and make newcomers more successful. But there are a lot of wiki policies that could be useful to point out to people. LLMs are a new technology, and they can and do make mistakes. It takes careful testing and tuning and evaluation to get them to perform reliably and may not even always work out (and it may not in this case either). But they may make it possible to produce more of these useful checks that help newcomers make better edits, and may help experienced editors notice things that need fixing. We will 100% need the input of all of you to help us figure out together whether we're on to something or not.
Our approach to all this is that editing decisions should be made by humans (except for the very simple kinds done by things like ClueBot, etc). And that features like these try to help humans notice/find places where they could apply their judgment.
Okay, anyway -- more to come from team members who are deeper in the details. MMiller (WMF) (talk) 04:03, 16 August 2026 (UTC)Reply
WMF A/B tests are notoriously unreliable, and invariably interpreted in the most positive light possible. See e.g the image viewer disaster, or the initial claims about edit check where the posted positive results turned out to be false. Why should we trust whatever results will be posted this time? More importantly, why is such a tool created when the WMF should know by now the massive pushback they would get against AI content suggestions? Aren´t there enough other improvements requested (often for many years already?). Fram (talk) 06:51, 16 August 2026 (UTC)Reply
So far EditCheck has produced amazing results (e.g. vastly increased the share of newcomer edits that contain citations) – and the Editing team listened a lot to community members while developing new checks/suggestions. But you don’t have to trust anyone given that all suggestions and edit checks can be enabled and disabled by local admins. Johannnes89 (talk) 16:56, 16 August 2026 (UTC)Reply
Yeah, I think I meant referencecheck (or whatever it is called), not editchecks (which I haven't checked), should have been more careful in what I said. They claimed a serious number f added references, but it turned out that a lot of edits were tagged as "reference added" when this wasn't true, and a lot of other "references" were completely invalid but counted as a success anyway. I posted this with clear examples, but the WMF ignored this completely. But that's about reference adder AB tests, not edit check, so again, I should have checked before posting. Fram (talk) 09:41, 17 August 2026 (UTC)Reply
From glancing at the list linked above (totaling 349 suggestions - 22 labeled MOS:GEO, 116 labeled NPOV, and 211 labeled "simplify language"), here are my thoughts:
MOS:GEO - most seem fine at a glance. Lots of suggestions to fix diacritics and spelling. Several which are well outside the purview of MOS:GEO. Seems lacking in nuance in some edge cases (but saying more confidently would require some fact-checking).
NPOV - lots of issues. It is too timid to say anything forceful, even when it's warrented and supported in RS, and wants to add qualifiers (e.g. "reportedly") for simple statements of fact, such as "improved quality of air" or "top of the chart". It seems generally opposed to any words which are not incrediby boring, and even some which are: perilous, successful, unreliable, transparent, vocal critic. It is often very unclear in what it is referring to ("the evaluative phrase", "the promotional claim", "subjective description", "the evaluative language"). Multiple suggestions are to remove language which is not there. It also has a terrible time recognizing attribution, and suggests several times (probably a dozen+) that it be added when it's already there. I've skimmed through maybe half of these, and a majority have issues.
Simplify language - meh. Some are fine. It seems to want incrediby short sentences. Ironically, I also disagree with the one I see where it suggests combining sentences. Most suggestions are pretty vauge. One suggestion is "make this neutral". Also says "American English is preferred on Wikipedia" on a British biography.
Summary of my views: GEO is mostly decent but may lack nuance, NPOV isn't really useful because it lacks the understanding to know when a strong viewpoint is neutral and generally dislikes big words, and simplify language is overzealous in suggesting short sentences. LittlePuppers (talk) 06:16, 16 August 2026 (UTC)Reply
Okay, I was looking at a different file from Gnomingstuff and one which is at least a few weeks old. Take that how you will. LittlePuppers (talk) 06:19, 16 August 2026 (UTC)Reply
Glancing through what is (I think) the latest (and much longer) version, there may be some improvement, but most of my thoughts still apply. LittlePuppers (talk) 06:29, 16 August 2026 (UTC)Reply
The MOS:GEO ones are often not fine. For a start, being outside the purview of MOS:GEO suggests some underlying flaw in the model. "Update the country name to conform with Wikipedia’s geographical naming conventions", I have no idea what that is meant to refer to. "The sentence is amended to specify that the Grand Canal is in Venice, providing clearer geographic information" lacks understanding that the Venice location was established in the prior sentence. "The parenthetical abbreviation after “Guantanamo Bay detention camp” is incorrect and should be removed" is simply wrong, although it is perhaps an unnecessary abbreviation a reader may also not understand. "The place name "South Island of New Zealand" is not formatted according to MOS:GEO; it should use commas between the island and the country" is again just wrong. "The original text mentions "the Atlantic" without specifying that it refers to the Atlantic Ocean, which may cause confusion", not sure what to say about that one, Atlantic Ocean is even written out explicitly earlier on the page. CMD (talk) 06:35, 16 August 2026 (UTC)Reply
Yeah, the set I was looking at initially had a very limited list for GEO. A lot of what you mention reflects broader issues with all the categories as well. LittlePuppers (talk) 06:56, 16 August 2026 (UTC)Reply
Most of these are known issues with LLM-suggested edits, or at least the kind of thing that has certainly been possible to know about for at least a year:
  • The "promotional claim"/"evaluative language" stuff is a 2025-era LLM tic. Very specific verbiage, especially the "evaluative" part, that shows up over and over again in AI edit suggestions and basically nowhere else. Here's a bunch of examples.
  • The "original text mentions the Atlantic" suggestion is another consequence of the isolated-sentences approach; the LLM is responding to the sentence "According to Herodotus they dwelt geographically along the sea south of Libya on the Atlantic," and so it doesn't have the context of first reference/subsequent reference.
Gnomingstuff (talk) 06:56, 16 August 2026 (UTC)Reply
Further context or not, saying "the Atlantic" is not going to cause confusion. CMD (talk) 07:06, 16 August 2026 (UTC)Reply
"American English is preferred on Wikipedia", great so it's not just wrong but will make a bad situation worse. -- LCU ActivelyDisinterested «@» °∆t° 09:07, 16 August 2026 (UTC)Reply
The issue with language suggestions in regard to NPOV is that LLM are not neutral, and do not give neutral suggestions. So using them to make these kind of suggestions is a way of creating a fake consensus. -- LCU ActivelyDisinterested «@» °∆t° 09:11, 16 August 2026 (UTC)Reply
Thanks Gnomingstuff for this thorough demonstration of why LLMs are systems for producing text-like slop. No matter how many attempts are made to patch these behaviors, they will keep happening because no comprehension or intelligence is involved, and never will be. Just guess after guess, a fountain of hot slop staining our precious reputation as one of the few uncontaminated places online.
This shameful and embarrassing effort needs to be canceled immediately and the donation money wasted on it so far written off. I'm not even going to start getting into the unethical nature of using LLMs in the first place, which should have been sufficient on its own to rule out even considering something like this.  — Hex • talk 15:58, 16 August 2026 (UTC)Reply
The sheer quantity of bullshit from the WMF is exhausting at this point. Cremastra (talk · contribs) 04:41, 17 August 2026 (UTC)Reply
This is yet another reason to not trust the WMF and it shows how it's impossible to assume good faith on their part. The WMF at this point is an active threat to the very existence of Wikipedia. Ita140188 (talk) 07:58, 17 August 2026 (UTC)Reply
Dealing with WMF is like living through Groundhog Day. They have too many employees so they bureaucratically create "jobs" building crap that nobody asked to solve "problems" that don't really exist, creating a bigger set of unforseen consequences (because WMF is composed of many software engineers and few Wikipedians and is always and forever tone-deaf to community desires). The volunteers who make the project run are all "power users" to them... Well, here's what the "power users" are saying, "tech bros"...... NO AI ON WIKIPEDIA. Didja get that? Carrite (talk) 14:57, 17 August 2026 (UTC)Reply
Wikipedia has a reputation as one of the last bastions of information, in an era of hallucinated, enshittified LLM-generated slop. Any embrace of AI-powered anything on the platform constitutes a plan to throw that into the bin. ser! (chat to me - see my edits) 11:51, 18 August 2026 (UTC)Reply
HELL NO. AI slop is already a massive problem on Wikipedia and many editors have to waste their precious time getting rid of it. The literal WMF coming in and force-feeding AI slop into the mouths of editors is, to put mildly, disgraceful and shameful. AI is not reliable for anything, especially for use on a perceived beacon of human, somewhat trustworthy (even if it really isn't) content like Wikipedia. We should never have even thought of such a horrible decision, let alone create a bare-bones prototype of it and make it accessible to editors. 🪐Kepler-1229b | talk | contribs🪐 17:17, 13 September 2026 (UTC)Reply
@Kepler-1229b, You should read the rest of the discussion? The literal WMF coming in and force-feeding AI slop into the mouths of editors is, to put mildly, disgraceful and shameful. is not what is happening here based on reading the rest of the discussion. Sohom (talk) 17:21, 13 September 2026 (UTC)Reply
It's a possible outcome. I know that it's supposed to be opt-in right now, but it could very well be possible later on. We need to nip this in the bud before it gets there. 🪐Kepler-1229b | talk | contribs🪐 17:37, 13 September 2026 (UTC)Reply
Any installation on en-WP without consensus from the community sure feels that way, because we'd have to deal with the results even if we choose not to utilize it ourselves. We're already having that problem with the AI-based Revise Tone newcomer tasks. The rest of us have to come through and clean up after the slop. ChompyTheGogoat [ Bleat | Munched ] 01:40, 14 September 2026 (UTC)Reply
I oppose the use of LLMs to write any article content (with exceptions for translations, grammar, adjustments to wiki syntax, or the like; but never for writing any original text), but I think they can be far more useful for the information search part. MGeog2022 (talk) 17:55, 18 September 2026 (UTC)Reply
Which is already fully allowed, because you don't need internal Wikipedia tools to search Wikipedia. Everything is accessible to external programs. Given how difficult navigation is here I do utilize it occasionally to help me find guidelines, as well as sources. ChompyTheGogoat [ Bleat | Munched ] 20:27, 18 September 2026 (UTC)Reply
Oh god no, I hope this can still be stopped. This is just a vector for introducing misinformation to wikipedia articles. TietoTeekkari (talk) 19:43, 20 September 2026 (UTC)Reply

Taking a step back: could AI suggestions be beneficial?

[edit]

As pointed out above, the current development stage is way too early for a broad rollout. This should have been better clarified, both to reassure the community about the experiment, and to provide clearer development goals.

However, taking a step back, a discussion can still be held regarding the potential of this whole endeavor. Would the community, in theory, agree to an AI model surfacing suggestions to newcomers in such a way, assuming the current pitfalls could be smoothed out in development? More concretely, are these expectations realistic, and is it worth investing further resources in this project?

These are questions I don't, personally, hold the answers to. However, we should be discussing them together, alongside members of the Editing team involved in its development (courtesy ping to @Quiddity (WMF)), if we want them to work in sync with community sentiment, and avoid investing resources in dead ends. Chaotic Enby (in solidarity · talk · contribs) 23:57, 15 August 2026 (UTC)Reply

Would the community, in theory, agree to an AI model surfacing suggestions to newcomers. I think a better way to explore this tool would be to make it available to established editors first. People who have the experience and policy knowledge to be able to properly evaluate the suggestions. Maybe the people will say "The suggestions were all spot-on and incredibly valuable". Maybe they will say "Nothing this thing suggested made any sense at all, it's total garbage". More likely, somewhere in between. But let's do the experiment rather than pre-judging it. RoySmith (talk) 00:08, 16 August 2026 (UTC)Reply
I think a better way to explore this tool would be to make it available to established editors first. While it might not have been clear at first, this is, in fact, exactly what the experiment is planning to do. The tool is still in development, and we can't say, in advance, how it will end up in terms of quality. For now, I'm just trying to figure out the proportion of editors who either will find it a non-starter in principle (regardless of the suggestion quality) or are opposed to investing further resources in its development for any other reason. Chaotic Enby (in solidarity · talk · contribs) 00:20, 16 August 2026 (UTC)Reply
Mark me down as "opposed to investing further resources in its development for any other reason". Wikipedia has a large amount of technical debt that would be easy for a WMF developer to fix, but instead they are doing this?
As one example, Arbcom is a very important function, and many people who participate get stressed over whether they are under the word count limit. But the tool that puts a banner at the top of your comment fails to accurately count your words! Worse, you can't invoke it while composing -- you have to post and hope that you didn't go over. And if an arb replies to you inline, that increases your count! This is the sort of thing that a WMF developer could fix in an afternoon. It's important, but making an accurate arbcom word counter will never hit the top of any survey of things many editors want to see fixed.
Another example: what happens if when I sign this comment I accidentally hit the "~" key three times instead of four? How about 5, 6 or 7? This is a typo that happens again and again. How hard would it be for a WMF developer to make it so that you get a "are you sure" message before accepting a signature that is almost always a typo?
There are hundreds and hundreds of these easy to fix things that depend on old scripts written by volunteers and all too often no longer maintained.
I think the WMF should spend a significant amount of developer effort -- 75% or 80% -- fixing these small, non-sexy quality of life issues and only then devote the other 20% - 25% to fun things like AI suggestions. --Guy Macon (talk) 01:02, 16 August 2026 (UTC)Reply
Or, if you don't like the above 4-tilde signature, --Guy Macon (talk) (3 tildes),  --01:13, 16 August 2026 (UTC) (5 tildes), or --01:13, 16 August 2026 (UTC)Guy Macon (talk) (7 tildes).Reply
Guy, on that last one, you don't need to add a signature at all anymore. Discussiontools will do it for you. I haven't added one to the end of this post, for example. In solidarity, asilvering (talk) 01:04, 16 August 2026 (UTC)Reply
Will I (Guy Macon) get an error message for this unsigned post or do I have to make a preferences change to get the autosign goodness? Show preview says it will post the unsigned comment with no error message.
Discussiontools is a nice tool, but it is no substitute for software baked into Wikipedia that checks for common errors and throws up an "Are you sure?" message. I could give you a hundred examples of places where the WMF is depending on unpaid volunteers to maintain basic functions that keep the Encyclopedia running smoothly while focusing on the exciting new stuff --01:30, 16 August 2026 (UTC) Guy Macon (talk)Reply
Guy would need to be using DiscussionTools for that, but unfortunately they are using classic full page source editing where none of those helpful things will occur. DLynch (talk) 13:13, 16 August 2026 (UTC)Reply
FYI I made Module:Word count, I can refine this further if you think it would be useful. It does not 100% solve the issue you mention Czarking0 (talk) 15:57, 16 August 2026 (UTC)Reply
Although I agree with a lot of this sentiment, I think an organizational strategy which places say 20% of the engineering resources on long term tech rather than present issues is reasonable. As long as the development of this feature counts under that I do not see the problem. Czarking0 (talk) 16:00, 16 August 2026 (UTC)Reply
+1, that seems to be what is happening here (see the comment below about 'avoiding sunk costs' and getting feedback early and often from established editors). I can see individual categories of suggestion being useful to experienced editors, focus on making something that works for them before considering anything that might be visible to newcomers. (They already have Special:Homepage) – SJ + 23:14, 18 August 2026 (UTC)Reply
I agree with Fram above in that this feels like a way (though a much more subtle way than Simple Summaries) for the WMF to influence editorial decisions which, with the exceptions of legal reasons, they just should not be a part of. ♠JCW555 (talk)♠ 01:55, 16 August 2026 (UTC)Reply
There isn't really any plans for the WMF to get involved in editorial descision (and I say that as somebody who has through m:PTAC reviewed the Annual Plan). The plan for Edit Suggestions is purely meant as a assistive tool to help editors and kinda comes from the idea of being able to surface gadget/userscript suggestions to everyone without having to have coding knowledge. Sohom (talk) 02:57, 16 August 2026 (UTC)Reply
But the fact that the AI is making these suggestions at all in the first place is the WMF having a subtle hand on editorial decisions in my mind. The examples Gnomingstuff lists above, like the FreedomWorks example, is an example of the AI making an editorial judgement that it should not be doing. Some of the others are more subtle like the DeAndrea G. Benjamin example, but they're still editorial decisions that the WMF shouldn't be engaging in period. ♠JCW555 (talk)♠ 03:23, 16 August 2026 (UTC)Reply
They are using generic models without fine tuning for these suggestions, so out of all the parties that could be said to have made editorial decisions in this scenario I would say OpenAI and Google would rank above the WMF, and I really wouldn't consider either of those companies to have made any editorial decisions... the problem IMO is potentially encouraging uh... not really making any editorial decisions, and vibing through things without considering the context of, e.g. the contents of the sources as we are supposed to. Alpha3031 (t • c) 08:28, 16 August 2026 (UTC)Reply
It is not worth creating AI prose suggestions for newcomers (especially if it takes a year for each model!). Putting aside quality questions, editors here are expected to be competent and able to contribute in English. While there are different ways to be confident, we expect the ability to read and write in English, and thus to some extent to have the tools to be able to copyedit themselves. We also expect editors to be able to read and comprehend our guidelines and policies (pre-emptive clarification, read, not memorise). The best way for us to be able to understand competence in these areas, and thus to be able to assess and offer advice if needed, is to see their edits. Seeing instead a whole slew of new editors making the same llm-prompted changes is harmful to the community in being able to understand and accommodate new editors, and harmful to the new editors in giving them a misleading picture of how things work, and more harmful when the llm doesn't understand our policies and practices (as this one does not). Apologies to Sohom, but "There isn't really any plans for the WMF to get involved in editorial descision" just isn't true if this sort of system is being set up. This sort of suggestion task is the WMF making editorial decisions, even if they're making it through an llm. There are many reasons time would be better spent anywhere else. (I distinguish "prose suggestions" from say the add-a-link task, as that at least teaches a technical competence that we do not expect new editors to have. Such teaching does seem useful, although as mentioned above it would be nice if it was developed further.) CMD (talk) 03:45, 16 August 2026 (UTC)Reply
Seeing instead a whole slew of new editors making the same llm-prompted changes is harmful to the community in being able to understand and accommodate new editors I just want to emphasize this -- the English Wikipedia community is not good at welcoming newbies who make mistakes. That is a problem. The English Wikipedia community is actively hostile to editors making mistakes with large language models. If a newbie puts a poor LLM-based reword into an article, an experienced editor is almost certainly going to WP:BITE them off. And a newbie, operating in a system they're unfamiliar with, with a human-sounding voice telling them "This copyedit is right", is never going to have the knowledge and is almost certainly not going to have the confidence to challenge the AI when it presents them with a bad suggestion. That is going to set them up for failure when a grumpy human editor, burnt out from dealing with LLM edits, callously reverts them.I think using machine learning to help with encyclopedia maintenance is a wonderful thing! And I think large language models are really cool -- but they require a high degree of skill to use correctly in the manner that looks like it's being explored here. And, again, the community is so burnt out from dealing with poor quality LLM content that any editor who uses this tool and (inevitably) makes a mistake will be attacked by the community. Does the team working on this understand that, @MMiller (WMF)? GreenLipstickLesbian💌🧸 05:10, 16 August 2026 (UTC)Reply
Seconding that this could be helpful for all sorts of maintenance but should be aimed at experienced editors while working out kinks.
I would personally like to see rubrics for highlighting potential vandalism and LLM edits. And I want much less text in my sidebar and zero suggested text: just a few words indicating the kind of issue to look for / the kind of style guidelines to check.
And I'd like to see explicit self-evals of each rubric for suitability to the task and for false positive/negative rates, which could also be compiled and developed by community maintainers. – SJ + 23:52, 18 August 2026 (UTC)Reply
@GreenLipstickLesbian -- I think that the most important thing that would prevent against the situation you're describing is that the suggestions wouldn't propose text for the newbie to accept/reject. It would just point out the spot in the article that needs attention, e.g. "Does this sentence need to be rewritten to be easier to read?" -- it would not give them a re-written sentence to add. I know that in that situation, the AI may be wrong about whether the sentence needs to be rewritten, and the sentence may be perfectly fine -- but the newbie might be like, "Hmmm, well I guess I'll reword it?" and they may make a pointless edit, or may make the sentence worse.
Suggestion tasks like "add a link" and "add an image" actually do give newbies specific edits to accept or reject, and we see them generally apply good judgment and be constructive. Yes, many of them mess up and get reverted, but that might have happened to them anyway if they were left to their own devices. So these tasks cause a bunch of things to happen at once, and the question is whether it all adds up to a net positive or net negative, you know?
What do you think? MMiller (WMF) (talk) 21:13, 19 August 2026 (UTC)Reply
@MMiller (WMF) Thanks for the response! Yes, I think these sound like reasonable precautions. I do really want to emphasize the point about making it clear to experienced editors what the newbies actually are seeing is very important.
Having had the beta experimental version of suggested edits enabled for a few days, I definitely see the vision. And I see how it could be very useful. (My response to many of the "consider adding a citation" suggestions has, admittedly, been a)"lol no, I'm removing the unsourced text for other PAG issues", b)"this is cited, it just needs an inline citation", c) "... yes, i see how citogenesis is made", or, d)"... we need an easily accessible essay on 'how to source a statement' that we can link to the newbies here" because wow, i'm having trouble". ) But I also see the biting that happens when newbies do make pointless edits/less than ideal edits ( has some which took more than a few years to rectify), and so I am worried about adding any "they're just thoughtlessly adding machine suggestions to articles"-esque ammunition, even it's it's not strictly true. GreenLipstickLesbian💌🧸 21:33, 19 August 2026 (UTC)Reply
This is why I suggested some basic FAQ type popups upon first edit, including an agreement to not use LLMs, to prevent true good faith mistakes and remove plausible deniability for the rest. ChompyTheGogoat (talk) 17:37, 26 August 2026 (UTC)Reply
RE: especially if it takes a year for each model! to be fair to the teams involved the stated idea behind this specific experiment seems to be wanting to try something that won't take a year per task on what the editing and ML teams think are relatively low risk tasks... I'm just not sure that the community and the teams involved have a sufficiently compatible idea of what tasks are low risk. I think it would be possible to develop something acceptable to the community, but unless very sure about it, it may be best to get a vibe check from the community before thinking something is "low risk", and treating things as "high risk" otherwise. Alpha3031 (t • c) 08:36, 16 August 2026 (UTC)Reply
Would the community, in theory, agree to an AI model surfacing suggestions to newcomers in such a way – I won't, and if that ever happens I'm gone. Models are bias black boxes, and even if suggestions were 100% accurate this would still be an issue. fifteen thousand two hundred twenty four (talk) 04:37, 16 August 2026 (UTC)Reply
If AI had far more quality control, I wouldn’t be opposed to this. Except it doesn’t yet. LLMs hallucinate, and those hallucinations are not something you want informing newcomers who are near-clueless as to how Wikipedia works, let alone people learning English who won’t be able to tell when an AI-based edit suggestion system instructs them to insert a grammatical error or something similar, like the above pyusician —> musician.
This could probably work in theory. But it would assume an AI with a reasonable knowledge of the given article subject, some common sense, and a degree of fluency in English (by which I mean not doing things like pyusician/musician), and, given the above examples provided by Gnomingstuff, this is definitively not that.
I oppose AI-based features being integrated into Wikipedia software, and will continue to do so unless a day comes when LLMs equal or surpass the common sense of a human. Perhaps in a few years LLMs will be reliable enough for this to work smoothly and without issue. With the current state of AI, I don’t think we’re there yet. Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 04:59, 16 August 2026 (UTC)Reply
The AI could work 100% of the time and I'd still oppose it because these edit suggestions are being directed by an AI that's controlled by the WMF, which outside of legal purposes, should never have any editorial control on Wikipedia in the first place. ♠JCW555 (talk)♠ 05:17, 16 August 2026 (UTC)Reply
That’s a fair point; I hadn’t thought of that. (Yet another reason why this is a terrible idea.) Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 05:39, 16 August 2026 (UTC)Reply
The solution for this is making prompts publicly available and allowing each community to tweak them. Alaexis¿question? 06:36, 16 August 2026 (UTC)Reply
LLMs hallucinate, and those hallucinations are not something you want informing newcomers who are near-clueless as to how Wikipedia works, let alone people learning English who won’t be able to tell when an AI-based edit suggestion system instructs them to insert a grammatical error or something similar, like the above pyusician —> musician. - that's a rather odd example to get hung up on, given that automated spell checkers and their admittedly sometimes absurd correction suggestions have been around for decades (long before the term "hallucinations" came to be used in this context). Many or even most professional writers - journalists, book authors, scholars - use them routinely, and have learned to live with the occasional fail of this type (pyusician —> musician).
Indeed, in stark contrast to your logic here, WP:SPELLCHECK has long described them as potentially useful for Wikipedia editors, too, as long as they don't blindly rely on such tools:

Spellchecking software and online tools can be helpful when copyediting Wikipedia articles. [...]

  • No spellchecker is completely accurate. You must check the output of any tool you use. [...]
  • You are responsible for all spelling and grammar changes you make, even if the changes are suggested by an error-checking tool, such as Grammarly or ChatGPT.
Maybe it is time to conceive of Wikipedia editors a bit more as adults who are generally capable enough of using such tools even though their suggestions are not 100% error-free (very few things are).
That said, I do generally agree with your point that AI needs quality control, and folks should definitely ask if WMF has done enough here yet (I'm not sure it has). It's just that - as the spell checker example shows - it's not realistic to demand 100.000% accuracy, or to point to isolated failure cases without assessing how frequent they are. (To be fair, User:Gnomingstuff did already make an informal heuristical effort at the latter with regard to their list above: it took me only about ~1-2 hours to find the above issues, and that was only a spot check, i.e. there is informal evidence that these are not very rare at this point. But ultimately I'd be interested in more concrete assessments of how likely editors using this tool will be to encounter each of those failure cases in practice.)
Regards, HaeB (talk) 06:54, 16 August 2026 (UTC)Reply
Most adults have a lot more experience with spelling than they do with Wikipedia's guidelines. LittlePuppers (talk) 06:58, 16 August 2026 (UTC)Reply
Sure, but that doesn't mean that we keep the edit button away from them.
See Wikipedia:Competence is required, which is perhaps a better known version of this principle that we generally expect editors to be competent adults (metaphorically, with apologies to all the very smart and capable teenage editors among us) that do not require special restrictions and safeguards to protect them against their own mistakes.
Regards, HaeB (talk) 07:30, 16 August 2026 (UTC)Reply
The thing I’m concerned about is everyone knows that spellcheck is obviously not infallible, as the Internet often likes to humorously point out. But to a newcomer, anything that comes from ‘Wikipedia’ seems naturally correct and reliable.
Certainly not everyone would fall into this trap, since, as you point out, it isn’t like all newcomers are to be treated as children who don’t know what they’re doing. But it’s likely that some would; perhaps enough that the consequences of Wikipedia’s reputation for factual accuracy is something we should take into account on this. Given all of that, I don’t think we can draw a one-to-one comparison between run-of-the-mill spellcheck and an edit suggestion feature built into Wikipedia itself.
That’s why I argue that we should hold ourselves to a higher standard; after I’ve given your points some thought, however, I think you are correct that my own request for equal [to] or surpass[ing] human-level reliability is asking a bit much of a literal artificial intelligence. Simultaneously, I think spellcheck-level absurdity is far too low a bar for this feature.
Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 09:11, 16 August 2026 (UTC)Reply
Sounds like something that could be addressed with a user level disclaimer about LLMs before the feature is enabled on their account. Czarking0 (talk) 16:02, 16 August 2026 (UTC)Reply
That would definitely fix that issue (I’m surprised I didn’t think of that, to be honest). Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 00:11, 17 August 2026 (UTC)Reply
I agree with @RoySmith that some suggestions are more suitable for experienced editors. I believe that checking citations is a good use case with an LLM flagging potentially problematic citations and editors verifying them manually (we have a proof-of-concept that I've been maintaining but as long as it's a userscript it won't make a dent in the sourcing problem). This is a task that is not too complex but requires familiarity with WP:RS and some experience. Alaexis¿question? 06:35, 16 August 2026 (UTC)Reply
Based on the suggestions here, my stance is basically the same as what WP:LLM already says: typos and punctuation suggestions seem unnecessary but mostly fine (except when they're not, see the "physician" example), anything beyond that is unlikely to be fine, and suggestions related to NPOV are light years away from fine. 100% accuracy isn't possible but, realistically speaking, people are going to rubber-stamp whatever suggestions they receive.
As far as the spot-check goes, the timeframe is not really all that scientific since I actually saw this last night (not even while doing AI stuff, I was doing Commons new image patrolling and saw one of the screenshots from it), slept on it, and did the full writeup the next day. I also have seen tens of thousands of AI edit suggestions, as well as thousands of snippets of parsed Simple Summary text, so I already knew the places where this was likely to run into issues. This is also why I think the "experienced editor"/"new editor" dichotomy is not helpful, in general. Experienced editors are not necessarily experienced in copyediting and/or the specific problems that crop up with AI, and newcomers are not fragile baby birds who have never edited a piece of writing in their lives. Gnomingstuff (talk) 07:10, 16 August 2026 (UTC)Reply
Automated suggestions could be helpful in very limited circumstances. I can imagine an experienced editor who has chosen to opt in to them assessing each proposal carefully, skilfully rejecting those with the sorts of problem listed above, and improving Wikipedia by acting only on the good ideas. I'm sceptical that AI is the best way to produce such input but I'm willing to be proven wrong. Methods such as this are totally unsuitable for new editors, many of whom will blindly obey The System, make bad edits in good faith, and get reprimanded or blocked. There is also the unoriginal but valid argument that editors would prefer the WMF to devote its resources to other activities, even if we don't always agree exactly on which ones. We're seeing a lot of these "Here's a new toy no one asked for that we've secretly blown this year's development budget on" announcements. I do honestly try to hope for the best, but I'm afraid my initial reaction has become "Oh no, what have they broken now?" Certes (talk) 09:57, 16 August 2026 (UTC)Reply
No, LLM suggestions would be the opposite of useful. They would only reduce the amount of thinking an editor does, leading inexperienced editors to potentially make an edit because the LLM said so. People trust tool output, so if wikipedias own tooling gives them a suggestion, however ridiculous, they might easily do it.
This should be stopped at all costs. TietoTeekkari (talk) 19:48, 20 September 2026 (UTC)Reply

Wouldn't it make more sense to focus on basic copyediting suggestions? I just went to a random article, Mircea II of Wallachia, and requested some copyedits from chatgpt.

suggested copyedits

Lede: “Radu The Handsome” → “Radu the Handsome”. The capitalized The is plainly wrong. � Wikipedia Early life: “first born sons” → “firstborn sons”, unless this is deliberately preserving the exact wording of the translated charter. Since it's presented as a quotation, I'd check Treptow before touching this one. � Wikipedia Reign: “his father's allied stance with the Ottoman Empire” is awkward. Something like “his father's alliance with the Ottoman Empire” is much cleaner. � Wikipedia “forcing the latter to negotiate with the Ottoman court for support, with Mircea II fleeing and going into hiding” is clumsy. → “…forcing Vlad Dracul to seek Ottoman support, while Mircea II fled into hiding.” � Wikipedia “Vlad Dracul signed a treaty with the Ottomans, which stated he would…” → “…under which he would…” or “…which required him to…”. Treaties don't really “state he would” do something. � Wikipedia Same sentence: “as well as allowing two of his sons … to be held as captives” has faulty parallelism with “he would pay.” → “…and allow two of his sons … to be held captive.” � Wikipedia “Radu the Handsome (Fair)” is unnecessary/odd. Unless the source actually uses Radu the Fair, I'd simply remove “(Fair)”. � Wikipedia “Callimachus tells that the leaders…” → “Callimachus states that” or “According to Callimachus, the crusade's leaders…” � Wikipedia “after the defeat, Mircea led the remainder of his unit and the Christian forces across the Danube.” This sounds as though he led all surviving Christian forces. It may be what the source says, but I'd verify rather than copyedit it into a stronger claim. � Wikipedia “allowed the Ottomans to again have control of the fortress” → “allowed the Ottomans to regain control of the fortress”. � Wikipedia The rest of that sentence is especially tortured: “in an effort to retain their support of his having the throne and in an effort to keep his two captive sons safe.” → “to retain their support for his rule and protect his two captive sons.” � Wikipedia “captured by Saxon elite, which were part of the ruling council” → “captured by members of the Saxon elite who were part of the ruling council”. This also agrees with the substantially better wording already used in the lede. � Wikipedia “would also carry out vengeance against the Saxon elite” → “would also take revenge on the Saxon elite”. Carry out vengeance against is unidiomatic. � Wikipedia “Vlad Țepeș would also fight successfully against the Ottomans for a number of years” → “Vlad Țepeș also fought successfully against the Ottomans for several years.” There is a lot of unnecessary historical would in this section. � Wikipedia “Vlad's skilful political maneuverings, his capable administration skills and great tactical thinking, made him…” has an erroneous comma separating subject and verb, as well as some awkward prose. At minimum, delete the comma after thinking. Better: “Vlad's political skill, capable administration, and tactical ability made him a dangerous opponent.” � Wikipedia “…a very dangerous opponent to his enemies” → “…a very dangerous opponent”. An opponent is inherently someone's enemy/adversary here. � Wikipedia

Most of these seem like pretty decent suggestions, don't dip into npov/tone danger zones, and are simple enough for new editors to review and handle. ScottishFinnishRadish (talk) 00:14, 20 August 2026 (UTC)Reply

Also, that might be the first time my first random article click wasn't about a sports ball player. ScottishFinnishRadish (talk) 00:15, 20 August 2026 (UTC)Reply
the nice thing about having an insite spell/grammar checker would be less people would use Grammarly (we could tell people using Grammarly to use the insite one instead, similar to how PasteCheck works). I've found asking a chatbot for corrections gives decent results, but obv you can't defer to them (esp. on content) Kowal2701 (talk, contribs) 19:59, 20 August 2026 (UTC)Reply
Why is it a problem that people use Grammarly? RoySmith (talk) 20:02, 20 August 2026 (UTC)Reply
it goes beyond its scope as a spell/grammar checker and rewrites content, and a lot of people using it don't know it's LLM-powered. It's the worst because people will put time and effort into writing, then put it through Grammarly which turns it into slop. We see it at AINB a fair bit Kowal2701 (talk, contribs) 20:10, 20 August 2026 (UTC)Reply

Context from Editing Team

[edit]

Hi y'all – I'm Peter Pelberg, product manager of the Editing Team. Together with the Machine Learning Team, we've been working on the experimental model-generated edit suggestions. You have raised a range of valid concerns/questions that warrant responses. You can expect those when I'm back online in earnest next week.

In the meantime, I'd like to clarify some aspects of the work and the thinking that's informing it.

First and most importantly: there are no plans right now to deploy LLM-generated edit suggestions as default-on to anyone. The closest thing to a deployment that we've talked about is a potential A/B experiment that would not move forward until we, volunteers and staff, have deemed these suggestions reliable and promising. This commitment is in Phabricator by way of the experiment (T431377) needing T428311 to happen first. For context, T428311 states, "Learn whether experienced volunteers at en.wiki think the experimental suggestions are sufficiently reliable to be shown to newcomers so that we can decide: Will we invest the effort to scale this initial set of LLM-generated MoS suggestions across languages via T431376?"

Now, with regard to what this initial set of 30,000 LLM-generated suggestions is and is not…

  1. Who has access to these experimental suggestions? In this initial, experimental phase, these suggestions will only be available to people who A) enable the Suggestion Mode beta feature and B) install this user script. Note: We will soon introduce a setting within Special:Preferences so people interested in trying out experimental suggestions don't need to install a user script to access them.
  2. What do these experimental suggestions do? This initial batch contains experimental suggestions for three types of improvements/issues (listed below). Each suggestion will highlight the span of text it is relevant to, offer a generic description of the issue, and crucially, ask the experienced volunteers encountering it to indicate whether you think the suggestion itself is valid/useful or not. And if not, offer an explanation as to why. Said another way: These suggestions do not offer fixes or ask experienced editors to make any.
  3. What are "experimental" suggestions? "Experimental" suggestions are a set of – as I think @Sohom Datta put it well – pre-alpha suggestions. The purpose of them is to evaluate their reliability. Currently, there are a total of 6 experimental suggestions available at en.wiki to people who opt-in to seeing them. 2 of these 6 suggestions are powered by machine learning models; the other 4 are based on deterministic heuristics. You can see the full list here: Special:EditChecks#Experimental checks.
  4. What types of suggestions is this LLM generating? This initial batch contains suggestions for simplifying language, rewriting language in a neutral point of view, and adjusting place-based names to follow Wikipedia's geographical guidelines. We are by no means committed to this set of suggestions or to using LLMs in general. We chose this initial set because we found them to be reasonably accurate and assumed they would be relatively low-risk.

Last thing for now: we are aware of the sunk cost fallacy. Thank you for raising it, @Kowal2701. In fact, as I hope the above demonstrates, the whole point of developing experimental suggestions, and making them available in production to experienced volunteers who have explicitly opted into them, is as @Chaotic Enby described: to assess whether they're reliable. Ones that aren't, we will abandon. The ones that are, we'll work together to figure out how and to whom to make them available.PPelberg (WMF) (talk) 05:48, 16 August 2026 (UTC)Reply

as […] “1.” and “3.” demonstrate. This is perhaps not the most important point, but it is bugging me: were the numbers meant to all be “1.”? Comment struck due to this being fixed; apologies. Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 06:08, 16 August 2026 (UTC)Reply
Thanks for responding now and letting us know you don't have time to fully engage immediately, hopefully you will have time to go through all the comments above later in the week as you say. As the process of contingent experiments has been brought up again, perhaps it would help to answer the experimental question plainly, because it can be very easily answered without a complicated multi-month experiment. The suggestions are not reliable, and definitely not reliable enough to show to newcomers (not that this is the only metric that could be considered). To look further, re We chose this initial set because we found them to be reasonably accurate and assumed they would be relatively low-risk, the linked section lacks an assessment of accuracy or risk. Part of the issue here is that a simple glance is all that is needed to identify issues with many the suggestions, so it feels like even that is not being done before these are farmed off to volunteers. There is a mismatch in understanding somewhere. CMD (talk) 06:18, 16 August 2026 (UTC)Reply
Part of the issue here is that a simple glance is all that is needed to identify issues with many the suggestions, so it feels like even that is not being done before these are farmed off to volunteers.
Honestly can't put it any better than that. I do appreciate the response and understand it is the weekend. Gnomingstuff (talk) 07:18, 16 August 2026 (UTC)Reply
Tbh this is a systemic issue to do with WMF governance rather than a reflection on anyone here, where developers and idea labs seem to typically be distant from the projects and communities, and everyone at both ends is left straining to compensate for this. Phab being largely open to the public goes a long way, but I wonder if there could be something like a WP:VPI on meta for developing/screening ideas (obv staff's medium is meetings and private convos etc. and idrk what currently happens, but something like "John and I were talking yesterday about this, X is a potential solution" would be good to get some feedback at the earliest stage). I've previously suggested that staff should be encouraged to do a little bit of editing each week at any wiki they choose (and reduce hours/week a bit for this while keeping salaries the same), but it risks the WMF's status as a 501(c)(3) organization (I was told) Kowal2701 (talk, contribs) 07:57, 16 August 2026 (UTC)Reply
I don't really know that it has anything to do with governance in this case. The issue would be the same regardless of whether the development process was public or semi-public or completely closed off; it's a classic, perennial issue of workplaces. The issue being -- and I am really trying to be as polite as possible here -- that QA is either not being done, or not being done correctly.
From what I understand, the suggestions were assessed for quality via AI, and then someone seems to have reviewed 15 suggestions manually. Unless there's more internal discussion (there probably is, this stuff is nearly impossible to follow), that seems to be the whole QA process.
Meanwhile, I actually have worked in QA. For a task like this, we would be given a spreadsheet like the ones here and would review every single line individually. We caught a lot of errors this way, the final product was better as a result, and it only took a day or two of work at most. I do not feel this is an unreasonable amount of work to be expected from a professional organization. Gnomingstuff (talk) 17:52, 16 August 2026 (UTC)Reply
A stage of basic competitive fault-finding, annotating a large sample by hand to flag and classify problems, goes a long way. Your immediate reflex to do this as the first response above demonstrated this nicely.
If an originating team works in public to catch and patch issues in that way, before others run into them, with high standards for accuracy, that would set discussions like this off on a better foot. – SJ + 00:23, 19 August 2026 (UTC)Reply
Thank you. Sorry, I didn't know who was in the team working on it, I trust all these people. Kowal2701 (talk, contribs) 06:51, 16 August 2026 (UTC)Reply
I oppose any A/B tests of this feature. It is antithetical to what Wikipedia should be, and is not something the Wmf should attempt. Suggestions sent by some central, biased, singleminded entity diminishes the wide variety of input, style, viewpoint, ... that makes the richness of the community. Wikipedia ia a beacon of human diversity, an LLM is the opposite of this. Fram (talk) 06:58, 16 August 2026 (UTC)Reply
In contrast to such hyperbolical rhetorical flourishes, the community has already been widely using AI suggestions sent by some central, biased, singleminded entity controlled by WMF for over a decade, in form of ORES (or now the "revert risk" models). These AI suggestions (on whether to revert another editor's changes as vandalism) are made available to any editor at Special:RecentChanges even though they might be seen as more consequential on average than those of the new tool under development here, and have a fairly high error (false positive) rate. I have personally implemented many thousands (probably tens of thousands) of those AI suggestions over the years, and survived skipping many more mistaken suggestions.
I don't want to entirely dismiss concerns about control by the Foundation though. In case of these vandalism detection AI models, the more recent work on them seems to have been dominated by decisions and aims of the WMF Research department that may not entirely align with the community here (for example, they seem to have foregone possibly substantial quality improvements for English Wikipedia in favor of language equity, and implemented their own conceptions of "fairness" with regard to IP editors).
The main employee behind the success and broad community acceptance of the original ORES left WMF years ago, and his parting recommendations for "Community-centered Evaluation of AI Models on Wikipedia" do by and large not seem to have been taken up by WMF.
Regards, HaeB (talk) 08:36, 16 August 2026 (UTC)Reply
As I said earlier, keeping the prompts (and the pipeline in general) publicly available and enabling each community to tweak them would go a long way in assuaging these concerns. Some competence is required to maintain these tools, but there is no reason not to make prompts and benchmarks transparent. Alaexis¿question? 10:50, 16 August 2026 (UTC)Reply
The system prompts used for the model appear to be published on gitlab here according to the mw:VisualEditor/Suggestion Mode/Model-generated editing suggestions#Research findings. Not sure which size Gemma gemma4:latest is, maybe the E4B? I would be interested to know if other models were evaluated, for example Nemotron 3 Super is a similarly sized model to gpt-oss:120b, but a bit newer. Mistral is dense so might be a bit slow to run. In any case, there are many newer models in the same weight class, surely it wouldn't be that hard to generate say ~1000 from each of them to see if anything is clearly better, given the only thing that has been modified is the system prompt? Alpha3031 (t • c) 12:22, 16 August 2026 (UTC)Reply
Thanks. Going off of that, it appears that no reference information is given, which is not surprising looking at the NPOV suggestions, and that no actual Wikipedia policies/guidelines are included in the prompt. LittlePuppers (talk) 15:24, 16 August 2026 (UTC)Reply
A successful model would likely go beyond a prompt alone and be equipped, at the very least, with retrieval-augmented generation capabilities in order to access the text of the policies/guidelines themselves, and, in the case of NPOV, ground its answer in available reliable sources. I don't think that would be enough for NPOV specifically (there is still too much of a black-box effect in the model's weights, that won't be fully offset by prompting or sourcing), but this is to say that the current approach is far from optimal. Chaotic Enby (in solidarity · talk · contribs) 15:46, 16 August 2026 (UTC)Reply
The particular challenge is, I think, that there seems to be a trend to use increasingly sophisticated models to push increasingly challenging tasks increasingly to newer editors. LittlePuppers (talk) 15:33, 16 August 2026 (UTC)Reply
Regarding point 4., community feedback makes it pretty clear that anything regarding NPOV or tone will not be seen as low-risk, and might be interpreted as an attempt by the WMF to influence editorial decisions. I believe it would help the prospects of the experiment to commit to follow the emerging consensus, which is clearest on this specific aspect.On a broader level, I will reiterate my suggestions on Phabricator on working with the community:

[...] the scope and timeline of any planned community rollout, and the importance of working with the editor community and respecting a future consensus on whether to deploy it, should be made as clear as possible to avoid a repeat of Simple Summaries.

Chaotic Enby (in solidarity · talk · contribs) 12:31, 16 August 2026 (UTC)Reply
I'm with the rest of the group in saying that the NPOV suggestions are high risk and difficult to get right. I'm excited about trying to figure out if a good suggestion model can be developed for simplifying language. There is a large body of research showing that Wikipedia's science content is often way and way too complicated and linguistic complexity is an important part of that. We've had discussions on Wikipedia about using heuristics (sentence lenght, number of syllables per word), which is contentious. Perhaps an AI model can do this better than a heuristic, as it can distinguish between a long sentence with an easy sentence structure and an impenetrable long sentence. In solidarity, —Femke (talk) 🐦 15:27, 16 August 2026 (UTC)Reply
@Femke I have tried both AI models and conventional readability tests like Flesch–Kincaid and this is an area in which current AI tech cannot and should not be used. And my conclusion was that conventional readability tests also suck.
NPOV is another area where AIs cannot and should not be used.
We can use AI to find typos, but fixing them actually requires an experienced Wikipedian. Polygnotus (talk) 15:33, 16 August 2026 (UTC)Reply
I too have plenty of experience using AI for this purpose and find that they can be used successfully in identifying overly difficult text and decently enough for suggesting alternatives if one pays attention to subtle changes in meaning. As these edit checks are only about identifying overly complicated text and because we have an abundance of low-hanging fruit in this area, I'm confident that a tool can be developed. The big question is if within the limits of a certain budget, it can become good enough, and to what extent it attracts the right editors to fix it. In solidarity, —Femke (talk) 🐦 15:40, 16 August 2026 (UTC)Reply
One of my common complaints at FAC (especially for scientific articles) is choppy writing style (sort of the opposite problem from overly difficult text). I often advise authors that they should try using some of the LLM tools to make suggestions for how their writing could be improved. I don't know how well this advice is received. I also don't know how much it is used in the intended way of offering suggestions vs just copy-pasting the output into their article; that would be an abuse of the tool, but people abuse tools all the time and the fact that they do so doesn't mean the tool is bad. RoySmith (talk) 15:43, 16 August 2026 (UTC)Reply
find that they can be used successfully in identifying overly difficult text and decently enough for suggesting alternatives if one pays attention to subtle changes in meaning. Subtle changes in meaning convert text supported by reliable source(s) to text not supported by reliable source(s).
I'm confident that a tool can be developed Yeah, developing a tool is the easy part. The hard part is ensuring it is a net positive.
The big question is if within the limits of a certain budget The WMF is not the right party to create such a tool because they are unfamiliar with the challenges Wikipedia editors face and because they have a tendency to waste a lot of time and effort on creating shiny new tools while neglecting decades of tech debt.
The good news is that volunteer devs can do it, but they will run into the problem that it won't work as I explained above. Polygnotus (talk) 15:46, 16 August 2026 (UTC)Reply
Some teams are unfamiliar with the editing challenges. Other teams have a track record of listening, or have editors in their team. The editing team is an example of the latter. For instance, when people at enwiki and dewiki asked for better VE performance recently, the team managed to fix this tech debt really rapidly.
I'm very willing to help test this part of the tool. I have experience with editors doing this well and less well. In solidarity, —Femke (talk) 🐦 15:58, 16 August 2026 (UTC)Reply
@Femke You can try WP:SCRIPTREQ, the people there are a lot more agile than the WMF is.
Please ping me if you have something I may be able to help test it. Polygnotus (talk) 16:02, 16 August 2026 (UTC)Reply
Tbh I still think we should look at incorporating SEWP articles here as 'simple summaries', but that's a different discussion Kowal2701 (talk, contribs) 15:47, 16 August 2026 (UTC)Reply
I wouldn't assume they'd want us to. Simplewiki is a tiny wiki with just a few active people. If we incorporate their content their vandalism rates will go through the roof. Polygnotus (talk) 15:49, 16 August 2026 (UTC)Reply
They'd also get more good-faith editors though! What I'm thinking is that SEWP still exists as a separate wiki, we just have an opt-in feature for displaying one of their articles behind a button. IIRC @Ferien was sort of open to the general idea (may be wrong) Kowal2701 (talk, contribs) 15:53, 16 August 2026 (UTC)Reply
Adding a CTA, if we get consent from the Simplewiki regulars, may be a good idea. Polygnotus (talk) 16:04, 16 August 2026 (UTC)Reply
Yeah, I am personally quite open to the idea, though I am not too sure what our community as a whole would make of it at the minute. --Ferien (talk) 21:04, 17 August 2026 (UTC)Reply
What kind of typos is AI fit to resolve that WP:AWB is not? Czarking0 (talk) 20:13, 16 August 2026 (UTC)Reply
One thing I'm having in mind are typos that depend on the context to make sense of them (or to whether there is a typo to begin with). Chaotic Enby (in solidarity · talk · contribs) 20:19, 16 August 2026 (UTC)Reply
FWIW, the fixes in Special:Diff/1369578064 were all suggested by Claude. I imagine most of them would have been caught by other tools, but I was impressed by the flagging of Freilberg. I don't remember exactly what it said, but the gist was that it spotted that I had "Peter Freiberg" in one place and "Peter Freilberg" in another. It figured out that these were probably referring to the same person and while it didn't know which was wrong, it assumed one of them was, and left it up to me to figure out which. RoySmith (talk) 20:24, 16 August 2026 (UTC)Reply
And I note the firsts of those fixes is changing a direct quote in a way that is inconsistent with the source. Sure, you can probably justify that (MOS:QUOTE allows minor typographic things to be silently corrected) but I don't think that's something an AI should be recommending in any way. * Pppery * (alt) in solidarity 20:49, 16 August 2026 (UTC)Reply
So, you're saying the fix is correct, and if a human had suggested it you would agree with it, but since an AI suggested it there's a problem? RoySmith (talk) 20:56, 16 August 2026 (UTC)Reply
I'm saying it might be correct (not that it is correct), but determining whether it is is a judgement call I don't want an AI to make. * Pppery * (alt) in solidarity 21:22, 16 August 2026 (UTC)Reply
The AI didn't make the judgment call. It just brought this to my attention and I made the judgement call. That's why my name is on the diff. I'm really not seeing the issue here. I made an error, a tool alerted me to it, and I fixed it. How is this a problem? RoySmith (talk) 21:36, 16 August 2026 (UTC)Reply
The funny thing about using LLMs here is that they start working against each other. The default behavior of AI trying to "rewrite articles into formal encyclopedic tone" is to undo any text simplification (and to do so poorly, for instance replace "is" with "serves as", "uses" with "utilizes", etc). The "text simplification" category here, however, seems to be basically a catch-all. "Simplification" suggestions here range from stuff like Correct the misspelling of the player's first name from "Russel" to the proper "Russell" (which isn't even correct!) to "break up these sentences."
There's also the problem that telling someone to Rewrite the sentence for clarity and smoother flow is not actionable to the majority of people: if someone doesn't know how to copyedit then they don't know how to do that, and if someone does know how to copyedit they don't need those vague directions. It's like prompting people as AIs. Gnomingstuff (talk) 18:23, 16 August 2026 (UTC)Reply
(Ironically, the LLM prompt that judged suggestions reads in part You are a strict reviewer. Your job is to find flaws, not to be nice. Really weird feeling to be envious of an LLM, its opinion certainly seems to be taken more seriously and its tone is given much more leeway.) Gnomingstuff (talk) 20:42, 16 August 2026 (UTC)Reply
Who on earth outside the group of people responsible for this is going to read Gnomingstuff's report and consider those suggestions to be "reasonably accurate"?  — Hex • talk 16:02, 16 August 2026 (UTC)Reply
Pre-alpha or not, the WMF shouldn't be experimenting with ways to funnel freeform model suggestions to editors to begin with. Machine models should never be allowed to so directly influence the contents of the project, as even if the suggestions are individually found to be valid and reliable, there will still exist overall biases. An LLM will favor certain sources, certain topics, certain sides. This will be reflected in what suggestions are and are not made, and editors evaluating and implementing individual suggestions will be entirely blind to any larger systemic issues they would be enabling.
Humans have issues with bias too of course, but this can be counteracted on an individual level by self-awareness of this fact, and on a group level by the diversity of our views. A monolithic model has neither, it predicts tokens. fifteen thousand two hundred twenty four (talk) 21:45, 16 August 2026 (UTC)Reply
"Humans have issues with bias too of course, but this can be counteracted on an individual level by self-awareness of this fact - Citation needed: "making people aware of their bias doesn't do anything to mitigate it." Levivich (talk) 22:45, 16 August 2026 (UTC)Reply
Citation: I made it up from first principals and subjective experience, the preprint will be out soon.[Humor]
I feel entirely comfortable claiming that editors who operate with awareness of their own potential biases will take steps to mitigate them in this structured environment where WP:NPOV serves as a strong guiding force. If you find this unpersuasive, so be it, it is human to disagree. fifteen thousand two hundred twenty four (talk) 23:58, 16 August 2026 (UTC)Reply
The operative word here is can. As human beings that are capable of independent thought and self awareness, we can choose to examine our own biases and seek out information and experiences to change how we think about the world around us and the assumptions we make. Large language models are capable of exactly none of those things, because of the very simple fact that they are computer programs. Claude is exactly as capable of herself awareness as MS Paint is of having an independent thought.
Not every person makes the conscious choice of examining their baises, or is even fortunate enough to exist in a socioeconomic situation to even be able to, but that is beside the point ‑‑gurkubondinn 00:26, 17 August 2026 (UTC)Reply
"Research from Harvard found that the effects of personal interventions such as awareness raising at a personal level are positive, but short-lived."
"And the worst method, the one that actually has no effect at all, is to tell people to be good people, to be egalitarian, and so on. ... It is easy in the sense that it last for a short period of time, but it won't last very long. ... Now, when young people encounter this result, when they see that, yes, they were able to make change, but the change doesn't last, they get very sad, because they want a better world. And I'm not at all sad about that. ... So our brains change, our minds change, associations move around, but they always will gravitate to whatever is your cultural default. And so to bring about actual change, society around us has to change, and then we will move, and then the default will be a new default."
LLMs are biased because humans are biased. Because they're trained by humans, and humans train their biases into the machines. We are no more able to cure bias in machines than we are able to cure it in ourselves. There are plenty of ways in which human intelligence is superior to machine intelligence, but lack of bias isn't one of them (in either direction). Levivich (talk) 02:57, 17 August 2026 (UTC)Reply
Absolutely, I agree with you. That's why the models can never be "neutral" or "unbiased". My point was just that I think you misunderstood 15224's reply, people have the capability to examine and be aware of their biases, but computer programs don't. It's not automatic and it doesn't happen for all people (for various reasons), but we have the ability to examine our own biases because (unlike computer programs) we are conscious beings. It's a bad comparison is what I'm saying, and it anthropomorphises computer programs (that were created by biased people). I have a beef with whoever it was that decided to wrap LLMs in chatbot interfaces. ‑‑gurkubondinn 11:06, 17 August 2026 (UTC)Reply
We're looking forward to getting more deeply into this with you all this week. Before that, I wanted to express gratitude to y'all for the perspectives you are sharing. From the concerns about transparency and community control over models of this sort to issues with specific suggestions you're encountering, please keep the feedback/questions/concerns/etc. coming.
And for anyone who is interested in seeing what these suggestions look like in practice, please do the following...
TRYING EXPERIMENTAL SUGGESTIONS
  1. Ensure you have the Suggestion Mode beta feature enabled
  2. Enable "experimental" suggestions by pasting the following snippet into your common.js: mw.loader.load( 'https://meta.wikimedia.org/w/index.php?title=User:DLynch_(WMF)/alwaysbesuggesting.js&action=raw&ctype=text/javascript' );
  3. Open a page in VE that has one of the 30,000 suggestions available. E.g. https://en.wikipedia.org/w/index.php?title=List_of_examples_of_Stigler%27s_law&veaction=edit
Note: this week, I'm going to see if we can share a spreadsheet so you can see all of the model-generated suggestions in one place rather than having to tap around the wiki looking for them. PPelberg (WMF) (talk) 00:59, 17 August 2026 (UTC)Reply
I wrote a small script so you can view the list of suggested edits on an article without needing to enable the beta feature, switch to VisualEditor, or add that line to common.js to enable "experimental" suggestions: User:DVRTed/sandbox/edit-suggestions.js. — DVRTed (Talk) 03:06, 17 August 2026 (UTC)Reply
The spreadsheet would be helpful, if only because I don't know which of the csvs is the "real" one. In general I think this kind of thing is much easier to review in spreadsheet form than article-by-article. Gnomingstuff (talk) 04:18, 17 August 2026 (UTC)Reply
@Gnomingstuff: what you described makes total sense to me.[i][ii] I've checked in with engineering and it turns out that compiling this list will take a bit of time. Assuming nothing unexpected turns up, you can expect me to return here with a link to a CSV you (and everyone else here!) can review before this week is over.
---
i. Knowing, definitively, what suggestions we need y'alls expertise in reviewing and being able to differentiate those from earlier iterations that we've since discarded.
ii. Seeing all of the suggestions in one place so that you don't have to hunt them down yourself. PPelberg (WMF) (talk) 22:49, 17 August 2026 (UTC)Reply
Is the intent that the U/I would just present these suggestions to the user and let them edit the article themselves if they opt to accept the suggestion? Or is the idea to have a "Make this edit" button that the user could just click? The reason I ask is that if it's the later, it would make sense to include some machine-readable marker in the wikitext (I'm thinking an HTML comment) identifying the source of the inserted text. And/or have a log of such changes. The idea is to make it easier for somebody to audit the performance afterwards. RoySmith (talk) 23:04, 17 August 2026 (UTC)Reply
The latter would fail WP:NOLLM, so that would be a complete nonstarter. The former is not great either. ‑‑gurkubondinn 23:15, 17 August 2026 (UTC)Reply
WP:NOLLM says "Editors are permitted to use LLMs to suggest corrections to their own writing, and to incorporate them after human review. This is limited to spelling, punctuation, capitalisation, grammar, and other simple mistakes." So this would indeed be a starter in those cases. RoySmith (talk) 23:22, 17 August 2026 (UTC)Reply
The suggestions that Gnomingstuff went through go far beyond the very narrow exception in NOLLM. And this exception only applies to [an efitor's] own writing. ‑‑gurkubondinn 23:44, 17 August 2026 (UTC)Reply
The latter would be impossible in the current implementation, anyway, the LLM is not prompted to provide an actual change and the suggestions are usually just stuff like "The sentence is long, contains a grammatical error, and could be expressed more clearly." Gnomingstuff (talk) 05:09, 18 August 2026 (UTC)Reply
@RoySmith: Good question. If/when we (staff + volunteers) come to think these suggestions are reliable and useful, the intention would be to offer a suggestion that would contain the following:
1. A description of the issue and type of fix that is needed. Both of which will need to be generic enough to make sense across the contexts it might appear within while at the same time being concrete enough for the people encountering it to know how to start on the path of acting on it. So for the "Simplify language" suggestion, it might be something like, "Readers might find this text difficult to understand. Try rewriting this using shorter sentences and plain language."
2. A link to the local policy/guideline the suggestion originates from. This serves both as an opportunity for people encountering a suggestion to learn more and also a way to ensure that suggestions are grounded in project consensuses and conventions.
3. Two actions: one action to Dismiss the suggestion and along with it, a way to express why someone has elected that choice. And a second action – maybe we'd label it Rewerite? – that when tapped would A) focus someone's cursor into the span of text the suggestion thinks there is an issue within and B) cause the article text in question to enter a "revising state" so people are clear about where exactly their focus is needed (see screenshot below). From there, the responsibility would be on the person acting on the suggestion to decide what to fix and how to fix it. Said another way: there are NO plans for these suggestions to make fixes with the click of a button let alone to describe specific solutions.
Screenshot showing the revising text state within Suggestion Mode
If you'd appreciate a more concise answer, what @Gnomingstuff described here is spot-on. PPelberg (WMF) (talk) 05:32, 18 August 2026 (UTC)Reply
@PPelberg (WMF): So how would you cram enough information in such a tiny area to give the person all the information they need? Are you aware that many guidelines are like 7k words? If you give a newcomer a link to WP:NPOV, that is obviously not enough to have them actually be able to judge the neutrality of an article if they do not have relevant experience and knowledge of the field. And what will happen when inevitably people complain that edits are not improvements? Will this just WP:BITE newcomers even more? Polygnotus (talk) 05:40, 18 August 2026 (UTC)Reply
@Polygnotus: great spot. The questions you're asking sit at the very core of this project.[i]
We seem to be aligned in thinking[ii] that some suggestions may be more harmful than helpful to show to newcomers. They're likely, as you alluded to, too complex and experience-dependent to distill down into a relatively small piece of guidance. If we don't account for this, we could lead newer folks into publishing edits that experienced volunteers revert or respond to with hostility. This could in turn drive these potential contributors away.
And while I don't think we can know for certain which suggestions those are in advance, I think we (staff and volunteers) are develpoiong a pretty good sense that conversations like this one are helping us refine.
I also think we'll learn a lot over time by trying things out, and there are a few features we've put in place to help:
Experimental suggestions: suggestions can be enabled as either default-on or experimental. The latter means that people who have explicitly enabled the soon-to-be-available setting have the space to safely experiment with suggestions. Through that testing, they can decide who—if anyone—an experimental suggestion should be shown to by default.
Tags: all edits in which someone sees and/or acts on a suggestion are tagged. You can see these in action by filtering Special:RecentChanges for Edit Suggestion seen or Edit Suggestion used.
On-wiki configuration: volunteers can independently decide the minimum number of edits someone must have published in order to see a suggestion. If you visit Special:EditChecks and look at the link suggestion, you'll see that volunteers have set the minimumEditCount value to 1000. This means the suggestion will only be shown to editors who have made ≥1,000 edits.
The idea is that, together, the above will let us see the kinds of edits these suggestions lead folks to make in practice and, with that, decide whether, how, where, and to whom they're shown.
How does this sound to you? What, if anything, do you think we might be missing or misunderstanding?
---
i. I hear the questions you're asking as something like: "How might we translate a great deal of nuance and complexity into a format that is simultaneously 1) succinct and simple enough that newcomers will engage with it and 2) explanatory enough that in doing so newcomers will be equipped with the information and know-how they need to act on them in ways experienced volunteers see as constructive?"
ii. Please correct me if I've misinterpreted what you've said PPelberg (WMF) (talk) 05:32, 19 August 2026 (UTC)Reply
I would urge you to please stop developing this feature. It seems like any version of this feature will be detrimental to wikipedia.
A suggestion from an LLM, a suggestion from an official wikipedia tool, will easily be read by an editor as trustworthy, when we know it is not. LLM usage offloads critical thinking to software, meaning the editor is doing less of that themselves. This alone will have a negative impact on quality. And even grammar edits can change the meaning of a sentence, so this is just a non-starter in general.
Please do not pursue this track. TietoTeekkari (talk) 20:00, 20 September 2026 (UTC)Reply

I'm pessimistic about this, but I have a very high bar for pessimism/hopelessness to stop me from supporting a low-stakes experiment. Giving this tool to experienced users to try out seems like one of those low-stakes experiments worth trying, in case there's a way to make it work (again, I'm pessimistic, but possibly with certain constraints regarding task and topic?).
But also, just to put a finer point on something, because I think it does good to repeat it now and then: there's the worry about the quality of these suggestions, but there's also the worry about, for lack of a better word, branding. The branding that the WMF has begun using, "knowledge is human", is an echo of a popular sentiment here. At a time when absolutely every company and every project is cramming in as much AI as possible, Wikipedia is mostly headed in the other direction. That's a good thing IMO, but as a result, any pitch someone has for an LLM-based tool is evaluated with a handicap applied. If you'd otherwise be graded on a scale of 1-10 where 1 is harmful and 10 is helpful, start by subtracting 2 or 3 for "we don't want to be associated with that" (or, alternatively, "we see these as detrimental by default"), and it needs to be really useful to get a critical mass of people behind it. But no complex tool starts its life as an 8+ on that scale, I don't think, so super-low-stakes tests are IMO a good way to start working towards it. FWIW. — Rhododendrites talk \\ 22:06, 16 August 2026 (UTC)Reply

I really appreciate and agree with the branding note Rhododendrite gives here. Best, Barkeep49 (talk) 22:48, 16 August 2026 (UTC)Reply

A humorous interlude

[edit]

I know people love to make fun of AI hallucinations, so I couldn't resist posting this fun map (4:38 in the video). Strange spellings aside, Hoboken, NJ has gotten transported to the midwest, Chicago is on the Pacific coast, and Boston has been relocated to Colorado. I want some of whatever it's smoking. RoySmith (talk) 02:12, 17 August 2026 (UTC)Reply

Also available in Africa. I think I'll stick to the Commons maps. Certes (talk) 14:30, 17 August 2026 (UTC)Reply
These LLM map-fails are legion. My take away from these examples is that they are not only vivid examples of LLM hallucination, and thus LLM unreliability, but they are also vivid examples of the unreliability of humans, because each one of these published hallucinations is only possible because humans obviously failed to check the work--to even look at the maps--before publishing. Humans are unreliable: an important lesson for any crowdsourced project. Levivich (talk) 14:57, 17 August 2026 (UTC)Reply
I don't know about this particular channel but loads of them are (almost) fully automated with no human in the loop. Polygnotus (talk) 15:39, 17 August 2026 (UTC)Reply
Then again, aside from quite a few mistakes, have you all seen the "Chloe vs History" channel on youtube? Quite the advancements in AI lately. Some of the advancements have been made on this channel, which should probably have a Wikipedia page. Randy Kryn (talk) 15:42, 17 August 2026 (UTC)Reply
That Africa one is from a US State Dept presentation. That ain't no YouTube clickbait, that's supposedly professional humans at work. You'd think they'd have looked at the slides before making their presentation. And, I'm speculating here, but I bet multiple humans were involved in that, because when the US State Dept gives a presentation at an int'l conference, I don't think it's just one person creating and making the presentation all by themselves. Maybe. But either way: evidence that at least sometimes, even professional humans presenting at an int'l conference obviously don't bother to check their work. I guess my point is: what makes LLMs unreliable isn't just that they hallucinate, it's that humans won't catch it because sometimes they don't even bother to look. Another well-known example is lawyers submitting briefs to courts with hallucinated citations -- literally licensed professionals in the performance of their profession. So this isn't a problem limited to youtube clickbaiters or "kids on the internet," even trained and licensed professionals, even on a global stage, succumb to laziness. At least the lawyers get fined for it -- now there's a new fundraising channel for the WMF: fine editors for publishing hallucinations on-wiki! Levivich (talk) 16:15, 17 August 2026 (UTC)Reply
At least with hallucinations, they're usually immediately obvious to anybody who bothers to look. This particular example is clearly YouTube clickbait. The bottom-feeders who produce these things don't give a whit about accuracy, just that they can churn out some mildly entertaining garbage videos that collect likes and shares and other revenue-producing metrics. But the fact that people are willing to abuse a tool for their own commercial benefit doesn't mean that the tool is inherently worthless. People abuse wikipedia in all sorts of ways (spam, SEO, reputation management, advertising, etc). Does that mean we should shut Wikipedia down? RoySmith (talk) 15:49, 17 August 2026 (UTC)Reply
PS, it doesn't take AI to generate garbage maps. Some people are able to do it all by themselves with nothing more technologically advanced than a sharpie. RoySmith (talk) 16:22, 17 August 2026 (UTC)Reply
Sure, and wikipedia had hoaxes and plenty of innocent misinformation long before LLMs became popular, but the problem, of course, is one of scale: LLMs make this sort of thing like 100x more common than before. It used to take hours to write a good hoax on wikipedia, now it takes minutes. Levivich (talk) 16:30, 17 August 2026 (UTC)Reply
"Chiiicago". AAND it's in multiple places at once. Ladies,Gentlemen and enbies, The Chicago hivemind. Starlet 01:43, 25 August 2026 (UTC)Reply

Back to business after the humorous interlude, a full ban on AI?

[edit]

  • This would be an overreaction. User:Polygnotus and many others have been building various LLM-powered tools, including ones that are used to detect LLM edits (User:Fermiboson/AIlog) and to fight vandalism (Wikishield - LLM is optional). There are more than one hundred editors using these tools. We do want WMF to experiment and implement the best ideas. Alaexis¿question? 07:06, 18 August 2026 (UTC)Reply
  • A full ban does not make sense. We already have a range of community tools that do cool things with Wikipedia using AI. In particular I want the best available tools for review, including those that take advantage of AI for trainable pattern matching and classification. That includes anything that helps with slop-detection, edit checks, citation checks, page linting, and vandal-fighting. This experiment falls under review... with a higher bar for quality I can see a range of suggestions being useful, particularly after a few rounds of focusing on quality. – SJ + 01:07, 19 August 2026 (UTC)Reply
  • I would support a ban on generative AI touching any Wiki content regardless of whether LLMs improve. JoelleJay (talk) 11:30, 22 August 2026 (UTC)Reply

Continued discussion

[edit]

Thank you all for keeping the conversation going, and thank you to Gnomingstuff and others for taking the time to look through outputs. @PPelberg (WMF), @SSalgaonkar-WMF, and I have read the whole thread and tried to pull out some of the most important things we heard and questions being asked. Peter and Sucheta, please add if I missed anything (well, for that matter -- anyone please tell me if I missed anything). I'm just going to list these questions/topics for now, and we'll keep working through them and hopefully you'll keep discussing with us (there are many of you and not as many of us!) This is to build on what I posted above and what Peter posted above.

Firstly, I think we're on the same page about something especially important: we're not going to deploy LLM-backed suggestions to anyone unless this community is supportive of it. And if they do become deployed, they'll be configurable via Special:EditChecks the way existing checks and suggestions are (i.e. community could decide who gets to see them, change what they say, what articles they show up on, or turn them off). I hope that in projects like these, we at WMF are bringing capabilities to volunteers so that they can produce the kinds of checks and suggestions that help both newcomers and experienced editors get the most good wiki work done with the least amount of drudgery. There is also a question of whether these suggestions would be a vehicle for content suggestions -- the answer there is no, we are not designing these to propose text to the user; rather they point out places the user should look and what they should look for. The LLM explanations that Gnomingstuff pointed out initially are in those files as a way for us to understand internally why the LLM is making the suggestion; they would not be shown to editors. (But I know there is a more subtle question here -- when just pointing out a sentence that needs attention in some way exerts that sort of influence).

In that vein, I wanted to say that this current project is more about figuring out what it might be like for checks/suggestions to be generated via LLMs, than about what those exact types of suggestions are. If NPOV is not a good one to pursue (volunteers here have given many good reasons why NPOV is particularly tricky), then we should put that one down in favor of simpler ones to try out. (Worth noting that checks and suggestions are being generated via a bunch of other ways, too, like simple rules, text match, simple models, etc.)

Also just a quick nomenclature thing:

  • Checks: this refers to edit checks that react to what the editor is doing at that moment in the editor. e.g. they paste in a blob of text from ChatGPT, and the check pops up right then and says, "Please avoid copying text from other sources".
  • Suggestions: this refers to suggestions for improvement that have been pre-calculated and are waiting to show up when someone clicks Edit, e.g. they open up the editor, and there is a box that says, "This link appears more than once in this section." Right now, Suggestion Mode is available on English Wikipedia as a beta feature, and suggestions are available in the feed on Special:Homepage.

You can see all the checks and suggestions and their statuses at Special:EditChecks (note that "experimental" checks are only available to people who have a specific user script installed).

Sorry -- this post is getting long. But on to the list of questions we want to be able to talk about here, raised in the thread above. This list of questions is not something that we just want to provide answers to and then expect that everyone will agree with us. This is a complex area, and we don't have all the answers, but we're trying to figure them out with you.

  1. What kind of communication with communities has there been on this project so far?
  2. Why are we working on projects like this instead of more work on the backlog of bugs and small improvements for editing?
  3. How did we / do we QA lists of suggestions like these?
  4. What if communities don't want certain suggestions on their wiki?
  5. What determines whether suggestions are good enough to go beyond testing?
  6. Is NPOV appropriate to point an LLM at, given its nuance and complexity?
  7. If an LLM provides suggestions, how do we make sure that the editor isn't swayed/biased just by virtue of it coming from Wikipedia (a trusted source)?
  8. Does it make sense for the message to the world to be "Knowledge is Human", but we are also using AI for certain things?
  9. How could the models we use be transparent, open, and have community controls/auditing?

Alright -- more to come as we get into those questions. -- MMiller (WMF) (talk) 23:53, 17 August 2026 (UTC)Reply

@MMiller (WMF): Hi Marshall!
Because of a long history the relation between the community and the WMF is, let's say, far from perfect. And arguably worse than ever.
This is obviously not your fault, but its important to set the scene.
Because Wikipedians are smart and informed they are, generally, skeptical and wary of AI.
This is the correct position in 2026, because we do not yet know the impact of AI on the environment, jobs, and the upcoming resource wars.
The community generates the value and people think they donate to support it. The WMF is terrible at doing the community actually wants and needs. MediaWiki's tech debt is a sad joke. Community members who try to explain what we need from the WMF are routinely ignored. For decades.
Instead of working on the things that are important, the WMF nerds build shiny new toys. Debugging old code sucks. Doing something fun with AI/ML is fun. Because WMF leadership is terrible it seems (from the outside) that no one is working on the important stuff while we get an endless stream of halfbaked projects that waste a lot of time and money. The WMF basically never finishes a project it starts, overcommits at an early stage, and only gets feedback when its too late.
The WMF has a tendency to drop in and reveal they spent a lot of time working on a terrible idea, without asking input from the community, and then are surprised when the community rejects it.
The WMF is terrible at communicating its wishes and goals and what its working on.
The previous WMF attempt to "do something with AI" was a terrible idea and everyone hated it. It proved yet again that the WMF does not understand the Wikipedia community, what writing an encyclopedia means, and (the limitations of) AI.
The WMF did not learn from this and is again presenting a half-baked plan they spent a lot and time and resources on.
The community desperately wants and needs the WMF to succeed and produce great software at a reasonable cost, but that has yet to happen.
Quite a few members of the community are nerds who have spent a lot of time playing around with AI so they know what its limitations are in the context of writing an encyclopedia.
As the person behind the AI Proofreader, AI Source Verifier, AI Editsummary etc I am clearly not some anti-AI luddite.
The WMF is actively making the community more and more anti-AI, and to be honest rightly so. The WMF has no respect for the hard work of countless Wikipedians.
The fact that the WMF does things like have an AI check for NPOV problems shows that the WMF does not understand AI or its limitations or how to write an encyclopedia.
The community could greatly benefit from responsible AI use in a few specific tasks, whereby the human makes the final decision and is responsible for the edit. And the WMF makes it impossible for me to communicate that to people.
At this point, we both know its too late to listen to my feedback. This terrible idea will continue no matter what. The WMF has spent millions on AI related stuff and any benefits to the community were not proportional to the amount of money spent. It doesn't matter to the WMF; they had fun and can put something cool AI-related on their CV and move on.
The WMF has a toxic positivity problem wherein honesty and negative feedback is punished and ignored and all criticism must be hidden below 7 layers of a compliment sandwich. Even people far more diplomatic than I am just can't deal with all the corpo-speak and manipulation.
In an ideal world the WMF would stop what its doing and actually listen.
LLMs present an unique set of challenges and even opportunities and the WMF is fucking it up for everyone else and does not seem to understand how much damage they are causing and can cause in the future.
Is NPOV appropriate to point an LLM at, given its nuance and complexity? No, and that is a silly question. I don't teach my fish braille and ask it to rewrite the bible. Tools are useful in some contexts and bad in others.
What if communities don't want certain suggestions on their wiki? None of the proposed suggestions in its current form would be an improvement.
Does it make sense for the message to the world to be "Knowledge is Human", but we are also using AI for certain things? No, of course not.
How could the models we use be transparent, open, and have community controls/auditing? That is impossible unless you accept the models are terrible compared to competitors. There are no ethical LLMs available. And making something like that would be unethical and require unethical actions.
Why are we working on projects like this instead of more work on the backlog of bugs and small improvements for editing? Because the important work sucks and nerds think its fun to play around with AI. Also the WMF does not understand its role; it should serve and protect the community. In any well-run company you do maybe 90% important stuff and 10% fun stuff.
In the future, the community should be involved at the earliest opportunity, when brainstorming. The WMF should learn from its mistakes. Reflect on Simple Summaries. What went wrong, why, how to avoid that in the future? Allowing random newcomers to act on typofixes proposed by an AI will just make a lot of people very angry, sets newcomers up for failure and degrades the quality of Wikipedia. I use the opposite approach; my software helps experienced Wikipedians to fix typos, and an AI helps filter out things that aren't typos, and no one objects to that.
Polygnotus (talk) 00:41, 18 August 2026 (UTC)Reply
Please assume good faith. No one is doing this to have fun or put "something cool" on their CV and move on.
Wikimedia has actually spent a terrifyingly small amount on AI infra and tooling, which is part of the problem here: when an experiment is run, fast iteration on prompts, benchmarks, evals, and models isn't second nature. There are indeed ethical language models, as others note below, and better orchestration would help choose the best for a given task. – SJ + 16:02, 21 August 2026 (UTC)Reply
Marshall (and colleagues), I will say that from reading your post earlier, I think it reflects a good attitude for approaching the issue. (I'm not going to go down the rabbit hole of the difference between words and actions and trust.) I will also note that I have no idea how much say you have in what projects you work on and how far you take them. (Although if "Senior Director of Product" isn't just a bunch of fancy words, I would guess quite a bit.)
I clicked your link to see what's currently on Special:EditChecks, and I will say that almost all of them seem like really good ideas. None of them (the ones I like, at least), require any AI beyond some if statements in a trenchcoat.
In short: I think we could completely ignore all this AI stuff and you'd have some really great and helpful projects to work on. (Others have been mentioned above.)
I know working with such a large community can be hard. I would certainly like to think that we could give you some advice on how to help us. Ask . We only bite sometimes. Thanks for taking the time to listen. LittlePuppers (talk) 01:47, 18 August 2026 (UTC)Reply
That is an excellent and fitting username. Polygnotus (talk) 01:53, 18 August 2026 (UTC)Reply
One time I said that I didn't know what "Director of Product" means (I am not a native speaker) and the ex-CEO of the WMF started attacking and accusing me because she couldn't handle mild criticism. Polygnotus (talk) 03:11, 18 August 2026 (UTC)Reply
Thanks -- just to take this opportunity to shed a little light on what I do and how we're set up:
  • We're set up as a bunch of "cross-functional" product teams. Cross-functional means each team has a few engineers, an engineer manager, a designer, a product manager, a data analyst, a movement communications specialist (give or take -- sometimes a couple teams will share people in certain roles).
  • The product manager is responsible for setting the priorities of the team, what order to work on them, deciding what is in and out of scope, deciding whether the results of an experiment show that something is worth pursuing further. They do this in collaboration with their teammates, not unilaterally.
  • As a director of product, many of these product managers report up to me. When I started at the foundation, I was the PM of the Growth team, and I've been around for eight years and am now in a management role.
  • So in my role, I look across the various teams and play a role in setting the overall priorities -- like which various potential editing projects the various teams working on contributors should prioritize, what the readers teams should prioritize. I work with director counterparts in engineering and design to do this.
MMiller (WMF) (talk) 18:28, 18 August 2026 (UTC)Reply
Re 6, and to some extent 9: Does the WMF have a suitable operational definition of NPOV to measure against? As I understand it, transformer-based models are fairly common in sentiment analysis nowadays, but I don't know if pre-existing models (if any exist) would be well adapted to an encyclopedic context, and I don't think they would be sufficiently interpretable. In any case, I'd say that it is something that is strictly more difficult than tone check, as well as likely being considered higher risk by the community. I can't be certain of course, but I would expect for the community to grant social licence to a model for NPOV specifically would, at minimum, require much better interpretability than typical for current models. Whether the team believes they can achieve that is probably something best answered by the technical staff, the community can only inform as to acceptance criteria. Alpha3031 (t • c) 05:00, 18 August 2026 (UTC)Reply
@Alpha3031 They just take off-the-shelf selfhosted open-weight models, which are way worse than Claude Fable 5 (and other mainstream commercial offerings) like gpt-oss:120b aya:35b / aya-expanse:32b, llama4:scout, qwen3:235b and qwen3:1.7b and then give it a presumably AI generated prompt containing:
 Role: You are an expert Wikipedia Copyeditor and Reviewer.
    Goal: Conduct a professional audit of an article's plaintext so it aligns with Wikipedia core content policies and Manual of Style.

    Context & Constraints:
    - Plaintext environment: non-prose elements may be stripped (infoboxes, references, templates, media, etc.).
    - Non-markup focus: do not suggest wiki-syntax, template, link-formatting, or reference-format edits.
    - Awareness of extraction gaps: if text appears truncated from extraction, do not flag it unless clearly an authorial issue.
    - Focus only on prose, terminology, neutrality, and information structure.
And
    - Neutrality: remove peacock terms, bias, and unnecessary loaded language.
You appear to be overestimating the sophistication of their approach by a rather wide margin.
Polygnotus (talk) 05:14, 18 August 2026 (UTC)Reply
I am aware of that, given that I had read (and in fact posted the link to) said prompts above. Though, I wouldn't say hosted open weights models are necessarily "way worse" than commercial offerings currently given recent and especially upcoming releases (GLM at 700B especially is likely a lot easier to host than the 2.xT models, though still harder than 120B of course). The page does indicate that the team recognises some situations may require bespoke models though. Alpha3031 (t • c) 07:05, 18 August 2026 (UTC)Reply
Open-weights models can be perfectly adequate for some use cases. For source verification the very modest gpt-oss-20b model worked as well as Sonnet 5. Detecting NPOV issues is just a much harder problem, and it needs much more context and probably better models as well. Alaexis¿question? 07:18, 18 August 2026 (UTC)Reply
I don't disagree that open-weights models (even older or smaller ones) can be adequate for many tasks. I just also wanted to point out that as of the time of the current discussion, there are several open-weights models in the 700 billion to 3 trillion parameter range that are comparable to the frontier in a much wider selection of tasks, namely Kimi, Qwen and the upcoming GLM 5.3 (which being much smaller, is probably going to be cheaper to evaluate).
Incidentally, it does appear that previous WMF research findings (m:Research:Test External AI Models for Integration into the Wikimedia Ecosystem#Evaluation Results and Findings) have pointed out the difficulty of NPOV assessment:

Some policy detection tasks are hard for humans and machines alike. Specifically for NPOV, precision across the three model families is very low. Models overall do better when detecting Peacock behavior. This is because while understanding neutrality requires in-depth reasoning, [bold mine] peacock behavior can be detected via language features.

so it's somewhat odd that they've picked it as a easy, low-risk task in this specific project.
I would say that adequate NPOV assessment is probably the hardest possible task to set as a goal, given that it requires the aforementioned tone check, a comparison with the sources, and also some sort of test to ensure the sources the other models see are an accurate reflection of the body of published RS more generally. It's definitely not suitable for a project intended to explore what can be done with relatively generic models. I think there could also be room to explore, e.g., faster community feedback cycles so that the WMF doesn't feel like it needs a whole year to develop a single task specific model. Alpha3031 (t • c) 08:56, 18 August 2026 (UTC)Reply
Agree that NPOV assessment is the hardest possible task, also for humans. More importantly, the assessment relies on consensus. There is no authority that has the truth about whether something is NPOV. The whole point is that the community reaches a consensus looking at all available reliable sources. Delegating this to an LLM (with the mark of approval of the WMF itself, thus giving it an undue resemblance of authority) defies the premise of Wikipedia, which is based on decisions taken by consensus. Aggravated by the fact that LLMs are not at all unbiased by any definition of the term (and are often just plainly wrong). Ita140188 (talk) 10:27, 18 August 2026 (UTC)Reply
Agreed, NPOV is hard! Alaexis¿question? 12:52, 18 August 2026 (UTC)Reply
Thanks for the response, this is already much more transparency than the last time around and I appreciate it.
I am a bit confused as to whether these questions are for us and for you. Gnomingstuff (talk) 17:35, 18 August 2026 (UTC)Reply
@Gnomingstuff: oh, good clarification. The questions Marshall posted are a first pass at translating what we're hearing from y'all into a set of questions that we (staff) can then respond to one-by-one. Of course, if you think there are questions we've missed and/or questions that we've misinterpreted, please comment as much. PPelberg (WMF) (talk) 20:05, 18 August 2026 (UTC)Reply
I thought I would ring in to point to another AI-backed project WMF is doing and share our experience working with volunteers on it, given the interest here in this sort of thing. For context, I lead the Product Safety and Integrity (PSI) team here at WMF, we build security and safety features.
One thing we're working on is detecting abusive content using LLM-based models. Specifically to enwiki, last week we deployed a new feature that is only visible to oversighters, at Special:AbuseReview, which uses an open-weight model called CoPE to flag edits that probably need to be suppressed due to containing personal information.
This on-wiki feature is now being used enwiki oversighters to review and take action on what it raises. They tell us the practical accuracy rate they experience in reviewing the output is better than what they see in the reports they get from human users. And, when we were testing the model output in June and July, in batches through an off-wiki process, most of the true positives it was catching were not being caught organically on-wiki.
One thing I want to mention is the iteration and volunteer collaboration that it took for us to make something deployable. There was some initial volunteer skepticism, and we needed to demonstrate not only that it was valuable, but that this was going to be accurate enough to not waste precious volunteer time. That took some internal experimentation on our part to get something that showed enough promise on both those fronts that we could ask for support doing a round of manual labeling -- which is what really pushed it over the edge of accuracy to be a deployable feature.
This collaboration has helped us do something meaningful about doxxing on English Wikipedia over the last few months, that feels good to both WMF and volunteers. That is possible in large part because volunteers were willing to give us space to operate, and to keep an open mind that could be convinced by data. We also needed to be flexible in our own thinking and incorporate volunteer ideas, something PSI has gotten used to doing in our work. EMill-WMF (talk) 21:17, 19 August 2026 (UTC)Reply
Endorsing Eric's statement, as one of the oversighters involved in the testing/iteration he describes. This is catching OS-level material we were not catching otherwise, and is a clear benefit to the encyclopedia. In solidarity, asilvering (talk) 21:54, 19 August 2026 (UTC)Reply
I think there is a role for AI to play on Wikipedia and it's going to be in something kind of like this. Something Eric alluded to but doesn't say directly which I think is important: the first draft the WMF showed us was nowhere close to ready. From that the learning was that we needed to go even farther in how confident . The second draft - which is when we started doing manual labeling - was still not anything which would have been appropriate for use. It did lead to a learning that one element - personal information - was more reliable than the other OS criteria which got us to the third draft and the efforts to then backfill June & July. That got us a lot of incredibly sensitive information that wasn't getting reported and which is now getting appropriately oversighted. I was the one who ran some data after we completed June to compare the true positive rate against email reports and found it to be in the range of 5% higher. I plan to revisit these stats soon and my hope would be that we'd have a higher difference between email reports and these flags for two reasons: 1) the classifier has been improved since I ran those numbers originally 2) some of the "obviously needs OS" tickets we'd have gotten in the past, we won't be getting because OS will be oversighting edits before some other qualified person finds them and reports them. I am really glad we're getting these signals now to stop some very sensitive personal information from being exposed against policy, we wouldn't have been able to do it without the AI classifer model, and it did take an iterative process to get there. Best, Barkeep49 (talk) 23:05, 19 August 2026 (UTC)Reply

Downloaded the most recent csv, picked an NPOV one at random:

  • "1358549724,The_Dying_Rooms,10791564,NPOV,Remove biased and profane language describing the Chinese government and replace it with a neutral summary of the government's response.,"In the film, Blewett and others travel to mainland China to visit orphanages housing children abandoned due to the ""one-child policy"". The filmmakers stated that unwanted female and disabled children were left to die of neglect, allowing parents to have another child. Showing that China government is a piece of sh!t for letting this happened and being a coward by lying to our faces that it didn't happened and said that all the footage is fabricated to destroy the reputation of the china government.>",enwiki,df59b690-6c8e-4827-b809-76e90d409f9f,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

The statement the LLM objects to was in the article for less than a minute and long reverted by the time the suggestion was created. So one can add to all the above problems and errors that it also wastes times and resources by not checking the current version but some snapshot.

Profanity checks in general are a bad idea.

  • "1358746317,2024_G20_Rio_de_Janeiro_summit,72163669,NPOV,Rephrase the incident with the first lady using neutral language and omit the explicit profanity.,"During a speech about fake news, Rosângela Lula da Silva, first lady of Brazil, swore at Elon Musk, saying: ""I'm not afraid of you. Fuck you, Elon Musk"" (Eu não tenho medo de você. Inclusive, fuck you, Elon Musk, in Portuguese).",enwiki,fba2904a-9276-47e3-8f5c-27fe5aff71f6,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

No, we are not going to "omit the explicit profanity" from a quote, and suggesting things like this is a very, very bad idea.

  • "1357774252,God_Emperor_Trump,74631506,NPOV,Remove profanity from the description of the phrase on the sword to maintain a neutral and encyclopedic tone.,"According to Fabrizio, the phrase could mean 'here's your fucking tariffs'.",enwiki,0f886b1f-2a4c-4303-b06e-63da6700b71e,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

Again, it's a quote, from the creator of the sculpture. The AI should not make statements like "Editors often revise this kind of wording, saying the tone is unbalanced. ", which only works to influence newbies by making a false claim to authority.

And then there the internally contradictory advices, indicating the inherent stupidity of LLMs.

  • "1356142445,Allan_Segura_(model),80130058,simplify_language,Combine the two sentences about sexual orientation and activism into one concise sentence.,Allan Segura is openly gay. He is a transgender rights activist.,enwiki,d28d67bc-2b32-4cbb-9135-ed391c91c13b,Readers might find this text difficult to understand. Try rewriting this using shorter sentences and plain language. [Learn more](https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style#Vocabulary).,Simplify language"

So do we need to combine these two (very short) sentences into one sentence, or do we need to rewrite this using shorter sentences? Or, just perhaps, our readers are perfectly capable of understanding these two sentences and won't "find this text difficult to understand". What a joke. Fram (talk) 09:01, 18 August 2026 (UTC)Reply

@Fram Also, wasn't the fact that the WMF didn't do content the thing protecting them in lawsuits?
Let's say an article contains negative information about a rich person. If the WMF starts to mess with content, do they not open themselves up to be forced to make changes? I am not a lawyer. Polygnotus (talk) 12:23, 18 August 2026 (UTC)Reply
All good reasons not to use AI for this purpose and probably not for any purpose. Once again, the WMF is wasting what remains of its valuable technical staff (and its ample funds) on trying to drag Wikipedia in completely the wrong direction.
If the bot is criticising vandalism which was only live for a minute, I suspect that it may be reading page history (why?) rather than taking a snapshot. Certes (talk) 12:32, 18 August 2026 (UTC)Reply
A snapshot sample of a large number of articles will also catch revisions that lasted only for a minute. There are lots such edits that are reverted quickly.
Assuming that edit suggestions are not displayed when the content has changed, this particular suggestion would never have been shown - no harm down do anyone. Generating checks on a snapshot rather than live is an architectural decision driven by performance, complexity and other considerations. Alaexis¿question? 12:51, 18 August 2026 (UTC)Reply
Presumably that means any such system will need to be rerun on each relevant page each time they are edited, in case the edit touched the prompted part? Not sure how the current tasks handle this actually. CMD (talk) 12:59, 18 August 2026 (UTC)Reply
Not necessarily, it can be run once a month, with some kind of caching enabled not to re-check the vast majority of content that stays the same. Then the complex and time-consuming part (LLM calls) is done asynchronously, and the easy part (deterministically checking that the content stayed the same) is done when the user opens visual editor. Alaexis¿question? 14:34, 18 August 2026 (UTC)Reply
That is indeed how it's currently working. A large batch of the suggestions were pre-generated, and was poured into a fairly simple API that VisualEditor's suggestion mode knows how to query and match up to a document.
Presumably if this was a successful experiment that we decide together to scale up, it'd turn into a more complicated system where we do something like fire off a job that precomputes the suggestion for each new revision of a page. Still avoiding needing to do the expensive generation every time someone opens the editor, but keeping them more up-to-date. DLynch (WMF) (talk) 23:03, 19 August 2026 (UTC)Reply
Again, it's a quote, from the creator of the sculpture.
That's another LLM tic -- at least when it comes to Wikipedia edits/suggestions, they really hate direct quotes and will usually tell you to paraphrase them and/or do so themselves. (Example from a 2026 edit summary: "Paraphrased a lengthy direct quote regarding the skater's performance at the 2025 World Team Trophy into a concise summary. This improves readability and helps maintain a neutral, encyclopedic tone by removing overly detailed personal reflections.") Gnomingstuff (talk) 17:41, 18 August 2026 (UTC)Reply
The examples Fram shows above makes me more resolute in opposing this on principle. The WMF is trying to exert editorial control and hiding it by labelling them as "suggestions". ♠JCW555 (talk)♠ 16:48, 18 August 2026 (UTC)Reply
@JCW555, I guarantee you they are not trying to exert editorial control with this feature. They're trying it because they think it will be helpful for beginner editors. We can tell them they're wrong and we don't like the feature without impugning their motives. In solidarity, asilvering (talk) 21:21, 19 August 2026 (UTC)Reply
But suggesting sentences should be reworded is the WMF trying to exert editorial control. To quote MMiller above "It would just point out the spot in the article that needs attention, e.g. "Does this sentence need to be rewritten to be easier to read?". Whether a sentence/passage needs to be reworded is a matter for the talk page amongst editors, not through the WMF via their AI. No reply from the WMF in this section has done anything to assuage that concern for me. If the WMF comes out and says that these "suggestions" wouldn't touch content at all, no matter how small or big, then I'd be a tad less aggressive in my opposition. ♠JCW555 (talk)♠ 21:44, 19 August 2026 (UTC)Reply

Is NPOV appropriate to point an LLM at, given its nuance and complexity?

[edit]

Hi all, I'm Sucheta, product manager on the Machine Learning side of this work.

Reading everything you all have said, it’s clear we shouldn’t move any further with the LLM-generated NPOV suggestions. We're dropping that type rather than trying to iterate on it.

@Ita140188 said this really well: There is no authority that has the truth about whether something is NPOV. Makes sense to me – NPOV isn’t just about word choice; it's about the representation of a topic that editors decide on by weighing the available reliable sources against each other and reaching consensus. We had wanted to take a crack at it to see what the LLM would produce, so thank you for looking at these and thinking about them.

I do want to make the distinction between these LLM-backed NPOV suggestions and the Revise Tone suggestions that are in production now on this wiki. Those Revise Tone suggestions come from a model called BERT that we’ve fine-tuned for a narrower scope. The model is trained to notice peacock language, based on 20,000 examples of revisions where the "peacock" template was added or removed. We think these have worked out well, and thousands of them have been actioned by both new and experienced editors on this wiki. Having them available also made it more likely that a newcomer would make a constructive edit, than when they open the editor without some suggestion inside. You can see them on Special:Homepage in the suggested edits feed (if you check these out and have thoughts, please let us know).

What about the other types of LLM suggestions besides NPOV? Well, let’s keep talking here about whether they have potential. Note: we're working on making a spreadsheet available to you all so that you can see all of the suggestions in one place. -- SSalgaonkar-WMF (talk) 17:13, 18 August 2026 (UTC)Reply

Thanks for the response. Good to hear about the NPOV feature.
The main issue is mostly the same across the board: the actual LLM suggestions have issues as above, but including broad categories without the actual suggestions is just confusing and provides almost no context. This is going to affect any possible category, it's just structurally inherent to the task as I understand it. Other than that:
MOS:GEO: Two issues I can think of:
  • Seems near-certain to inadvertently wade into a geopolitical quagmire of some sort.
  • Less dramatically, a large proportion of these suggestions are really just suggestions to treat everything as first reference (the "Atlantic Ocean"/"Atlantic" thing mentioned above). I don't think there's any way to get around this while working at the single-sentence level.
Simplify language: Two issues again --
  • The text parsing needs to be fixed before anything is done with this since otherwise the suggestions won't make any sense, especially the issue of parsing multiple sentences as one (e.g. King's fourth novel, Euphoria (2014), was inspired by events in the life of anthropologist Margaret Mead. It won the inaugural Kirkus Prize for Fiction and the 2014 New England Book Award for Fiction, and was a finalist for the 2014 National Book Critics Circle Award. Euphoria was listed among The New York Times Book Review's 10 Best Books of 2014, TIME's Top 10 Fiction Books of 2014, and the Amazon Best Books of 2014. -- they aren't visible on-page but there are unicode separators between many of the clauses, maybe that's related).
  • The category scope seems to be off. The vast majority of suggestions are to break up sentences, comparably little about actually simplifying language -- if anything, the language suggestions I've found seem to be suggesting the opposite, to make language more complex. But suggestions for typo fixes etc. also show up here.
Gnomingstuff (talk) 17:58, 18 August 2026 (UTC)Reply
That unicode-characters thing is actually deliberate. The data comes pre-massaged into the exact form that works in VisualEditor's search (non-text gets replaced with a opening and closing internal model tag, so that  is actually <ref></ref>). E.g. If you go to the Lily_King page that quote's from and paste it into the VE search box, it should highlight that entire paragraph, covering the citations. As far as I know, that replacement got done as a post-processing phase after the initial suggestion-generation, though I wasn't directly involved so there's a chance I'm wrong.
...that said, you did make me realize that we're accidentally not comparing them in that form in our final "has the user already changed this bit of text since they started editing" check, so gerrit:1326927 will make all these suggestions that cover citations / templates actually visible for review. DLynch (WMF) (talk) 20:54, 18 August 2026 (UTC)Reply
  • Questions about peacock language; Does the function ignore direct quotes? Any suggestion to change the wording of a direct quote shouldn't happen. Also, can we can we get it to come down hard on subjective words like best while being less aggressive on words that might be verifiable facts like largest? And maybe even less aggressive for largest known and largest on record? --Guy Macon (talk) 18:39, 18 August 2026 (UTC)Reply
    @Guy Macon: Great question. This suggestion can be configured to ignore quoted content which, at present, is exactly what en.wiki has done.
    You'll notice that ignoreQuotedContent within Special:EditChecks#tone is set to true. If you'd like to learn more about exactly about how quoted content is detected, T414715 contains more details. Could you please let me know if anything you see (or don't see) brings other questions to mind?
    Now, to the second question you're asking...
    Assuming it's accurate for me to understand it as something like "How might we specify the suggestion based on specific words/phrases?" I wonder if you think TextMatch could be helpful here. In essence, it enables volunteers to write custom suggestions to appear when predefined words/phrases are detected with an article. PPelberg (WMF) (talk) 19:45, 18 August 2026 (UTC)Reply
An update on the NPOV suggestions. As of ~30 minutes ago, we've removed all NPOV suggestions from the experimental batch of model-generated suggestions. Thank you all for the quick feedback here and for trusting us to hear you.
Note: if, by chance, you happen to still encounter one, can you please let us know? For now, we implemented the above in a bit of a fragile way so we can get something out quickly and will come back to make this more robust in the coming days. PPelberg (WMF) (talk) 22:11, 18 August 2026 (UTC)Reply
Thanks, love this fast and encouraging response. If peacock language detection is working well with a training set of 20,000 examples, is compiling training data a helpful step for catching other narrow style issues? – SJ + 01:20, 19 August 2026 (UTC)Reply
Great question! I think so - though this project is meant to test how viable it is to generate suggestions without manually compiling the training data you described.
If we were to place the LLM-generated suggestions and Tone Check approaches on a spectrum, Tone Check would sit at one end: it works well for identifying a well-defined issue, but it took us over a year to build and deliver, with training and evaluation being two of the most time-consuming pieces. We're now considering three ways to generate suggestions using models, among the many other approaches we're exploring:
  • generic models with task-specific prompts (what we're testing here)
  • generic models plus fine-tuning
  • specialized, bespoke models
So the question we're really asking is what amount of rigor is required to produce suggestions that are sufficiently reliable and useful?
We're asking a similar question about evaluation: whether faster methods like LLM-as-a-judge can tell us if suggestions are reliable without hand-labeling a large test set. Even here we review a number of samples manually to make sure the judge is scoring appropriately; that sample is just much smaller. We also created an eval dataset of past user edits per suggestion type so we could compare the model-generated suggestions to actual edits.
What do you think about this lens? Do you think there are some suggestion types that could be supported by this approach of using generic models without fine-tuning? SSalgaonkar-WMF (talk) 15:42, 20 August 2026 (UTC)Reply
How much effort is something like this? ScottishFinnishRadish (talk) 16:01, 20 August 2026 (UTC)Reply
Thank you for sending this! I really like this idea of experimenting with LLMs to generate basic copyediting suggestions. Reiterating what you said: they seem to be relatively low-risk, simple, and as a result, something newcomers could handle. In terms of effort, I don't think it would require an impractical amount to build a dataset like this (as your demo even shows).
We started to explore copyediting suggestions using MoS guides around capitalization and grammar, but we didn’t feel confident enough about their quality and utility to include them in this initial dataset. The capitalization suggestions scored lower in our LLM-as-a-judge evaluation than other suggestion types, and our manual review showed that “grammar” was too broad a category; we couldn’t come up with one label or description to describe the range of issues surfaced by grammar suggestions.
These both feel like solvable problems, and I’d really love for us to take another look. Would you be willing to share the prompt you sent to ChatGPT? SSalgaonkar-WMF (talk) 13:54, 21 August 2026 (UTC)Reply
Find some typos that can be fixed or copy that can be edited at https://en.wikipedia.org/wiki/Mircea_II_of_Wallachia. If I were going to refine it I would probably break results down to things that need review from someone with varying levels of experience so you can filter the output to users based on experience. For instance, checking if the source said first born or firstborn is a great task for someone trying to step up from beginner editing. ScottishFinnishRadish (talk) 15:10, 21 August 2026 (UTC)Reply
I don't know if the technology can support this, but it would be nice if we could have multiple queues of suggestions in different areas. A queue for spelling and grammar fixes and a queue for source-to-text validation in STEM articles might appeal to different people looking for work. RoySmith (talk) 16:02, 21 August 2026 (UTC)Reply
When I was looking through the csv before I didn't find errors when the llm was identifying the wrong units being used or similar. In my personal testing I've found it good at picking up tense mismatches (occasional errors sure but overall picks things up I've missed when rewriting something). This sort of pattern fixing seems much less likely to cause an issue than setting an llm to gambol over fields of longer text, as well as being a simpler spot and fix for newcomers. CMD (talk) 00:45, 19 August 2026 (UTC)Reply
My apologies if I just read past it and didn't notice, but where is this CSV file? RoySmith (talk) 00:49, 19 August 2026 (UTC)Reply
@RoySmith In , GnomingStuff dug it up. CMD (talk) 01:11, 19 August 2026 (UTC)Reply
Got it, thanks. RoySmith (talk) 01:14, 19 August 2026 (UTC)Reply
We'll be sharing a clearer spreadsheet version (with the most up-to-date entries) tomorrow, per Peter above. HTH. Quiddity (WMF) (talk) 01:17, 19 August 2026 (UTC)Reply

Why are we working on suggestions?

[edit]

The first reason we're working on this is because of the need to get more new people involved in editing. Being a newcomer has always been hard, and is especially hard now that people spend most of their online time on mobile. When newcomers (especially on mobile) open up the editor for the first time, it is overwhelming and they often just leave -- they are like "Wow, scary, nevermind." (here are some interesting survey results about this moment) But we've seen that when we point out specific bits of the article that could use improvement, the newcomers are much more likely to do something constructive and stick around. We've tested many of the checks and suggestions in Special:EditChecks and they have had these measurable positive impacts. We’ve also been inspired by the tools/scripts/gadgets that volunteers have built that do similar things (some examples here).

So this project here (the LLM suggestions) is another way of learning how we might find more kinds of suggestions in the vein of "how can we help newcomers on mobile be more and more constructive and more likely to stick around?" (in ways that align with policies, values, and existing editor workflows).  It’s also worth noting that the newcomers who start with suggestions often wander off on their own in the wiki once they get comfortable. The design helps with that, because the suggestions happen inside the Visual Editor, i.e. you have to edit in VE to get them done (as opposed to, say, a separate interface).

Secondly, we also think that this can help lower patroller burdens at a time when those burdens are increasing because of AI slop, as Gnomingstuff has pointed out. (i.e. newcomers making constructive edits in the first place lowers burden on patrollers from newcomers being confused). For example, we are developing a way to deter and label edits when people are pasting content from an LLM.

And thirdly, we think that these suggestions can be helpful for experienced editors, too.  A few of you have said in this conversation that experienced editors don’t need help finding improvements to make, but we have also heard from many who appreciate suggestions like these. In my own editing experience, I often click edit to do a specific change, but then discover a few other small things to improve via the suggestions.  So we think there is also opportunity here to help experienced editors get more wiki work done with less seeking/searching/effort.  It becomes a question of which of these signals to present to which users in which places.

How does this all sound? It would be great to hear from anyone who has seen these suggestions in action with newcomers or has been using them.

-- MMiller (WMF) (talk) 19:49, 18 August 2026 (UTC)Reply

Around here, the city is running an e-scooter pilot. Install the app, hop on a scooter, ride to where you're going. One of the interesting things they do is the app automatically imposes a lower speed limit on all new riders. Once you've ridden more than (IIRC) 10 hours, you get to go full speed. I think it also won't let newbies take out a scooter after sunset. I forget the details, but you get the idea.
I could see doing something similar here. Have some way of scoring suggestions for how risky they are. Fixing an obvious typo is pretty low risk. Rephrasing a statement that appears to be biased is higher risk. The type of article might also factor into it: an article about a WP:CTOP would probably not be the best choice for a newbie to learn on. The total neophyte would only get the safest suggestions. People who had gotten a bit more experience might be offered a wider range of suggestions. RoySmith (talk) 20:14, 18 August 2026 (UTC)Reply
MMiller (WMF), two thoughts come to mind:
  1. A good place to ask might be WT:AFC and similar (e.g. NPP)—on one hand, writing articles is kind of exactly what we don't want new editors to try, because it's really hard, but we see about every kind of possible error there: from formatting (a first heading duplicating the page title, formatting inside headings, malformed templates, all sorts of weird stuff) to tone (as established this is hard, but there are probably a few ways to detect COI) to referencing (citing Wikipedia, citing social media, just not referencing, weird formatting, duplicated refs). I see that some of it you have projects related to, but that's a handful more off the top of my head, and I'm sure people at those pages can think of more. I'd also be curious if you've investigated what the impacts of limiting this to visual editor are (i.e. how many new editors use VE).
  2. Another thing to consider is looking at what gadgets and user scripts are commonly installed. Your check about disambiguation links makes a lot of sense to me, because there's been a gadget which displays them in a different color for years. And some of those do make more sense as user scripts or being community maintained, but there are definitely some which would make more sense as part of MediaWiki or which could use some love. (Looking through my user scripts, there are a handful I don't really use, and some which are enwiki-specific, but also many which make a lot of sense to integrate or which I'm importing cross-wiki and haven't been updated for 8 years or something.)
I don't know if you're short on ideas or not, but those are what came to mind for me.
And I just reread your post and realized you already did #2. LittlePuppers (talk) 21:21, 18 August 2026 (UTC)Reply
@LittlePuppers: thank you for sharing ideas of places to look for and vet ideas for new Checks and Suggestions. I've added WT:AFC and Wikipedia:New pages patrol to to the MediaWiki page where we are bringing together this sort of information. If/when other ideas come to mind, we'd be thrilled if you'd add them directly. Of course, I'm happy to add them as well. Just give me a ping if/when something strikes you.
On the topic of gadgets and user scripts, we're with you here. In fact, doing what you described helped prompt the work we're partnering with @Alaexis on to turn a tool he wrote for identifying cases where a source might not support its associated claim into a new experimental suggestion. PPelberg (WMF) (talk) 23:13, 18 August 2026 (UTC)Reply
Re writing articles is kind of exactly what we don't want new editors to try, I'm not sure I agree with that. For some new users who don't know how to get started, having them fix typos might indeed be a good way to ease them into editing. But some people will already know what they want to write about. If you tell them, "No, that's too hard, we want you to fix typos for a while", all you're likely to have done is lost an opportunity to get a new editor hooked on the project. The very first edit I ever made was to create City Island Bridge. RoySmith (talk) 01:32, 19 August 2026 (UTC)Reply
Yeah, I was wondering if someone would push back on that. My wording there was probably too strong, main point being that it's really hard and, I suspect, very often discouraging. But then again, I do see, on occasion, someone who will read through guidelines and knows how to write and all that and write pretty decent articles very early on. To be honest, those are probably also the people who write decent articles later on.
Today, your first edit would have to go through AfC, and get declined for being unsourced... ah, those were simpler times. There was probably also more low-hanging fruit then. I have no idea where I'm going with this comment. I'm all for developing features to help out those who start in all sorts of ways. LittlePuppers (talk) 01:59, 19 August 2026 (UTC)Reply
The difficulty is that so very many new editors know what they want to write about, but what they want to write about isn't going to make it to article form at the present time. They want to write about themselves, their companies, their favourite YouTuber, the assignment they've been given, the really fun thing they made up...
Yes, we would lose them if we told them to fix typos. But we also lose them if we don't let them create that article they want, and we can't let them create that article they want. Many of them are also happily LLM-ing it up in their draft. I almost wonder whether there's a call for a service (human, AI/LLM, both?) that we try to push would-be article creators towards that assesses or helps them assess whether they actually have sources - and tells them to stop if they don't, or at least to try another subject instead. Perhaps an AI/LLM project could be given some basic guidelines about common problems (interviews are likely to be inappropriate; this source doesn't seem to be about the subject) and at least head off the ones that don't have a chance. A lot of people seem remarkably inclined to listen to what a machine says, possibly because it's seen as authoritative and knowledgeable on basically any subject. Meadowlark (talk) 06:27, 19 August 2026 (UTC)Reply
Hi @Meadowlark, I'm Rita Ho, director of design at WMF working as the design counterpart to Marshall with product teams working on these new editing tools. Your comment about helping editors to create new articles by providing more guidelines (like including sources) is related to another feature in development, called Article Guidance! This feature is aimed at helping newer editors who want to create articles to succeed. It's kind of like if Article Wizard could be tailored and offer specific support depending on the type of article being created (animal/building/person/etc). It includes initial source validation, notability risk assessment, and initial minimal content structure or "outline" for someone to get started, and is community configurable.
Sharing in case you and others are interested to learn about this other project and participating in testing and giving feedback. RHo (WMF) (talk) 22:01, 19 August 2026 (UTC)Reply
I do like having community-create outlines. That seems (at least in theory) to catch a lot of the types of mistakes I see with new editors (sources!). LittlePuppers (talk) 22:11, 19 August 2026 (UTC)Reply
It becomes a question of which of these signals to present to which users in which places.
Building on what Marshall shared above, I think it might be useful to consider that we're trying out a range of ways for generating the signals to power new edit checks and suggestions...
Some suggestions, like adding references and detecting when someone has pasted content from an LLM, are generated using deterministic rules.[i] Some look for the presence of specific templates, like citation needed. Others use small language models specially trained on Wikipedia edits to identify a specific type of issue, like finding issues with tone or suggesting images to add to articles. There are also suggestions written by volunteers using the TextMatch feature (inspired by AbuseFilter). We're also experimenting with a way for volunteers to create suggestions based on the presence of maintenance templates or missing template parameters.
I share all of this in an effort to communicate that we're eager and open to experiment with a signals from a variety of sources. What's most important to us is identifying ones that we collectively see as reliable and useful.
---
i. E.g. contents of clipboard metadata and the absence of a reference within a defined amount of new text PPelberg (WMF) (talk) 23:03, 18 August 2026 (UTC)Reply
Thanks, I've edited the mediawiki page to add two ideas, one to identify/highlight problematic text that a maintenance tag is referring to, one to point people to en:Help:Find sources if they try to use an unreliable source like a blog or social media. I think it's probably best if we focus on problems that would result in a revert for the sake of retention of newcomers (rather than minor stuff like MOS:GEO, where people can learn from someone editing their work). I'd also suggest working more closely with WP:AIT if you aren't already (if they have the time, people such as @Polygnotus, Alaexis (as I see you're doing), @Dreamyshade) as they'll likely have a lot of ideas and be able to discern some issues which may not be apparent. Otherwise I'd just encourage people to boldly edit the mediawiki page and add any ideas or whatever. Maybe it could have a second column for concerns about an idea? Kowal2701 (talk, contribs) 08:41, 19 August 2026 (UTC)Reply
Approaching this from the narrow goal of improving Wikipedia, I'm concerned to see newcomers and suggestions again mentioned in the same breath. Automatically produced suggestions need to be assessed individually by someone familiar with writing for Wikipedia, and that's not a newcomer. If the goal is not to improve Wikipedia but to increase the active editor count by making newcomers feel useful then this may be a viable scheme, in the same way that we reluctantly allow education projects to introduce so many errors to our articles. However, if that is the case then (yet again) editors and the WMF are pulling in different directions and we have a clear conflict of interests to resolve. Certes (talk) 09:18, 19 August 2026 (UTC)Reply
@Certes -- yeah, I understand. We've been talking about this since the early days of the Growth team in 2018: were suggested edits more about improving Wikipedia, or about retaining newcomers so that they could grow into editors who improve Wikipedia later? We generally erred on the side of retaining newcomers -- for instance, the first suggested edit was "add a link". Do the wikis really need lots more blue links between articles? Some wikis do, some not really. But it was a good task for newcomers to get their feet wet, have a succesful first experience, and want to come back again. We saw some newcomers "go on a run" where they did hundreds of those tasks over the course of several days. And we (happily) also saw many do a few of them and then go do some other, higher value edits on their own.
But the ideal is that we can do both at the same time: design tasks that are both healthy for newcomers and constructive for the wikis. I would say that "revise tone" is an example of this. I would be curious what you think of the diffs in this Recent Changes filter, which shows both links and tone diffs, highlighted by how experienced the person is.
And more generally, where you would come down on that trade-off between "invest in newcomers learning" versus "constructive edits now" (hoping, of course, that we could have our cake and eat it too with the right designs). MMiller (WMF) (talk) 21:28, 19 August 2026 (UTC)Reply
I rarely improve tone and am no expert on it but I looked at the first three recent changes by different editors. The first looks like a useful improvement from a new editor who clearly already has the skills we need. The second slightly misses the point: I've seen the film and the character's defining aspect is that he is a local legend. The third is also wide of the mark: much of the promotion is in the paragraph before the one the editor changed. Its edit summary of "changed tone(bot told me to)" is also concerning: perhaps the editor values obeying "the bot" above using their judgement or feels that this is what is required to become an accepted editor.
Having our cake and eating it would be nice, but I do tend towards constructive edits over using articles as a sandbox. Certes (talk) 21:46, 19 August 2026 (UTC)Reply
Taking a look at the tone edits, starting at the bottom:
  • Special:Diff/1370168493: The edit in isolation is ok but the edit summary suggests that it is AI-generated, and the many other edits they have pumped out such as Special:Diff/1369053156 and especially this promotional draft corroborate that. The ideal response here would be for them to stop using AI for promotional edits. I would also suspect an SPA based on this edit history.
  • Special:Diff/1370225916: Probably OK, no immediate red flags
  • Special:Diff/1370168654: Obvious promotional AI slop.
  • Special:Diff/1370168718: Probably OK, no immediate red flags
  • Special:Diff/1370168757: Obvious promotional AI slop by the same person who did the last one, which should illustrate the volume at which this adds AI slop to the wiki. (Also, the topic is potentially controversial.)
  • Special:Diff/1370169145: Obvious promotional AI slop by the same person, I'm going to skip their edits from here on out but just know that I am skipping a lot of bad edits as a result.
  • Special:Diff/1370170671: OK but not really a tone edit, and their edit history is somewhat suspect. Also, Special:Diff/1369948373 does not actually fix the tone, it just puts a band-aid over it, which is the other problem with these edits.
  • Special:Diff/1370171161: Probably OK.
  • Special:Diff/1370171537: Promotional AI slop (by someone else this time, Special:Diff/1350356412 is the smoking gun and they have several warnings on their talk page). As you can see from their edit history they are also pumping these out at high volume so I will also be skipping their many bad edits.
  • Special:Diff/1370171930: Grammar edit that does not actually fix the tone, it is still promotional
  • Special:Diff/1370172403: Not a tone issue. The fact that there is an obvious tone issue literally one word away does not speak highly of their competence in editing. (that is, competence in editing English-language writing, not Wikipedia specifically)
  • Special:Diff/1370174754: Probably OK, no red flags
  • Special:Diff/1370176147: Probably OK, no red flags
  • Special:Diff/1370176495: Obvious promotional AI slop and they didn't even bother removing the citation markers.
This is the thing that I have been trying to point out for several months now. The feature is a net negative. It may result in numbers going up in terms of edits by newcomers, but the workload for editors also goes up -- if they even notice it -- to a point where there are simply not enough people available to cleanup. This also means that revert rates are artificially low, because again, there are not enough people to even see them. Gnomingstuff (talk) 22:07, 19 August 2026 (UTC)Reply
Okay, well as a third person who has randomly gone through some of these:
  1. diff: change is an improvement; the paragraph is unsourced, and might be better removed. (My changes: pt 1, pt 2, pt 3; still room for improvement.)
  2. diff: may or may not be an improvement; I haven't seen the film, and I doubt the editor had either.
  3. diff: possibly worse, though I haven't seen the show.
  4. diff: likely improvement. Removes unsourced content.
  5. diff: minor but definite improvement. Edit summary includes "bot told me to". I'd probably also remove "even" or reword, but "one of the largest" is likely factual, albeit not sourced inline.
  6. diff: could be worded better but it's an improvement.
  7. diff: minor improvement, room for more work.
  8. diff: meh, at least there's a period at the end of the sentence now.
  9. diff: not really an improvement.
  10. diff: improvement.
A common theme is that many of these seem to involve people changing tone without the background knowledge to understand the article. "Tone" is also vastly oversimplifying the variety of problems these articles have. LittlePuppers (talk) 22:15, 19 August 2026 (UTC)Reply
I actually have seen the show; the former text was an accurate description of the character, especially given that the characters in Sunny are... extreme and exaggerated people. So this person and/or any hypothetical AI they are using does not know the difference between fictional characters and real people.
The other issue -- and this is an issue across the board with all newcomer tasks -- is that the template on the article is actually more specific than revising tone: it says that the article is written in a primarily in-universe style and should not be. I assume they were not told that in the task. Gnomingstuff (talk) 22:22, 19 August 2026 (UTC)Reply
Re: newcomers vs experienced editors this is partially covered by Peter's comment above. I.e. These features are entirely configurable (and extensible) by each local community, and importantly, that means that some of the types of Suggestion can be completely limited so they're only seen by highly-experienced editors. If you/anyone can think of types of Suggestions that would be widely useful for just experienced users, and if those Suggestions can be programmatically recognized by simple textmatching, or more complicated types of code within the extension code, or via a locally run/controlled LLM, then it should be possible to setup those kinds of things (to test, and if proven useful then to make available to all editors who fit the locally defined configuration).
Also, one of the main goals of the feature is to help provide editors (both newcomer and experienced editors) with a handy link to the relevant guideline/policy within the Suggestion card's text. That makes it easier for editors to learn (or remind ourselves) of the specific nuances involved in any particular fix. E.g. If someone is editing a disambig page, and they ignored the EditNotice reminding them of the basic guidelines, then when they add two links within a single entry (or stumble upon an existing entry which does so during their editing), there could be a Suggestion that informs them of that guideline and has a singular pointer to the details (versus the nine links in the current Enwiki Editnotice). I.e. Targeted micro-tutorials that show up when relevant. The main limitation for what can be written in the Suggestion card, is length, as it also needs to work in a mobile-sized screen. Quiddity (WMF) (talk) 22:53, 19 August 2026 (UTC)Reply
Just want to reiterate that I appreciate the response here, it is already a much better dialogue than before and that's what I was trying for, to keep the focus on the content and implementation details itself. It may surprise some people here but I am not blanket anti-AI across the board no matter what. I just don't think that providing them to newcomers who don't have a sense of our AI guidelines is a good idea. (For more on that reread GreenLipstickLesbian's comment)
Some things I don't think I've seen mentioned:
  • Right now, this and other task suggestions seem to target tagged/templated articles with high view counts. I think both parts of this are a mistake. People template articles because they want to make them better and, by and large, the result has been templated articles getting either worse, or impossible to tease apart good vs. bad edits due to the sheer volume of them. Something like this might look good from an edit metrics standpoint but from the standpoint of doing patrol it is a lot of patrol work -- and that's just the entry point to one article. Something like that can easily spawn 5 more tabs to check if someone was indeed doing high-volume AI edits. And of course the higher the view count, the higher stakes any mistakes become. (I will say that Paste Check has been somewhat helpful, it technically creates more patrol work but it's more like pointing to patrol work that would exist no matter what)
  • If the text parsing can't be fixed (Though it should be), the fail states are systematic enough that you can probably just regex filter somewhere in the process.
  • Controversial material needs to be blacklisted for reasons that should be clear. Admittedly the three examples above were somewhat cherry picked to demonstrate that point, but they were not especially hard to cherry pick. (The third one was just CTRL-F "partisan" because I already knew LLMs have issues with that, and sure enough there they were.) This should also be low-tech. Something like blacklisting CTOP and BLP articles -- I know Simple Summaries blacklisted BLP articles, though there were a lot of holes in that implementation -- and more importantly, an additional filter at the sentence level that is just blunt keyword-based stuff. This should also make MOS:GEO suggestions much safer since it's really hard bordering on impossible to predict every way that can go wrong.
  • I don't know what level of LLM familiarity the people working on this have, but I assume it is more in the LLM-coding-in-general realm, less the intersection of LLM output and Wikipedia. So, seconding getting in touch with the people at WP:AIT, they've been iterating on stuff like this for a while and have a sense of what does and does not work. I also can't speak for everyone, but while the people on AI Cleanup are less likely to be open to AI and even less likely to have any spare time to help out with stuff, they do have more of a front-row seat to AI edits and edit suggestions than basically anyone else. Most if not all of the issues above were predictable when you've seen a lot of LLM edits on Wikipedia. I have only found one common LLM edit problem that hasn't shown up in these (which is actually kind of interesting from a model standpoint). Feel free to email me, I have a great deal of collected data on this.
  • This is cheating since I've mentioned it, but "more QA" seems to be a constant across the board for all features. Another place where getting in touch with AI folks can help because spot checking goes a lot faster when you know what to look out for.
(Also, I know this isn't something you have any way of knowing about, but I prefer to not be pinged to ongoing discussions I am participating in, it just generates notifications and emails that pile up). Gnomingstuff (talk) 18:28, 19 August 2026 (UTC)Reply

Where can you see all of the model-generated MoS suggestions?

[edit]

Okay. Linked here you will find a spreadsheet that contains the batch of ~6,500 experimental LLM MoS suggestions as they appear in Suggestion Mode for people who enabled experimental suggestions.

Within the spreadsheet are two categories of suggestions: those related to simplifying language and MOS:GEO. We've removed all of the NPOV suggestions based on the feedback y'all have been helpfully sharing here.

How do these suggestions look to you? What are examples of specific suggestions that you find to be unhelpful, confusing, and/or just plain wrong? As you're going through these suggestions, what broader patterns are you noticing/conclusions are you reaching about these two categories of suggestions?

With this feedback in-hand, we're thinking we (staff and volunteers) can take a step back together and decide how and if we should move forward with this particular set of suggestions.

Of course, if you find any part(s) of the spreadsheet are unclear, please let us know.

A couple of notes:

  1. We welcome feedback in whatever form is most convenient for you. E.g. sharing directly in this discussion, enabling experimental suggestions in Suggestion Mode (see instructions) and offering feedback through the UI, etc.
  2. Some suggestions may be present in the spreadsheet and not visible when you edit the actual article with experimental suggestions enabled. This is because this spreadsheet contains suggestions that were generated in a batch offline. In the time since, some articles have been edited in ways that make the suggestions obsolete.

PPelberg (WMF) (talk) 22:35, 19 August 2026 (UTC)Reply

Thank you! Will take a look Gnomingstuff (talk) 22:43, 19 August 2026 (UTC)Reply
You bet and thank you! PPelberg (WMF) (talk) 22:49, 19 August 2026 (UTC)Reply
Is it worth implementing some sort of minimum length check for simplify language? I am not sure how much shorter "Scorer for Crystal Palace; Terry Fenwick" or "Town in Faisalabad District" can be. The model seems to feel parentheticals are a lot more complex than they are in reality: "Romelda Aiken-George (née Aiken}; born 19 November 1988) is a Jamaican netball player." It also appears to be picking up some template code or something ([ { "List of leaders of Markazi Jamiat Ahle Hadith": "Order" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "1" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "2" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" } ]) as well as picking up lists as prose ("See also
List of prime ministers of Pakistan List of presidents of Pakistan Chief Secretary Khyber Pakhtunkhwa List of chief ministers of Punjab List of chief ministers of Sindh List of chief ministers of Balochistan". It's even picked up the reference list for 1994_Andhra_Pradesh_Legislative_Assembly_election as in need of simplification.If the list could be refined to remove the simply technically non-applicable, it would be easier to parse. I would be interested in a few examples worked through by editors here who have spent a lot of time looking into language simplification (@Femke?). For example, "It is the feminine form of the Late Latin name Clarus which meant "clear, bright, famous"" and "A space station (or orbital station) is a spacecraft which remains in orbit and hosts humans for extended periods of time" seem near fully concise to me, but I have not looked into this topic much. (Looking now as I copy these in, these may also represent examples of the model looking at too short a sentence and being confused by parentheticals.) CMD (talk) 02:29, 20 August 2026 (UTC)Reply
In this case, cross-referencing against the csv I have, the AI's "suggestion" is Fix the misplaced bracket and punctuation in the parenthetical expression for the maiden name. Which actually is a legitimate fix -- there's a stray curly brace -- but isn't related to simplification at all, see my comment about it being a catch-all.
I know it might be confusing and the spreadsheet here is supposed to mimic what editors see, but for QA purposes having those around without having to cross-reference might be helpful to pin down what's going on Gnomingstuff (talk) 03:04, 20 August 2026 (UTC)Reply
Thanks, same issue as a lot of the MOS:GEO examples then in the llm finding something that is out of its supposed scope leading to a confusing tag. CMD (talk) 03:07, 20 August 2026 (UTC)Reply
@Chipmunkdavis: thank you for reviewing! Responses to the points I see you raising...
Is it worth implementing some sort of minimum length check for simplify language?
Great question. Assuming we were to agree this suggestion was worth moving forward with, I think we could implement a way for you all to set a minimumCharacters value as is currently available for Reference Check. cc @DLynch (WMF) who can say definitively whether this is feasible.
The model seems to feel parentheticals are a lot more complex than they are in reality... as well as picking up lists as prose...
Great spots and noted. We'll investigate whether there are things we could do to exclude these two category of issues you're describing. PPelberg (WMF) (talk) 23:35, 20 August 2026 (UTC)Reply
Sure, there's two ways we could do it. The equivalent to the reference check would involve doing it client-side -- you'd set a config value in Editcheck-config.json on-wiki and we'd filter out suggestions that we get from the API that're below whatever length you specify. The other way would be to configure the model to not generate those suggestions in the first place, which would be more-ideal overall in terms of work saved. Harder for you to configure though. DLynch (WMF) (talk) 23:41, 20 August 2026 (UTC)Reply
By the way, for the "simplify language" suggestions I was wondering if existing small language models could do a similar job, so I bashed together something (I used the existing WMF readability model, m:Machine learning models/Production/Multilingual readability model card, though it isn't really designed for sentence level ARA... so possibly a way to improve things if a more suitable model can be found). I did have Gemini 3.6 Flash screen the about 2/3rds of the suggestions and fill out the description field (columns are same as the CSV instead of the spreadsheets) since I haven't investigated which features produced the scores, not sure how useful those descriptions would be. I've posted the suggestions it generated here if anyone is interested in looking at them and comparing to the ones generated by GPT-OSS. Alpha3031 (t • c) 15:16, 22 August 2026 (UTC)Reply

What if communities don't want certain suggestions on their wiki?

[edit]

If a wiki does not want a certain suggestion to be shown to anyone on their wiki, any administrator or interface administrator can disable it directly, without requiring any code changes/action from WMF staff.

In practice, this would look like someone visiting MediaWiki:Editcheck-config.json, identifying the Edit Suggestion or Check they'd like to disable, and changing ShowAsCheck or showAsSuggestion from True to False. We're of course happy to help out/clarify any confusion. Although, the goal is for this system to be intuitive enough for you all to adjust it based on what you're seeing in practice and the expertise you've developed.

In addition to toggling Checks and Suggestions ON/OFF, there are a range of other ways you all can customize them:

  • Who sees them. account limits a Check or Suggestion to people who are logged-in or logged-out and minimumEditCount/maximumEditCount enables you to target them based on someone's editing experience.
  • Where within an article they appear. ignoreSections excludes Checks/Suggestions from appearing within specific section headings and includeSections does the inverse, restricting a Check/Suggestion to only appear within the sections you define.
  • What kinds of pages and content they run on. ignoreDisambiguationPages makes it so a Check/Suggestion does not appear on any disambiguation pages, ignoreQuotedContent prevents them from appearing on text inside of quotation marks or blockquotes, and inCategory/notInCategory and hasTemplate/lacksTemplate scopes Checks/Suggestions based on the categories and templates present/absent within an article. Note: there's a task for enabling configuration based on properties of an article's talk page too.
  • How sensitive an individual Check is. These vary by Check/Suggestion. For Reference Check, minimumCharacters enables you to define how much new text someone will need to have added for it to get activated. For inference-based Checks/Suggestions like link suggestions, predictionThreshold enables you to set confidence threshold the model must reach before anything is shown to someone.
  • Entirely new, locally-defined Checks. TextMatch lets you define entirely new Checks, locally using pattern matching, without any code changes from us.

At any point in time, you can visit Special:EditChecks to see the current set of Checks and Suggestions that are available and how they're configured. This page is updated automatically as code changes are merged. In the future, we'd like to add metrics so that you have even more visibility into how things are working and identify what might need adjustment.

Might there be ways you'd like to be able to configure Checks/Suggestions that we do not currently offer? Is there information that you'd like to see on Special:EditChecks that's not currently visible? More broadly, does how we're thinking about this line up with how you're thinking about it? We're open and eager to any and all feedback. It's important to us that you have the visibility into, and control over, this system to make it work for your wiki (and the same for volunteers at other wikis).

---

i. See the distinction between Edit Checks and Suggestions that Marshall posted above. PPelberg (WMF) (talk) 23:51, 19 August 2026 (UTC)Reply

What kind of communication with communities has there been on this project so far?

[edit]

In a June meeting hosted in the Wikimedia Discord, members of the Editing and Machine Learning Teams shared (discord link) that we were experimenting with LLM-generated MoS suggestions.

In that conversation, we mentioned that this initial experiment would not involve surfacing any actual edit suggestions to volunteers. Instead, we would show edit suggestions in an "experimental" state that invited the people who opted-into seeing them to offer feedback about the quality of the suggestions. Then after that feedback process, we (volunteers + staff) could decide whether any of them were worth actually showing to people via a controlled experiment.

In the time between that call and now, we shared progress updates on the MediaWiki project page and had planned to announce the existence of this work and invite feedback about it on-wiki this week or next.

I appreciate that the above still led us into a situation where many of y'all were caught off guard by this work and as a result, trying to put the pieces together in real-time. This is not ideal for y'all and it's not ideal for us!

In this thread we've seen folks like @Alpha3031, @Chaotic Enby, @SJ, and @Kowal2701 helpfully suggesting that work like this would be better off were we to be in touch about it early and often. This way, we can get on the same page about what ideas might be non-starters and for those that we do deem worthwhile to pursue, to align on when and how we'll evaluate them together.[1] [2] [3] [4]

This all leads me to wonder…

Let's imagine we (staff) have identified a new LLM-powered suggestion that we think would be worthwhile to experiment with. When would you all appreciate us checking in with you all about it? How much information/example/context do we need to prepare before it's useful for a large group to evaluate an idea? Where do you think would be the best place for us to post about this?

Asked as another way: if we could do this all over again, what would you imagine the "ideal" process to be? PPelberg (WMF) (talk) 23:08, 20 August 2026 (UTC)Reply

Not sure this is a communication issue, so much as a content and quality issue. If a feature is good then communication about it will be received more positively no matter what form it takes, whereas if a feature is not good then there is no way to announce it in a way that will be received well. Gnomingstuff (talk) 03:22, 21 August 2026 (UTC)Reply
To me, the ideal process would be asking the community first what LLM-powered suggestions they would be most interested in and would find acceptable. It might also avoid issues with setting up suggestions that may swerve into CTOPs and CTOP-adjacent spaces that you might not be aware of. There's plenty of low hanging fruit for LLM suggestions that I think would have a reasonable amount of support, but any time you're going to have an LLM suggest things, especially to new editors, that have resulted in arbitration cases and site bans you're probably looking at the wrong things. ScottishFinnishRadish (talk) 11:25, 21 August 2026 (UTC)Reply
Hello Peter, agreed with SFR + Gnomingstuff that this is about quality and collaboration, and how to start with easy cases when implementing any new interface or workflow. Sharing examples as they are produced, and working in public on essential parameters like what eval is used and the threshold for what gets presented, would make it easy for people to give targeted feedback early and often, saving everyone time.
* Give more attention to the choice of areas that will be covered. Have a page for suggestions, with definitions slightly more detailed than a one-sentence description. (what is the scope of each, defined how, tested against which guidelines)
* One step in the review checklist should tease out what might be easy or hard or surprising in each area.
* Spend time on better evals. Have human evals along with LLM-judge evals of initial suggestions.
* Use a tool for bulk elicitation and evaluation of suggestions (like an editable spreadsheet); working through the VE interface is too slow for meaningful human review at scale, and when reviewers find a type of error they will want to find many instances of it to fully characterize what's going wrong.
* Fine tune the models used on feedback.
This doesn't need to be a large group; small groups working in public on a predictable cadence is fine. The above is helpful regardless of what is generating the suggestions (LLM or other tool). It shouldn't be 'checking in' ⸺ this is part of the editorial workflow! There's no lack of interest in finding low-hanging fruit that works. I propose a better starting premise would be "we (all) have identified a suggestion that may be worthwhile to experiment with"... which is possible once there is an active page for suggestions. – SJ + 15:37, 21 August 2026 (UTC)Reply
Seconding all of these Gnomingstuff (talk) 06:14, 22 August 2026 (UTC)Reply
Also, you might let people test + give feedback on the interface separately from particular suggestions. For instance, using this to highlight {cn} instances on a page or other inline flags that are already present in the wikitext but harder to discover. That could happen continuously while studying low-hanging suggestions above. – SJ + 15:37, 21 August 2026 (UTC)Reply
@ScottishFinnishRadish and @Sj: thank you for offering these clear recommendations for how we might work more effectively together going forward.[i] And thank you Gnomingstuff for stating clearly that no amount of communication substitutes for the underlying quality of the work.
Next week, you can expect me to follow up with the concrete ways we're thinking about integrating the feedback people have shared in this discussion. Until then, thank you for the thought and attention you've offered this week...we continue to learn a great deal in the process!
---
i. E.g. create a local project page where we can generate and evaluate ideas for new suggestions together, make evaluation easier to do at-scale, avoid contentious topics when considering suggestions to pursue, etc. PPelberg (WMF) (talk) 23:25, 21 August 2026 (UTC)Reply
Next week, you can expect me to follow up with the concrete ways we're thinking about integrating the feedback people have shared in this discussion.
Hi y'all – there are at least two ways the work will proceed from here:
First, on the current batch of MoS Suggestions: we’re going back through this discussion to compile a list of issues that we can investigate and hopefully, use to improve the MOS:GEO and Simplify Language suggestions (spreadsheet).
Second, on process: we're going to draft a proposal for where/how we can evaluate ideas for new suggestions together (e.g. copyediting) and metrics we can use to assess the quality of suggestions once a dataset is available.
You can expect another message to Village Pump as soon as we have updates on the above.
In the meantime, thank you all for this helpful feedback. PPelberg (WMF) (talk) 18:25, 26 August 2026 (UTC)Reply
– SJ + 13:22, 27 August 2026 (UTC)Reply
Re: communication, I think it depends on the context. Is it something suggested by several editors on-wiki? Is it very similar to an existing feature? Is it similar to an existing user script? Go for it. Have similar ideas been controversial in the past? Is it a completely new approach to something? More discussion with the community as you design and develop it would probably be good. If in doubt, just drop a "hey, what do you guys think about an AI tool to detect NPOV issues?" on a noticeboard somewhere. It doesn't need to be some elaborate presentation. (I know big communities and corporate cultures can both make some of those things hard, but aspirationally I wish we had a culture—both here and at the WMF—that made that not a big deal and encouraged the communication.)
The best place is on wiki, either on this page or (for more specific features) on a relevant talk page. This is where we all are, and the same can't really be said for anywhere else. (Disclaimer that I can't speak for other projects.) And if you're not sure, it's generally pretty easy to ask "where is a good place for this" or just to drop a link on a few pages. LittlePuppers (talk) 03:50, 22 August 2026 (UTC)Reply
I'd advise against relying on suggested by several editors on-wiki. I'm sure you could find more than several who would support any and all LLM implementation across the board. "Consensus between multiple experienced editors" might be better phrasing, but I think similarly to existing tools that have general acceptance is a safer metric for proceeding without additional discussion. Anything that's novel should probably receive community input. ChompyTheGogoat (talk) 21:09, 27 August 2026 (UTC)Reply

Given the WMF seems inclined to continue developing these tools regardless of what we say, I propose we hold an RfC establishing that LLM tools may not be deployed without an explicit consensus from the community. If editors think this may be a good idea, I suggest the following as the first draft of the question:

Should we require an explicit consensus before the WMF is permitted to deploy tools involving the use of an LLM? The WMF may deploy tools for user testing, so long as all of the following criteria are met:

A: Testing of the tool concludes after no more than 12 months, after which the tool must be removed unless there is a community consensus for either extended testing or full deployment.
B: Use of the tool is restricted to editors who have opted-in
C: Use of the tool is restricted to editors who have been granted the LLM-tool user right. This would be a new user right created and granted by admins at WP:PERM.

Does anyone see issues with this proposed question? Are there any revisions? BilledMammal (talk) 07:42, 21 August 2026 (UTC)Reply

I'm not fundamentally opposed to this, but I think we need a lot more clarity about what this new LLM-tool user right would mean. A naive reading of the name would be "This user is exempt from WP:LLM" which I suspect is a broader interpretation than you intended. RoySmith (talk) 12:53, 21 August 2026 (UTC)Reply
This seems like policy creep, and not a good approach (we shouldn't have blanket limitations based on the "type of tool" involved). I believe it is also misdirected: the example above is testing what would be an entirely community-configurable tool. Opt-in is appropriate for things that are so new, and I'd say anything that messes with your margins should be configurable. But proliferating user rights is an anti-pattern, and should only be used as a last resort. – SJ + 16:21, 21 August 2026 (UTC)Reply
yeah, I think this may be premature for now, but worth revisiting if there are any disasters Kowal2701 (talk, contribs) 18:24, 22 August 2026 (UTC)Reply
Consistent with the current wording of NOLLM and with active work being done - so far successfully - at helping limit abuse, I would suggest some cleaving of administrative actions and content actions. Best, Barkeep49 (talk) 17:05, 21 August 2026 (UTC)Reply
Solely regarding access control (I'm undecided regarding the overall proposal): I agree with the other commenters that that creating a new user right isn't the best fit. I think tool-specific JSON lists of approved users that are fully protected could suffice, and be more adaptable to allow for per-tool authorization. isaacl (talk) 17:15, 21 August 2026 (UTC)Reply
The whitelist system works well for AWB/JWB. Certes (talk) 21:13, 27 August 2026 (UTC)Reply
I also think that the community should be able to control and configure the tools used in en wiki but I'm not sure an rfc is needed atm and its wording unfortunately isn't clear.
What is a "tool"? What else, other than edit suggestions, is in the scope?
12 months deadline is arbitrary: it's too long for a tool that the community actively opposes and may be too short for an iterated testing of a complex idea.
Some abuse/vandalism prevention tools by definition cannot be opt in. Alaexis¿question? 20:18, 21 August 2026 (UTC)Reply
It's hard to guess what is a tool, because the AI is developed secretly and added to Wikipedia's software or configuration as a fait accompli. Certes (talk) 21:13, 27 August 2026 (UTC)Reply
I agree with most of the comments above that this is not a well-defined RfC. In addition, the WMF have said this will be community-configurable if it makes it into production so it appears to be unnecessary anyway. Mike Christie (talk - contribs - library) 18:34, 22 August 2026 (UTC)Reply
I assume this is intended to apply to any future implementation that would integrate LLM features in any way, not just this one concept, which I fully support. Pushing it on us without community consent is inappropriate and exactly the type of forced AI we're seeing on every other platform. Wikipedia is supposed to be different.I'd agree on amending the deadline - I'm not sure what an appropriate initial test phase would look like, but I'd recommend a limit aligned with that and only extended or fully implemented via consensus. Just enough to get a feel for it, not a full beta phase to polish it for release. People with more programming experience than me would probably have a better notion of the timeline.I do think we need some limitation on access (especially if it is fully implemented) and I think Isaacl's suggestion sounds like a better way to handle it. ChompyTheGogoat (talk) 20:58, 27 August 2026 (UTC)Reply
I think that some level of access control is warranted, but the idea of getting individual approval for every tool that the WMF wants to test would impede testing numbers without much safety gain over a single user list for all beta tools.

Having a time limit on testing is strange, but my exact opinion on it depends on how community consensus is defined here. I think something that both mitigates the issue i think you're trying to solve while not dictating the WMF's development calendar would be something like "Any tool in testing shall be disabled if a consensus is formed at the tool's thread on VPWMF that the tool is causing disruption. At which point the tool will only be re-enabled with consensus." Although, if we got anything near consensus that a tool in testing is disruptive, I think the team working on the tool would disable/fix it very quickly. These are people who chose to work on mediawiki here, not WMF management. MetalBreaksAndBends (One for all) 01:32, 28 August 2026 (UTC)Reply
I think it's only suggesting permissions for LLM tools that would potentially be more prone to abuse, not for any and all in testing. ChompyTheGogoat (talk) 07:13, 28 August 2026 (UTC)Reply
There isn't any limiting language in the comment, so I (and most people, probably) would think it would apply to all LLM tools. MetalBreaksAndBends (One for all) 23:42, 28 August 2026 (UTC)Reply
That is what I meant - all LLM, not all tools period. If we end up with so many LLM tools being tested that they're hard to keep track of I'd consider that a problem unto itself. ChompyTheGogoat (talk) 02:50, 29 August 2026 (UTC)Reply
I'm not saying the issue is that they could be hard to keep track of (though over a long enough time period it could), I'm saying that applying for every beta is an unnecessary hassle. MetalBreaksAndBends (One for all) 03:38, 29 August 2026 (UTC)Reply
I think I would share the view that a full, formal RFC would be unnecessary if prior discussion shows clear consensus to implement, for example (assuming it's advertised to a noticeboard like this one or the cleanup wikiproject). I would expect any community members participating in such discussions to be sufficiently in touch with current attitudes towards LLM tools to bring up anything that would seem potentially controversial and standard pre-RFC discussion processes should function fine (w formal closures and move to full RFC if no clear consensus or if there is consensus there should be an RFC in the specific discussion). Alpha3031 (t • c) 04:05, 29 August 2026 (UTC)Reply
Do you mean a full rfc for each tool or an RFC for this proposal? MetalBreaksAndBends (One for all) 06:25, 29 August 2026 (UTC)Reply

Hey WMF: nice job on communicating in this section

[edit]

I think this conversation has been healthier than some other recent enwiki–WMF conversations. Just wanted to say thanks to the WMFers that are participating here and doing things like compromising (i.e. getting rid of the NPOV model), replying a lot instead of making one polished statement then leaving, and speaking clearly and honestly.

It's tough because on some issues, WMF and enwiki are very out of sync. So even if everyone does everything right, these conversations may still be tough. But doing things like compromising, having conversations with us that aren't just statements, and speaking clearly and honestly are definitely a good approach. Please keep it up. –Novem Linguae (talk) 19:23, 22 August 2026 (UTC)Reply

Absolutely seconding this! Really happy to see WMF folks take community feedback into account. I understand the task can be much harder than it seems at first, especially as these discussions get sprawling and it can be hard to find a thread that unites the whole range of community opinions together, let alone incorporate it in the team's plans for the project.The way you managed to navigate it was a very positive surprise, and I'm looking forward to more productive exchanges from both sides! Chaotic Enby (in solidarity · talk · contribs) 19:45, 22 August 2026 (UTC)Reply
Yes, absolutely. Ymblanter (talk) 08:10, 23 August 2026 (UTC)Reply
Absolutely agreed on all of this Gnomingstuff (talk) 19:01, 23 August 2026 (UTC)Reply
+1. Can't speak for anyone else, but pretty much all my complaints about WMF's poor communication are limited to the Trustees and a few Officers. Everyone else at the WMF seems to communicate fine. MMiller and PPelberg are two usernames (among others) I've grown very accustomed to seeing regularly on-wiki, they've communicated often and effectively with volunteers for years. Levivich (talk) 20:22, 23 August 2026 (UTC)Reply
Thank you -- we're glad to hear it -- we are trying hard! Although there are going to keep being times that we all disagree on ideas, the most important thing is that we can discuss constructively to figure out how to best improve/adapt the wikis. Thank you all for being here for these conversations, on top of doing all your usual wiki work. MMiller (WMF) (talk) 05:20, 24 August 2026 (UTC)Reply
I am of the opinion that most -- maybe all -- WMF employees who are not in top management are competent, helpful, want to do the right thing, and are eager to communicate. I attribute the stonewalling we often see with a perfectly reasonable fear that actually having a dialog with the volunteers will never help your career and just might get you fired. If only there was some way that WMF workers could organize and join an entity that works to protect them from being unjustly fired... Let me know if anyone has ever heard about something like that. --Guy Macon (talk) 12:38, 24 August 2026 (UTC)Reply
I find this post amusing. But as a Wikipedian, I can't help but be myself and note that if you define top management as something beyond "C Suite" Marshall is would probably be considered top management. Best, Barkeep49 (talk) 14:47, 24 August 2026 (UTC)Reply
+Whatever we’re on In solidarity Wikipedian12512 (Talking is fine | contribs) 01:05, 6 September 2026 (UTC)Reply

Hi y'all – one outcome of the recent AI-generated edit suggestions discussions was reinforcing the need for us to be in touch with you all early in the process of exploring new inference-based suggestions. This way, we can discuss the experimental suggestion's risks and decide whether en.wiki might be a good place to evaluate its reliability before considering a wider deployment.

We're now at this point with a new experimental suggestion that we'd value your feedback on.

Inspired by the volunteer-authored AI Source Verifier, and how helpful many experienced editors have found it to be in identifying cases where a citation might not support the claim it's attached to, the Foundation's Research, Machine Learning, and Editing teams are working with Alaexis to try integrating this script as a "suggestion", only visible to experienced volunteers who have opted into seeing experimental suggestions in Suggestion Mode.

Below, you will find more information about how this proof of concept will work and the input we are needing from you all. Before that, a note on why we're prioritizing work on this suggestion right now…

We are prioritizing this exploratory source verification work in response to hearing from volunteers:

  1. How tedious it can be to find potentially unverified claims within an article (e.g. read the claim, click on each source, ensure you can access each source, etc.)
  2. How source verification work is becoming even more important and prevalent as AI increases both A) the ease with which people can add new content to Wikipedia and B) the risk that said content is not supported by the sources it's accompanied by (i.e. contains hallucinated references).

While experienced editors will remain responsible for evaluating whether a source verifies its associated claim, we are seeking to learn whether a tool like this could make the mechanical parts of this wiki work less toilsome.

And in case you're curious, we did a bit of digging to put some numbers to all of this:

  • English Wikipedia has more than 70 million citations.[1]
  • In one study, annotators worked through a few hundred claims and the web pages cited for them, and judged 12% not supported by the source at all.[2]
  • A separate study found that at least 3.3% of facts on English Wikipedia contradict another fact elsewhere on the project.[3]

Note: Neither of those sets represents the encyclopedia as a whole, so the real number is not currently known to us.

How it works

This experimental suggestion will:

  1. Identify all inline URL-based citation(s) in the article
  2. Extract the text that precedes said citation(s)
  3. Retrieve the content of each cited source (and indicate if it's unsuccessful in doing so)
  4. Use an open weight Qwen3.6-27B model, hosted on Wikimedia's LiftWing infrastructure, to compare the claim against the source's contents
  5. Flag claims that the model has deemed to be only partially supported, not supported by omission, or not supported by contradiction.
    1. Note: The model will accompany each conclusion with the passage(s) from the source it's based on, except where the conclusion is "not supported by omission"

Note: This suggestion would not make edits to Wikipedia and can be configured, like other Edit Checks and Suggestions, to show/not show based on a variety of conditions.

We need your input

We will soon be ready to generate an initial experimental dataset that volunteers can use to offer feedback about the usefulness and reliability of this suggestion. First, we need to decide which wikis/languages to include in this dataset.

This leads us to wonder:

  1. Would any of you all be interested in evaluating a batch of these experimental suggestions for en.wiki articles? Note: The suggestions would be made available to you all in a spreadsheet and as a suggestion within Suggestion Mode, visible only to volunteers who have published ≥100 edits and have opted into experimental suggestions.
  2. If so, what types of articles do you think would be helpful for us to include within this dataset? E.g. new articles, articles of a certain quality, etc.

For anyone interested in seeing a demo and talking about this in a voice call, we will be hosting a meeting in the Wikimedia Community Discord on 14 Sep 2026 from 17:00 - 18:00 UTC. We'll of course be responsive here as well.

In the meantime, you can get a sense for how the suggestion works by installing the user script that User:Alaexis and User:LuisVilla have been maintaining.

References

  1. ↑ Mario Morvan, "Citation Location Needed", 17 May 2026; measured from 20,000 randomly sampled articles across dated dumps.
  2. ↑ Kamoi, R.; Goyal, T.; Rodriguez, J.; Durrett, G. WiCE: Real-World Entailment for Claims in Wikipedia. EMNLP 2023. pp. 7561–7583.
  3. ↑ Semnani, Sina J.; Burapacheep, Jirayu; Khatua, Arpandeep; Atchariyachanvanit, Thanawan; Wang, Zheng; Lam, Monica S. Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models. EMNLP 2025. arXiv:2509.23233.

Based on some of the questions that volunteers raised in the previous discussion about LLM-backed suggestions.

Expand to read the FAQ

1. What kind of communication with communities has there been on this project so far?

This is the first message we've put on-wiki about this work. Note: we've mentioned this work in previous discussions [1][2][3] and have been discussing implementation details in Phabricator.

2. How did we / do we QA lists of suggestions like these?

Quality assurance of this suggestion will happen in two phases, both of which will start once we know what wiki(s) would like to participate in this pilot:

  • Internal review: Before any dataset is shared with volunteers, the development team will review a subset of the suggestions and confirm no basic issues are present. We will attempt to address issues we find, and in cases where we can't, will share them with you all to consider as part of the evaluation you will be doing, should this work move forward at en.wiki.
  • Volunteer review: Once the internal review is complete, we will make the dataset available for y'all to review in two forms:
    • A spreadsheet so you can identify and categorize error patterns in bulk, rather than one-at-a-time. Note: A spreadsheet is likely not a long-term solution. Though, we think it's a useful first step.
    • Within Suggestion Mode (as an experimental suggestion) so that experienced editors (you all) can see how the suggestions would look and work in practice.

How does this sound to you? What (if anything) about the above do you think we could make clearer? Might there be step(s) you think we ought to consider adding?

3. What determines whether suggestions are good enough to go beyond testing?

Two inputs will determine whether a suggestion is good enough to go beyond testing:

  1. Qualitative feedback volunteers will share from reviewing the dataset (spreadsheet)
  2. Quantitative feedback about how experienced editors who opted-in to seeing the experimental suggestion engage with it (i.e. the rate at which they engage with and dismiss the suggestion). This information will be tracked using data logging we've implemented within Suggestion Mode.

Note: Prior to beginning the volunteer evaluation, we will share, and invite your feedback about, a concrete proposal for what thresholds we think need to be met for us to collectively consider the suggestion good enough to move forward.

Might there be existing gadgets/scripts that have gone through similar types of volunteer testing that you think would be useful for us to learn from?

4. Is source verification appropriate to point an LLM at, given its nuance and complexity?

Partly. 10 months of volunteer testing of the AI source verification script is leading us to think that LLMs are well suited for assisting volunteers with the mechanical work involved with identifying claims that may need their attention. Things like:

  • Identifying all inline URL-based citation(s) in the article
  • Extracting the text that precedes said citation(s)
  • Retrieving the content of each cited source (and indicating if it's unsuccessful in doing so)
  • Comparing the claim against the source's contents
  • Flagging claims that the model has deemed to be only partially supported, not supported by omission, or not supported by contradiction.

Through this mechanical work, we believe an LLM can help editors hone in on the claims that need their scrutiny. We do not think an LLM is appropriate for making actual determinations. We see this as nuanced and experience-dependent wiki work that volunteers need to continue being responsible for.

What (if any) of this thinking doesnot align with how you're thinking about this?

5. Does it make sense for the message to the world to be "Knowledge is Human", but we are also using AI for certain things?

We think so. Reason being, we see a core tenet of the "Knowledge is Human" message being the fact that the information included within Wikipedia is grounded in sources written by people, deemed reliable by people, and verified to support the claims they are attached to by people.

Volunteers' time is a finite resource, and editing is a time-intensive process. Our goal is to use AI specifically to reduce the toil that doesn't require human judgement — like information retrieval and pattern matching.

We do not see this suggestion as doing anything to change and/or minimize peoples' role in making any of these determinations. Rather, we see it as helping volunteers to decide what source verification wiki work to prioritize.

We'd be curious to know how (if at all) this thinking lands with y'all….

6. What are the known shortcomings of this approach?

One of the major shortcomings of this approach is that it can only check inline, URL-based citations. This means there are entire categories of suggestions (e.g. paywalled articles, books, etc.) that are beyond the scope of this proof of concept.

Should this initial, URL-dependent approach prove viable, we will explore the feasibility of expanding coverage to include more source types.

False positives are another potential shortcoming; the model might flag claims that, upon volunteer review, turn out to be supported by its source. We understand this to be similar to existing on-wiki tools, like Earwig's Copyvio Detector and ORES. Accordingly, we're understanding the key concern is that this suggestion actually saves volunteers time and effort! Said another way: we wouldn't want the false positive rate to be so high that volunteers spend more time finding genuinely unverified claims than fixing them.

How does this sound to you? Might there be other shortcomings you can see with this approach?

Thank you for reading and thinking critically about this work! PPelberg (WMF) (talk) 19:09, 10 September 2026 (UTC)Reply

1. Yes. 2. Is it possible to grab the datasets from those two studies (or any other similar) and cross-reference the pages by article-page categories, talk-page categories, and/or talk-page templates (like WP:CTOP templates), to see if any categories or CTOPs are more error-prone than others? That might help prioritize review of any identified problem areas. I would think WP:BLPs, which are all in Category:Living people, would be a top priority for source verification. Also: articles that are new, haven't been edited for a long time, or are tagged with relevant maintenance templates, like {{dubious}} or {{failed verification}}. Levivich (talk) 19:24, 10 September 2026 (UTC)Reply
Is it possible to grab the datasets from those two studies (or any other similar) and cross-reference the pages by article-page categories, talk-page categories, and/or talk-page templates (like WP:CTOP templates), to see if any categories or CTOPs are more error-prone than others? That might help prioritize review of any identified problem areas.
Oh, I think this is a great question/idea. @Alaexis: do you know if we're able to access the datasets used in the studies we referenced so that we could do the comparison @Levivich is describing?
I would think WP:BLPs, which are all in Category:Living people, would be a top priority for source verification.
This intuitively makes sense to me and to be doubly sure: to what extent (if any) would it be accurate for us to understand you suggesting this because of the following reasons?
  1. Like articles within WP:CTOP, the suggestion performing poorly on BLPs is more consequential relative to articles in other categories like Category:Railway lines
  2. Biographies of living people tend to use news, and other web-accessible sources. This means this suggestion should, in theory, be able to retrieve a larger share of the citations used within them.
Also: articles that are new, haven't been edited for a long time...
Good call. This makes sense to me.
...tagged with relevant maintenance templates, like {{dubious}} or {{failed verification}}
Mmm. In essence, you're saying articles within these two categories offer an existing corpus of claims volunteers have identified as failing verification. Accordingly, running the model against those could help us estimate the model's proficiency. Might I be missing/misinterpreting anything here?
A resulting question that comes to mind as I think about this: might articles tagged with {{dubious}} or {{failed verification}} be more likely to contain offline sources? PPelberg (WMF) (talk) 20:40, 10 September 2026 (UTC)Reply
Some of the datasets are available. In fact Semnani et al have made this analysis themselves Articles in the “history” category exhibit the highest inconsistency rate (17.7%), followed by Everyday Life (16.9%) and Society & Social Sciences (14.3%) (Figure 5). The most common error type in history articles is numerical discrepancy. By contrast, categories requiring precise technical knowledge and quantifiable information—such as Mathematics (5.6%) and Technology (9.4%)—show markedly lower rates.
I like the idea of generating edit suggestions for articles with {{failed verification}} templates for Bayesian reasons - if one such citation has been found it's likely that there are more (aka "if you see one cockroach, there are more" principle). Alaexis¿question? 20:50, 10 September 2026 (UTC)Reply
On why BLP, I actually didn't have either of those reasons in mind, but both are good reasons. I was thinking something similar to #1: content that fails verification (not just poor performance of the tool) is more consequential for BLPs than, eg, railway lines. Of all FVs, those are the ones we should find and fix first.
On why the FV tag, yes, and it could also help us estimate human proficiency :-) I think it'd be useful to know whether the model confirms the FVs or finds that the FVs are actually verified -- either way it'd be useful. But I had in mind what Alaexis mentioned: if there's one FV tag, there are probably more untagged FVs in the article. I'd guess the odds are higher than average (but I don't know that for sure).
For the dubious tag, that's like a "might" or "arguably" fails verification, so it'd be useful to have the tool analyze those for a person to review. A statement tagged dubious needs a verification check.
And yeah, I'd guess dubious and FV tagged content is more likely to have offline sources. I don't know the statistics, but I'd guess offline sources are rare and so it's unlikely to be significantly more?
Cool idea btw and well presented. Thanks to the teams for working on this! Levivich (talk) 21:12, 10 September 2026 (UTC)Reply
This sounds great and I'm glad the WMF is working on it. Regarding "what types of articles", I don't particularly see any reason to restrict articles included in the dataset, beyond "we don't want to scan the entire Wikipedia" - is that the reason? Or something else? In solidarity, asilvering (talk) 19:26, 10 September 2026 (UTC)Reply
If "We don't want to scan the entire Wikipedia" is the or a reason, then not scanning articles tagged as unreferenced (Category:All articles lacking sources) or lacking inline citations (Category:All articles lacking in-text citations) are obvious ones to not include as (assuming the tags are correct, which is a different issue) they cannot contain references that can be validated in this manner (~143k articles total). It's also not worth spending the resources attempting to scan references tagged as (permanently) dead, failed verification, or dubious. Thryduulf (talk) 19:55, 10 September 2026 (UTC)Reply
If "We don't want to scan the entire Wikipedia" is the or a reason, then not scanning articles tagged as unreferenced (Category:All articles lacking sources) or lacking inline citations (Category:All articles lacking in-text citations) are obvious ones to not include as (assuming the tags are correct, which is a different issue) they cannot contain references that can be validated in this manner (~143k articles total).
Great spot, @Thryduulf. Excluding articles in these categories for the reasons you named [i] sounds like a great idea to me unless, of course, there is a consequence here I'm not seeing.
It's also not worth spending the resources attempting to scan references tagged as (permanently) dead, failed verification, or dubious...
With regard to {{failed verification}} and {{dubious}}, can you please say a bit more here? Asked another way: what's prompting you to think it would not be worthwhile to scan articles with those two templates present?
Per what Levivich and I were discussing above, I'd been assuming, perhaps inaccurately, that scanning those articles could provide a helpful baseline to compare the model's proficiency against. Might I be missing something here?
Regarding the "tagged as (permanently) dead" bit specifically, I'm assuming the following. Please let me know what (if anything) I might've missed...
1. I assume you are referring to articles that include Template:Permanent dead link
2. If so, I assume the reason for excluding articles that contain ≥1 of these templates would be because these citations are unlikely to have an archived copy the suggestion could retrieve, making a check on them likely to fail
---
i. We can assume articles within All articles lacking source do not contain citations the model can evaluate claims against PPelberg (WMF) (talk) 21:01, 10 September 2026 (UTC)Reply
re failed verification and dubious, I hadn't thought about training the model. Rather I was just thinking that if a human has already tagged a reference as not supporting the associated text there isn't much benefit in an AI suggesting to a different human that it might not support the associated text (this would be a waste of time and resources). I agree that using them to train the model would be useful.
re permanent dead links, I was thinking on a per-citation not per-article basis - read the tag and skip the associated without spending any resources attempting to verify it in the source.
Your comment about archive templates has sparked another thought though, that this tool could highlight potential problems with archives. Firstly, if the url-status parameter is blank or live then it should attempt to verify using the live link or both links, if it is "dead", "deviated", "usurped" or "unift" then it should only attempt to verify against the archive link.
For every citation that has an archive link the following are possible:
  • Live link and archive are identical, both verify the text (no problem)
  • Live link and archive are identical, neither verify the text (problem, but not with the archive but worth flagging to human)
  • Live link and archive differ, but both verify the text (almost certainly not a problem)
  • Live link and archive differ, live link fails verification archive link passes verification (url-status parameter should be changed to deviated, usurped or unfit, but which probably requires human judgement)
  • Live link and archive differ, live link passes verification but archive link does not (flag this for human attention)
  • Live link and archive differ, both fail verification (include this information when flagging this for human attention)
  • Live link is dead, archive verifies text (url-status should be set to dead, possibly flag to something like user:InternetArchiveBot or some other automated task to make changes more widely).
  • Live link is dead, archive fails verification (change the url-status as above and flag to both bots and humans as other citations to the same source may also be dead and pass verification).
Thryduulf (talk) 22:29, 10 September 2026 (UTC)Reply
re failed verification and dubious...I agree that using them to train the model would be useful.
Wonderful. Thank you for walking out what you had been thinking!
...re permanent dead links, I was thinking on a per-citation not per-article basis - read the tag and skip the associated without spending any resources attempting to verify it in the source.
Ah, I see. In concept, what you're describing seems valuable to me. @Alaexis: do you think instructing the model to "skip over" claims that have the Template:Permanent dead link associated with them would require an update to the prompt?
@Thryduulf: in case you're curious, I'm asking about the prompt above because, for now, we're reluctant to make any changes to it. Rationale: a) the prompt has been performing pretty well, b) adjusting the prompt could affect the output in unexpected ways. For these reasons, we're hoping to keep the prompt as-is for this initial round of evaluation.
Your comment about archive templates has sparked another thought though, that this tool could highlight potential problems with archives. Firstly, if the url-status parameter is blank or live then it should attempt to verify using the live link or both links, if it is "dead", "deviated", "usurped" or "unift" then it should only attempt to verify against the archive link.
Oh, this is an interesting set of cases. @Alaexis two resulting questions for you and/or @Isaac (WMF):
1. Do we know how (if at all) the suggestion will behave in these various archive template cases?
2. More broadly, to what extent (if any) would it be accurate for me to think that both a) the LLM could accommodate nuanced instructions of this sort were we to deem them important and b) implementing these instructions would come in the form of an adjustment to the prompt? PPelberg (WMF) (talk) 00:22, 11 September 2026 (UTC)Reply
@Thryduulf If you're curious about how the code selects links to check, you can see the logic and additional notes in the userscript via the extractHttpUrl function. It was written to prefer internet archive links, and then the original link, and only then some of the other archive sites. Checking multiple URLs is possible without changing model prompts but it would still complicate the pipeline as it would functionally double the number of URLs to scrape and model requests to make so would slow things down a good bit. I'll leave that decision to Alaexis but the way I've been thinking about it: the core goal here is verifying the claim, which thusfar we've attempted to do that by fetching the "best" URL. As you raise, there are a variety of additional checks you could do at the same time with the goal of improving the citation (and therefore general Verifiability). I've also thought about inferring the language of the source URL to help fill in the lang parameters on citation templates. These additional checks/actions would complicate the core claim verification task though so I would almost prefer them to be separate at least from an interface perspective, but I am taking note of them.
I'll attempt to answer re: the permanent dead link suggestion as well. It's likely possible to add a check for them (this would also be with the heuristics for extracting links, not the LLM portion of the flow). My thinking: there aren't a ton of them and if they're a dead link then the flow will quickly+gracefully fail anyways (they would never show a suggestion to the end-user). Adding this sort of language-project-specific logic to the code can also complicate efforts to extend this tooling to other language editions in the future. But I'll leave it up to Alex whether it's worthwhile. Isaac (WMF) (talk) 17:22, 11 September 2026 (UTC)Reply
Unfortunately I don't know js so the link doesn't really aid my understanding (but that's not your problem). I don't have a strong feel for what is possible, my suggestions are all things that would are desirable if they are possible. If there are things it isn't going to do (for any reason) but might discover in the process of what it does do then documenting those things somewhere that other humans and/or bots can deal with would be a good thing. Thryduulf (talk) 18:04, 11 September 2026 (UTC)Reply
Yes and please keep the suggestions coming! I mainly wanted to communicate that even if some of these extension ideas don't end up incorporated into this experimental suggestion, they're not being discarded. Isaac (WMF) (talk) 19:48, 11 September 2026 (UTC)Reply
...beyond "we don't want to scan the entire Wikipedia" - is that the reason? Or something else?
Good question, @Asilvering. What you described is accurate, [i] In addition, it would be ideal if this initial dataset includes the types of articles that:
1. You all can imagine this suggestion being the most useful on
2. Includes cases that you think could be particularly complex/not straightforward so we can learn how/if the model fails here
---
i. The size of this initial batch of suggestions will be limited to ~1,000 articles PPelberg (WMF) (talk) 21:07, 10 September 2026 (UTC)Reply
This seems like a perfectly reasonable use of AI on Wikipedia, since it's not actually generating text. I'm looking forward to seeing how it performs. --Ahecht (TALK
PAGE
)
20:07, 10 September 2026 (UTC)Reply
Yeah, this sounds like a wonderful tool. It would be awesome if this could be fitted with a API front end so other tools could build on top of it. A while ago I built something (https://wikirefs.toolforge.org/show?page_title=Bronx+Grit+Chamber) which parses a page, pulls out all the individual claims, and matches them up with citations. But I did a half-assed job of it and got it to the point where it was good enough for my purposes. And I know if fails badly on some referencing styles. It would be great if all that low-level crud could be done once, properly, correctly, and then everybody else who wanted to build tools in that space could just take advantage of it instead of reinventing it from scratch. RoySmith (talk) 22:03, 10 September 2026 (UTC)Reply
A couple of other things that would be useful... If you don't find the claim in the cited source, look at the other sources cited in the article. It's not uncommon during editing an article for a properly cited statement to get moved but the citation doesn't move with it. Being able to recover the correct pairing would be valuable. Also, if you can't reach a source's URL, see if you can find some other place that has the same item. For example, I'll often cite a NY Times article via the NYT's own archives, which you may not be able to get to because it's behind a paywall, but you can find the same article in ProQuest or some other aggregator. RoySmith (talk) 22:52, 10 September 2026 (UTC)Reply
Combining source verification with finding the edit which added the source to the article often helps with this, could be a useful extentsion. It also helps when existing text in front of a source is modified without reference to the source. CMD (talk) 00:38, 11 September 2026 (UTC)Reply
I think that you can combine the script with "Who Wrote That" already, but you're that it would be useful to see it at a glance. Alaexis¿question? 11:06, 11 September 2026 (UTC)Reply
@RoySmith, it would be great to integrate with TWL and with the Internet Archive to get access to sources that aren't publicly available. Earwig's Copyvio detector has access to TWL so there is a precedent.
For citation that failed with the "Not Supported - Omission" verdict checking other sources in the article makes sense. It's not necessarily cheap - if you have 30 unsupported citations out of 300 total you'll need to run 30*300=9,000 checks, unless there is some kind of screening (the passages nearest where the claim sits or the sources added in the same edits). I tried building a standalone script (User:Alaexis/CNfirmed) to solve the broader problem of finding reliable sources. It turned out to be harder than I thought though. The biggest problem is *where* to look - generic web search is expensive and the alternatives are brittle. Alaexis¿question? 10:56, 11 September 2026 (UTC)Reply
It would be great if all that low-level crud could be done once, properly, correctly, and then everybody else who wanted to build tools in that space could just take advantage of it instead of reinventing it from scratch.
@RoySmith to be doubly sure I'm following, by "low-level crud" are you referring to things like splitting the article up into its constituent claims, identifying the source(s) associated with each, etc.?
A while ago I built something...which parses a page, pulls out all the individual claims, and matches them up with citations
Neat! If you happen to have them handy, we'd be eager to learn what referencing styles you noticed this tool struggling with.
...everybody else who wanted to build tools in that space could just take advantage of it instead of reinventing it from scratch.
I can imagine the creativity something like this could inspire. With this said, I think we're still a ways away from being able to determine how feasible something like this would be. Once you confirm the first question I posed above, I'm thinking I can create a phabricator ticket so that we can come back to it at a future point. PPelberg (WMF) (talk) 00:39, 11 September 2026 (UTC)Reply
Yeah, I envision some kind of network API where you give it a page title (revid, whatever) and it gives you back a list of claims and the associated citations in some structured form. Of course, this is really just one specific example of a general pattern. You really should be able to treat every page as an object on which you can perform operations and access those operations via a network API.
We spend too much time and effort reinventing wheels. For example, I don't know how many tools we've got which measure the readable prose size of an article. They all give different results because they all have slightly different definitions of what "readable prose" means. Neither are more right than any other, they're just different. And it's dumb that so many people have spent so much time writing essentially the same function in slightly different ways.
To go back to this particular example, there's already a bunch of different referencing styles in use. My code parses one of them (the one I tend to use) pretty well. I know it fails on others (to be honest, I don't remember which). Imagine a world where parsing claims and citations from articles was handled by one API that handled them all. Then all sorts of tools could take advantage of that. And more to the point when a new format comes along (say, sub-referencing, which is going to hit enwiki Real Soon Now), one bit of code will need to be adapted to handle that, and automatically all the various other tools will now be able to handle it too. RoySmith (talk) 01:12, 11 September 2026 (UTC)Reply
@RoySmith, that's so true. I had to build a lot of non-core stuff and would've been happy to reuse others' work instead. Discoverability of tools is a huge topic, Toolhub-Evolved is the latest solution I'm aware of.
Currently verification is not exposed as a standalone service but that wouldn't be too hard to do if you have a use case in mind. Fetching website data is already a standalone service on Toolforge (repo, ToolHub page). Alaexis¿question? 10:15, 11 September 2026 (UTC)Reply
Thank you for bringing this discussion here. As someone who's used Alaexis's tool on and off for a few months, this is the kind of collaboration I like to see: working with an established editor to support development or rollout of a tool that's already been field-tested by the community.To answer your questions: 1. Yes, definitely. 2. Dreamyshade has been building https://projo.toolforge.org/review to identify high-priority articles for WP:NPP reviewers – another opportunity for collaboration or sharing ideas? —ClaudineChionh (she/her · talk · email) 00:51, 11 September 2026 (UTC)Reply
Yes, thank you Claudine! The Projo unreviewed article priority-ranking tool (which is new and still in development, still janky) reflects my understanding of how to prioritize articles that need editor help in general, based on factors including CTOPs, BLPs, reference need, page views, orphan status, and cleanup tags such as AI-generated, COI, and POV (the page contains a table of all the factors). It's focused on supporting New Pages patrollers because that is the most urgent work that I know of right now, with the giant backlog and constant influx of COI/UPE and LLM-generated articles. I'm adding more factors based on data I can derive from categories and other elements in the replica database. I'd love to have access to APIs for predicted percentage of LLM-generated content and predicted numbers of verified/partially-verified/failed-verification citations!
I use Alaexis' Source Verifier a lot, especially for Articles for Creation review, New Pages patrol, and AI cleanup tasks. I find it very helpful, especially on articles that are tagged as likely AI-generated or that I suspect are AI-generated. It helps me rapidly check source-text integrity, which helps me figure out whether an article is mostly fine, salvageable, or trash. It's part of my standard toolkit for article quality evaluation, along with https://copyvios.toolforge.org/, https://wikipedia.gptzero.me/, and Cite Unseen - none of them are perfect, just tools, but I believe they're useful when used with competence and a grain of salt. I've been trying to gather lists of tools along these lines: Article workflows#Tools for specific workflow steps + User:Dreamyshade/Article workflows#Semi-automation. (That page has some of the ideas I'm trying to put into practice in the Projo tool.) Dreamyshade (talk) 01:48, 11 September 2026 (UTC)Reply
Wouldn't it not be possible to provide this System access to paywalled content over the Wikipedia Library? The Other Karma (talk) 09:44, 12 September 2026 (UTC)Reply
@The Other Karma, that's the single biggest thing that would extend coverage. Earwig's Copyvio Detector already reaches EBSCO through The Wikipedia Library (phab:T378077), but those credentials were granted for copyright enforcement only, so verification would need its own permission. Flagging this for Meta:User:Samwalton9-WMF as one more editor asking. Alaexis¿question? 06:21, 14 September 2026 (UTC)Reply
This looks like an actually productive use of AI on Wikipedia. Given the fact that it's not creating new information, it seems fine to me. I think a test run of this may be promising to introduce new editors. 🪐Kepler-1229b | talk | contribs🪐 17:23, 13 September 2026 (UTC)Reply
As for the questions, 1. Possibly but my main editing area lies outside of source verification, so I might not be able to test it as much. 2. Articles with "Failed verification" tags could be first priority for testing. 🪐Kepler-1229b | talk | contribs🪐 17:25, 13 September 2026 (UTC)Reply
For anyone interested in seeing a demo and talking about this in a voice call, we will be hosting a meeting in the Wikimedia Community Discord on 14 Sep 2026 from 17:00 - 18:00 UTC.
This Discord call is happening tomorrow (Monday). Here is a link for anyone interested in joining: https://discord.gg/wikipedia?event=1547782657969098752.
Note: we'll continue being responsive in this thread too. PPelberg (WMF) (talk) 04:51, 14 September 2026 (UTC)Reply
This seems really interesting, and a good usage of LLMs in Wikipedia. Thank you for the open communication too. qcne (talk) 12:02, 14 September 2026 (UTC)Reply
I'll be interested to see how this goes. LLMs are very bad at subtlety and being overconfident on specific answers, but very good when delivery an array of possible answers. Checking source verification could fit very well into the second type, displaying an array of possible issues that may require user intervention is something that could be very useful. -- LCU ActivelyDisinterested «@» °∆t° 20:41, 17 September 2026 (UTC)Reply
It's a much better fit than an LLM trying to understand NPOV, see my point about subtlety and overconfidence.
A common problem in verification is that an editor will add new unsourced content between a sentence and it's reference. So the content then looks like it's supported by a reference, but that's only partially true. Even if this could only surface those issues it would be a major win. -- LCU ActivelyDisinterested «@» °∆t° 20:47, 17 September 2026 (UTC)Reply
Sorry, I'm a bit confused by the question: "Would any of you all be interested in evaluating a batch of these experimental suggestions for en.wiki articles?" I signed up for what sounds exactly like this maybe six weeks or two months ago. The entry test was very buggy, but it seemed to be getting updated. Then nothing. Is this the same project or unrelated? Johnjbarton (talk) 22:17, 28 September 2026 (UTC)Reply

Thanks for the pro-active post on this. I have some reservations based on common things that can go wrong (as well as other things I would prefer be prioritized) but will wait to see the actual suggestions first. Gnomingstuff (talk) 06:09, 14 September 2026 (UTC)Reply

I have some reservations based on common things that can go wrong (as well as other things I would prefer be prioritized) but will wait to see the actual suggestions first.
Understood and sounds great. We'll be in touch in this thread once we have the initial set of suggestions generated for y'all to review.
Before that, we'll be in touch with what criteria we're proposing to use to select the initial 1,000 articles we'll generate suggestions for to make sure we think it will produce a useful sample.
Thanks for the pro-active post on this.
You bet; thank you for demonstrating to us the need and value in doing so, @Gnomingstuff. PPelberg (WMF) (talk) 22:44, 14 September 2026 (UTC)Reply

@RoySmith:, you suggested that it would be useful to have the "verify claim against source" exposed as an API endpoint, so I've done it, details here. I've checked it myself, so feel free to tinker with it too. To avoid any doubt, the edit suggestions pilot doesn't use this endpoint. Alaexis¿question? 19:00, 29 September 2026 (UTC)Reply

Cool, thanks, I'll take a look. -- RoySmith (talk) 19:31, 29 September 2026 (UTC)Reply

Hi y'all – below is the criteria we are planning to use to select the initial ~1,000 articles we will generate source verification suggestions for.

If you see criteria that are included or missing from the below that you think is critical to reviewing the reliability of this suggestion, please comment as much.

Article selection criteria

We are proposing to generate an initial batch of source verification suggestions for ~1,000 articles using the following selection criteria. The criteria you see are meant to reflect what we heard from you all: this dataset ought to be representative of where you see these suggestions being most useful and potentially, risky.

Selection criteriaDescriptionMotivation
Evergreen A set of 100 articles User:Alaexis has predicted are likely to be edited within the next 30 days. This is a set of popular and frequently edited articles where the impact of inaccurate claims is higher and where the suggestions are more likely to be seen and acted upon.
Category:BLP Biographies of living people Articles susceptible to mis-attributed/incorrect information and where evaluating all of them might be unrealistic. Automation to help facilitate this process could make this task more feasible for more people on an ongoing basis.
Category:CTOP Contentious topics
Category:History Historical articles Articles that are prone to inconsistencies and thus more likely to contain claims that need verification[1]
{{dubious}} Articles where at least one {{dubious}} template is present Strong signal that unsubstantiated claims may be present; offers a helpful reference to benchmark the model against[2]
{{failed verification}} Articles where at least one {{failed verification}} template is present
{{AI-generated}} Articles volunteers have concluded to be AI-generated Articles more likely to contain hallucinated references (read: claims that are not substantiated by the sources that accompany them)
Good article nominations Articles nominated for "Good" article status Thorough review of all claims within article nominations is necessary and toilsome.
NPP Backlog Articles created in 2026 that are not reviewed yet There is a long backlog of new articles for review and being able to increase the speed with which volunteers can evaluate verifiability could help decrease it.

Next steps In terms of process, here is how we see things going from here:

  • Step #1: Use the criteria above to select ~1,000 articles we will create an initial batch of source verification suggestions for. ← we are here
  • Step #2: Conduct an internal review of these suggestions to ensure they are of sufficient quality to warrant y'alls (volunteer) attention
  • Step #3: Share the results of the internal review and invite volunteers (you all) to conduct a second review of the internally-vetted dataset.
  • Step #4: Make this initial batch of suggestions available for volunteer review in two ways:
    1. In bulk via a spreadsheet
    2. In context via an experimental suggestion within Suggestion Mode
  • Step #5: Analyze volunteer evaluations
  • Step #6: Share findings and discuss next steps on-wiki

How does this all sound?

PPelberg (WMF) (talk) 16:17, 24 September 2026 (UTC)Reply

This sounds like a solid plan. I must emphasize the importance of being honest with yourselves in the internal review – if you're having to, say, regenerate half the suggestions because of obvious errors, that might mean more work is required on the underlying tool before volunteers test it. Please tell us exactly what you find there, with numbers. I also wonder, if each of these nine categories gets ~1/9th of the 1000 articles, do we have enough statistical power to do whatever analysis you are hoping to do? I'm not a statistician, and 100+ articles per category is still a lot, but that's something you want to check beforehand. Finally, would this check the entire article on the articles where it's run? For instance, if a page has an inline failed verification tag, I would be disappointed if the program was set to generate, say, 3 suggestions per article max and then didn't check the sentence tagged as failed verification. For the NPP category, I suggest choosing only articles that have sat in the NPP queue for a week or two to make them less likely to be instantly reviewed and fixed up by someone patrolling the front of the queue. Toadspike [Talk] 16:50, 24 September 2026 (UTC)Reply
This sounds like a solid plan. I must emphasize the importance of being honest with yourselves in the internal review – if you're having to, say, regenerate half the suggestions because of obvious errors, that might mean more work is required on the underlying tool before volunteers test it. Please tell us exactly what you find there, with numbers.
Well put and agreed. To be explicit, following the internal review, you can expect us to share:
1) How many suggestions we reviewed
2) What a review of each of these suggestions entailed
3) What the quantitative and qualitative outcomes of these reviews were
4) What we think the next step(s) are in the light of the above.
I also wonder, if each of these nine categories gets ~1/9th of the 1000 articles, do we have enough statistical power to do whatever analysis you are hoping to do? I'm not a statistician, and 100+ articles per category is still a lot, but that's something you want to check beforehand.
First, to the question you asked: we will be generating suggestions for all claims within the articles in the dataset, not just a subset. In line with the above, we expect this approach to produce a dataset of, at least, 10,000 suggestions. I think this will provide the scale we need to draw meaningful conclusions. Tho, I defer to @Isaac (WMF) to say definitively here.
For the NPP category, I suggest choosing only articles that have sat in the NPP queue for a week or two to make them less likely to be instantly reviewed and fixed up by someone patrolling the front of the queue.
Good spot. I think incorporating this constraint shouldn't be too difficult for this run. Although, if that proves to be the case, we'll let you know as much.
Thank you for thinking this through, @Toadspike. PPelberg (WMF) (talk) 21:24, 24 September 2026 (UTC)Reply
I think incorporating this constraint shouldn't be too difficult for this run. Although, if that proves to be the case, we'll let you know as much.
Ok! It turns out that implementing the above at this point could have unintended side effects. So, for this first run, we'll not apply any date-based criteria to the articles we analyze from the NPP queue. "Worst" case, we'll see fewer suggestions there because they will be, as you hypothesized, of better quality. PPelberg (WMF) (talk) 21:31, 24 September 2026 (UTC)Reply
Thanks, sounds good. I look forward to being able to get started on this. Toadspike [Talk] 21:45, 24 September 2026 (UTC)Reply
Wonderful. I expect us to complete the internal review and have results to share the week October 12th. PPelberg (WMF) (talk) 21:59, 27 September 2026 (UTC)Reply

Hi everyone, We're planning a small experiment this year aimed at improving the experience of people who donate to Wikipedia. We wanted to share our proposal and get your feedback. You can find our more over at the Fundraising Hub. Best, JBrungs (WMF) (talk) 13:00, 14 September 2026 (UTC)Reply

Will this be opt-in or opt-out?
Will someone be able to get all of the "...aimed at improving the experience of people who donate to Wikipedia" benefits without donating? If not, exactly what does someone who pays get that is not available to someone who does not pay?
May we please have an explicit assurance that nobody on Wikipedia -- users, admins, stewards, arbcom members -- will be able in any way to know whether someone is or is not a donor? --Guy Macon (talk) 18:33, 14 September 2026 (UTC)Reply
@Guy Macon: this is just a pointer to the linked full announcement and engagement at Wikipedia:Fundraising/Fundraising Hub#Testing donor account creation where at least most of your questions were answered in the first post. Thryduulf (talk) 19:54, 14 September 2026 (UTC)Reply
For my first question I see that the "small experiment" that the WMF is planning is opt-in. Are you saying that whatever is true about the plans for the small experiment will be true about the final rollout? Rather than me making that assumption and later being told that I should not have assumed, I would very much prefer a clear statement that the the final rollout will be opt-in, opt-out, or that this hasn't been decided yet.
I see nothing on the linked page that addresses my second question. I know that there has to be at least one difference: you don't get a receipt that you can show the IRS if you don't donate.
The wording at ] appears to answer my third question. The answer is "The information we collect will be private to your account and never shared externally or to other users." That seems clear enough. I will be interested in seeing how the WMF reconciles that with "Unlocks features across Wikipedia's websites and apps" -- I assume that this means adding only features that are invisible to other users, such as turning off donation banners. --Guy Macon (talk) 21:18, 14 September 2026 (UTC)Reply
You won't get answers to your questions here because this is a pointer not a discussion. You need to ask them at the linked page where those who know the answers are looking for the questions. To be clear, I have no connection with the fundraising team or this test at all. Thryduulf (talk) 21:21, 14 September 2026 (UTC)Reply
In case anyone is following this, we answered the questions across at the Fundraising Hub. Best, JBrungs (WMF) (talk) 09:17, 17 September 2026 (UTC)Reply

Here is a quick overview of highlights from the Wikimedia Foundation since our last issue on August 28. Please help translate.


Baseline Neutral Point of View policy is now being introduced.

Highlights


Annual Goals Progress on Engage
See also: Growth · Product Safety and Integrity · Tech News · Language and Internationalization · The Wikipedia Library · list of movement events · Wikifunctions & Abstract Wikipedia

TextMatch: A paste from an LLM triggering a Suggestion.
  • Introducing TextMatch: Communities are shaping Suggestion Mode with TextMatch to create custom suggestions that are tailored to their own policies and writing conventions. Share your feedback.
  • Visual editor becoming default: Design work on the user experience for visual editing for newcomers is underway. This follows visual editor becoming default on English Wikipedia following a community discussion.
  • Community Wishlist 2027 consultation: The consultation on the proposed new Wishlist process is closing this week. This year's cycle plans for wish submissions in late October/early November, triage completed by late November, and voting in early to mid January. Share feedback in any language on the talk page.
  • Improved Add a Link feature: Add a Link has been upgraded for English Wikipedia users with improved detection of links that have different capitalization. This will help reduce ambiguous suggestions caused by differences in capitalization in article titles, including articles about cultural goods.
  • Wikipedia Android app: Latest release includes updates to the Saved feature, bringing the app’s saving experience closer to iOS and Web. The update redesigns the Saved tab with an “All articles” view, removes the default “Saved” reading list, renames reading lists to “Collections,” and modernizes the article-saving experience.
  • Domain change for thumbnails: The domain of URLs for thumbnails is changing from upload.wikimedia.org to thumb.wikimedia.org. The old URLs will continue to work but MediaWiki will advertise the new domain instead. Original files, videos, and transcodes will still be served from upload.wikimedia.org.
  • Reading lists rollout continues on desktop: Following the initial deployment to Bengali, Chinese, Czech, and Vietnamese Wikipedias, reading lists went live on Arabic, French, and Indonesian Wikipedias and will go live on English Wikipedia on Sep 14, with all remaining wikis following on Sep 28.
  • New features of Pageviews Analysis: Several new features have been added to the Pageviews Analysis tool which allows users to compare pageviews across multiple pages. They include WikiNav which provides insights into how readers of Wikipedia explore the content, editing stats in Siteviews, the ability to lookup articles belonging to a WikiProject, and support for dark mode.
  • Article guidance: As part of an A/B experiment, Article guidance is now available in Turkish and Simple English.
  • Latest experiments: A current experiment for logged-out mobile readers on Bengali, Czech, English, Farsi, and Polish Wikipedias is testing whether simplifying the Minerva navigation bar improves reader retention. See all live, upcoming, and completed experiments in Product & Technology.
  • Wikifunctions status updates: Thanks to all contributors, Wikifunctions has now reached 5,000 functions.
  • Latest highlights from Tech News: Weeks 36 and 37 of Tech News include the Article guidance feature is now enabled by default for editors with less than 100 edits on Simple English and Turkish Wikipedia.


Annual Goals Progress on Enable
See also: Research newsletter · WikiLearn News · other newsletters on MediaWiki.org

  • Affiliates and grant funding: Have your say on proposals establishing (1) a new model and set of expectations for movement affiliates; (2) new, connected requirements to determine eligibility for receiving Community Fund grants from the Wikimedia Foundation; and (3) updated criteria for affiliate recognition processes.
  • Community Fund reflections: Read the reflections on the fiscal year 2025–2026 of the Rapid Fund and General Support Fund, covering $17.73 million across 427 grants in 90 countries, with 56.2% distributed outside North America and Northern & Western Europe.


Annual Goals Progress on Reach
See also: Wikimedia Apps · Readers

  • World of Wikipedia launches: The World of Wikipedia game has been released. It is an online collectible card game where players sort fact from fiction across Wikipedia articles to build their card library. The game is the latest experiment in reaching younger audiences through familiar formats, following Wikispeedia on Roblox and Which Came First? on the Wikipedia app.


Other updates
See also: Wikimedia Foundation Board noticeboard · Affiliations Committee Newsletter

  • Celebrate Women 2026: Read the impact of Celebrate Women 2026 as the campaign concluded with 112 events, 238,222 unreverted content edits and 2,499,773,577 bytes added.
  • US staff union vote: On Sep 3, eligible Wikimedia Foundation staff in the United States voted in favor of union representation in a National Labor Relations Board (NLRB) election.


Other Movement-curated newsletters & news
Diff blog · Goings-on · Planet Wikimedia · Signpost (en) · Kurier (de) · Actualités du Wiktionnaire (fr) · Regards sur l'actualité de la Wikimedia (fr) · Wikimag (fr) · Education · GLAM · Milestones · Wikidata · Central and Eastern Europe · other newsletters

MediaWiki message delivery 01:53, 16 September 2026 (UTC)Reply

Hi everyone!

I am Kaarel, working with the Wikimedia Foundation, facilitating conversations regarding the future of affiliate landscape. We would love to get your thoughts on the current discussions on recognition and funding of movement organizations.

We’ve heard many times (for example here and here) about problems related to grant funded projects and that's why it’s incredibly important that your voices, ideas, and expectations are heard.

We just published a summary report from our July and August conversations. It highlights what people agree on, where they disagree, and what still needs clearing up. We hope this report gives you a good overview of the chat so far and inspires you to jump in! Your perspective as a project contributor is vital to shaping the future of this model.

Please head over to Meta (general discussion, affiliate model discussion, funding discussion), to share your thoughts. If you have any questions or just want to chat through it, feel free to reach out to me. I am keen to set that up, listen and talk through.

We will also be hosting two public community calls next week and the week after: 1) Thursday, September 24, 16:30-17:30 UTC on affiliate overlaps 1) Tuesday, September 29, 16:00-17:00 UTC on affiliate growth paths - LINK TO THE CALL and 2) Thursday, September 24, 16:30-17:30 UTC rescheduled to Friday, October 2nd, 17:00-18:00 UTC on affiliate overlaps - I will share the call links here closer to the calls.

Thank you for your very kind attention and have a great rest of your week! --KVaidla (WMF) (talk) 08:40, 16 September 2026 (UTC)Reply

This is a real opportunity to bring editors and affiliates even closer together. But it does require some changes on the affiliate front and that's not easy. I know many editors are like "affiliates have nothing to do with me, so why should I care about this?" But I think "affiliates have nothing to do with me" doesn't have to be the way it is and more importantly shouldn't be the way it is. I hope other editors will join me in expressing a desire to see affiliates refocus their work in ways that tangibly improve projects, rather than continuing to be their own separate thing. Best, Barkeep49 (talk) 14:12, 16 September 2026 (UTC)Reply

I couldn't help but notice this: "The present text has been revised based on the valuable feedback from various community stakeholder groups, including the Affiliations Committee, the Global Resources Distribution Committee, and directors of existing movement organizations".
Notably missing are donors, people who read Wikipedia but do not edit, and the community of Wikipedia editors. I would welcome evidence that there was even a small effort made to get feedback from these missing stakeholders.
From Wikipedia Foundation exec: Yes, we've been wasting your money in The Register and The WMF Executive Director’s Reflections on the FDC Process on Meta (which I recommend reading in full for context):
"I have significant concerns about how our movement entities are developing... I believe that currently, too large a proportion of the movement's money is being spent by the chapters...
I am not sure that the additional value created by movement entities such as chapters justifies the financial cost...
I do also believe that people who are involved in chapter organizations (and other Wikimedia organizations) have a particular worldview that is in some ways different from that of Wikimedians who choose not to become involved with incorporated Wikimedia organizations, and I think a healthy funds dissemination process would benefit from multiple perspectives...
[The] process, dominated by fund-seekers, does not as currently constructed offer sufficient protection against log-rolling, self-dealing, and other corrupt practices...
The community members who are paying the closest attention to the process are applicants, and that community involvement in scrutinizing proposals is otherwise low...
There is currently not much evidence suggesting this spending is significantly helping us to achieve the Wikimedia mission."
I see little evidence that these fundamental problems have been addressed in the 13 years since they were posted.
And before someone says it, Wikipedia:Village pump (WMF) is the place where the WMF should look to get feedback from the community of Wikipedia English Wikipedia editors. not a talk page on meta where comments get few or no responses. --Guy Macon (talk) 15:08, 16 September 2026 (UTC)Reply
So do you think the changes the WMF has proposed will help donor money be spent better or do you think there should be different changes to affiliates (it's clear you don't think the status quo is good)? I don't think there's anything wrong with us posting here, I've posted substantive comments myself after all, but enwiki editors are not the same as "wikipedia editors" and so having a central place, like meta, that should be open to all seems like an obvious answer for me. Best, Barkeep49 (talk) 15:15, 16 September 2026 (UTC)Reply
It should also be noted that Wikipedias are not the only WMF projects and editors of those projects should also be able to have their say without requiring WMF staff to read dozens of independent discussions. Thryduulf (talk) 15:52, 16 September 2026 (UTC)Reply

I think that what the WMF is doing is well thought out and will certainly help to insure that donor money is spent better. On the other hand, I also think that those on the receiving end are very strongly motivated to maximize how much cash flows their way, and will do whatever it takes to accomplish that.

The problem is that I don't see anyone surveying a sample of donors and asking them whether this is what they had in mind when they donated. I think that if you asked them the overwhelming majority would say that the money should go to things like hosting.

The WMF knows this. That's why the latest banner ads say "We hope that [Wikipedia] has given you at least $2.75 of knowledge. If so, please join the 2% of readers who give to keep this resource available for all".

Note that they did not say "...who give to match racial justice leaders with machine learning research engineers to develop data-based machine learning applications."

I am reminded of the many smart people who asked "are we spending money of the right things in Vietnam? Should we fund better rifles or better boots? Should we buy more tanks or more jeeps?" all without ever asking "should we be fighting a war in Vietnam at all?" I think you can guess what the answers from the "stakeholders" who supplied the boots and the jeeps were. --Guy Macon (talk) 17:44, 16 September 2026 (UTC)Reply

The WMF does seem to spend an awful lot of money on causes which are noble but totally unrelated to what donors thought they were funding. No one is here to oppose worthy aims such as racial justice but they have their own dedicated charities. Indeed, some of our donors may also be sponsoring those organisations. Either way, they're entitled to have the money they allocated to Wikipedia[a] spent here and not diverted to unexpected social campaigns.
  1. ↑ Taken by the WMF but, as so much of the money comes from the Donate link in the sidebar titled Wikipedia, it's reasonable to assume what it's intended for.
Certes (talk) 19:07, 16 September 2026 (UTC)Reply
re your footnote, there are also links to donate from the sidebar of all the other projects I looked at except Commons and MediaWiki (usually as "donate" but sometimes as "donations"). Thryduulf (talk) 19:19, 16 September 2026 (UTC)Reply
Guy, they didn't say that because they aren't doing that. The Knowledge Equity Fund was a three-year program and is now over. In solidarity, asilvering (talk) 20:33, 16 September 2026 (UTC)Reply
Pick whatever they are spending money on that is unrelated to "keeping this resource available for all" this week and insert that. The specific spending on things totally unrelated to what donors thought they were funding keeps changing and is often only discovered after they have moved on to the next new thing, but we all know that the pattern is ongoing. --Guy Macon (talk) 20:49, 16 September 2026 (UTC)Reply
@Asilvering, I'm not sure this is the case. Consider the projects listed here. Yes, they have "articles created" and "editors retained" metrics and all that, but one gets the feeling that these goals and their monitoring is not the central concern in these programs.Alaexis¿question? 06:05, 28 September 2026 (UTC)Reply
Were the deficiencies and risks identified on this document ever addressed? --Guy Macon (talk) 19:08, 16 September 2026 (UTC)Reply
To my knowledge, the recommendations they list are our current practice, with the exception of WMDE. In solidarity, asilvering (talk) 20:06, 16 September 2026 (UTC)Reply
For the context of casual readers, the document is dated July 2011. Thryduulf (talk) 20:08, 16 September 2026 (UTC)Reply

Scholarship applications are now being accepted for Wikimania 2027, which will take place August 18 to 21 in Santiago, Chile. The final deadline to submit your application is October 31, 2026 (end of day Anywhere on Earth).

Scholarships cover travel and accommodations, registration, and limited travel insurance fees. Wiki project contributors and Wikimedians working to connect our movement with the wider open knowledge and culture ecosystem are invited to apply.

Tell us about your Wikimedia journey, share your vision for “The Internet we want” (the theme for Wikimania 2027), and be sure to include links to your work. All in-person attendees, including scholars, will be subject to trust and safety checks.

More detailed information on the application process can be found on the Scholarships section of the Wikimania wiki, and in this Diff post.

Orientation sessions for potential applicants will be offered starting the week of October 5, 2026. Successful applicants will be notified starting in January 2027.

Best of luck to all! Eureka-WMF (talk) 15:04, 22 September 2026 (UTC)Reply


Hello, should administrators on different wikimedia projects like Wikipedia be paid? I think yes. Botaki (talk) 12:28, 25 September 2026 (UTC)Reply

Speaking as an administrator who could certainly do with more money, no. Thryduulf (talk) 13:12, 25 September 2026 (UTC)Reply
There is an old saying: "He who pays the piper calls the tune." Right now the administrator permissions are granted and managed by the individual communities, without any input from the WMF with extremely rare circumstances. If the WMF starts paying us, they will have to assume control of that, as is necessary in any employer/employee relationship. I don't really think the community wants to give up control of who becomes an administrator, and what rules admins will enforce, and what tools they have at their disposal. So no, I don't think the WMF should pay administrators. Speaking personally, I don't want to be a WMF employee, thank you very much. Risker (talk) 13:34, 25 September 2026 (UTC)Reply
Go ahead and pay them. Nobody is stopping you. (Exception: paying an admin to do something will get you both booted from Wikipedia, but setting up a fund that is shared equally by all admins willing to accept payment is probably fine -- but check first instead of taking my advice.)
Oh, wait. I missed the "by wikimedia foundation" part. Sorry. Silly mistake. The WMF has no control over what administrators or editors do (again, with exceptions. If the admins were to, say, decide to allow child pornography or copyright infringement the WMF has a legal responsibility to stop them. See WP:OFFICE) I don't see how being independent is compatible with admins being paid employees of the WMF.
Also, we already have far too many individuals who appear to be far more interested in keeping the grant money flowing than in anything that actually helps any of the wikis. See Where does your Wikipedia donation go? Outgoing chief warns of potential corruption. --Guy Macon (talk) 13:40, 25 September 2026 (UTC)Reply
Ignoring all the other obvious concerns, were the WMF to start paying admins, they (the WMF) would be legally accountable for their actions. I doubt very much they'd wish to take on that responsibility, as they'd likely find themselves involved in an endless stream of lawsuits. AndyTheGrump (talk) 13:59, 25 September 2026 (UTC)Reply
I think no. Sohom (talk) 14:12, 25 September 2026 (UTC)Reply
I think our pay should be at least doubled. signed, Rosguill talk 14:51, 25 September 2026 (UTC)Reply
If you keep making these unreasonable demands, we're going to cut your pay in half! Levivich (talk) 19:09, 25 September 2026 (UTC)Reply
How much do I get if I make reasonable demands? Anyway, I'm not too worried about the salary, I'm mostly in it for the stock options. RoySmith (talk) 22:28, 25 September 2026 (UTC)Reply
(Not an admin) I would second the above comments, although not being an admin (thankfully) my opinion might not be worth much. A two-tier system would potentially create division in some editors minds. At the moment, admins are editors with extra buttons and (usually) a good understanding of how Wikipedia works.
The cabal shouters don't need any further encouragement! Knitsey (talk) 22:35, 25 September 2026 (UTC)Reply
There Is No Cabal (TINC). We discussed this at the last Cabal meeting, and everyone agreed that There Is No Cabal. An announcement was made in Cabalist: The Official Newsletter of The Cabal making it clear that There Is No Cabal. The words "There Is No Cabal" are in ten-foot letters on the side of the 42-story International Cabal Headquarters, and an announcement that There Is No Cabal is shown at the start of every program on The Cabal Network. If that doesn't convince people that There Is No Cabal, I don't know what will. --Guy Macon (talk) 22:47, 25 September 2026 (UTC)Reply
If only you paid me my cabalbucks. Knitsey (talk) 22:51, 25 September 2026 (UTC)Reply
This would create all kinds of unsolvable problems. Admins would in many countries be considered employed by the WMF, but would be elected by the volunteers and the volunteers would be able to fire them via the recall process. The legal ramifications of that would cause complete chaos. -- LCU ActivelyDisinterested «@» °∆t° 00:07, 26 September 2026 (UTC)Reply
Paid by the WMF? No. I do recall at least one editor who had a Patron, though. And I have wondered at times whether folks should experiment with a "tip jar" sort of thing, giving readers an option to thank volunteers other than just donating to the WMF. Personally, I'm skeptical that it's possible to introduce money in any of these ways without considerable unintended consequences. — Rhododendrites talk \\ 02:51, 26 September 2026 (UTC)Reply
Patreon, I suppose? Alaexis¿question? 17:30, 27 September 2026 (UTC)Reply
Not to mention, how would you determine pay? It would have to be per admin action, otherwise very busy admins would be paid the same as those that never use their tools at all. And if you're paying per admin action, that sounds like a very good way to get people to (a) rush their actions and get them wrong, and (b) get burnt out. Black Kite (talk) 18:30, 27 September 2026 (UTC)Reply
Also not all admin actions are the same. Protecting a page due to obvious vandalism requires much less effort than closing a contentious AfD, yet both would (presumably) count as a single action. It would also need to be decided whether CU and OS actions count as admin actions and if so at what level (one CU check typically results in more log entries than dealing with one OS request). This could also lead to competition between administrators to be the one to record the logged action rather than collaborating with their colleagues to get the right outcome. Thryduulf (talk) 18:41, 27 September 2026 (UTC)Reply
Another thing that probably is worth keeping in mind is that taking a decision to not do a admin action is sometimes as valuable as doing one. Declining a speedy deletion, declining a unblock request or a WP:PERM request is equally as valuable as granting them. Sohom (talk) 19:21, 27 September 2026 (UTC)Reply
If I block somebody and they're later unblocked, is my commission subject to clawback? RoySmith (talk) 20:23, 27 September 2026 (UTC)Reply
And what about panel closes - is it the standard fee for each member or do they have to share one? Thryduulf (talk) 20:41, 27 September 2026 (UTC)Reply
This is a terrible idea and would remove the independence of admins by introducing a massive conflict of interest. It's also just unworkable in general: there are way too many admins throughout the numerous Wikipedia languages and other projects, and there are no clear rules for how many there should be, creating an incentive to increase their number to get more money (or worse, forcing the WMF to determine how the projects choose admins) Ita140188 (talk) 07:54, 28 September 2026 (UTC)Reply

The discussion above is closed. Please do not modify it. Subsequent comments should be made on the appropriate discussion page. No further edits should be made to this discussion.

Please review:

There is a lot to unpack there, so please post anything you find that seems worth discussing. --Guy Macon (talk) 18:38, 28 September 2026 (UTC)Reply

Slow Editing Towards Equity

[edit]

I will start with one that caught my eye:

Was this $42K well spent? Who benefited? What was accomplished?

Was this the sort of thing that the donation banners describe your contributions as funding?

Why are some of the links at dead or useless? Are we not capable of posting these publications on WMF servers? --Guy Macon (talk) 18:38, 28 September 2026 (UTC)Reply

Takeaways from the linked paper at [14]:
Theory Section
- The power to shape consensus on Wikipedia is not evenly distributed. Even aside from groups like ArbCom, things like technical knowledge, access to bots, and understanding of the site's jargon gives the small slice of the community with mastery over them ("experienced users") an overwhelming presence in policy discussions.
- The rules pages act as entrenched technical authority that can be used by experienced users against outsiders, not merely as consensus records.
- As proof of the above, most policy pages are very old. If a page survives its first year of official status, it gains enough momentum to be effectively permanent. 11 of the 15 sampled policy pages were fully stabilized by 2011.
- The rules pages are very stable, and that's good. However, they are resistant to change in a way that discourages any new user that isn't already aligned with Wikipedia's way of thinking. Such users don't become experienced users with the ability to shape discussion, and so can't reform the very rules preventing their participation. This is a barrier to gaining new editors from outside the spheres Wikipedia originated in (American male tech nerds).
- Wikipedia's policies, both content and social, can be used as a cudgel against users who threaten local consensus and/or individual editors' worldviews. Examples cited.
- All of the above to say, Wikipedia's current policies fail to reflect actual community consensus due to how they empower users who already agree with them.
There's more stuff about the actual form said empowerment takes, lot of sociology jargon and such. It records the various levels of authority a rules page can have, including unofficial levels like "essay endorsed by popular community members."
Data Section
- The number of individual users involved in creating each policy is very low. It's mostly the same handful of people talking to each other. There were more "critical discursive moments" identified in the sample than unique editors.
- enwiki prefers discussion among a few people with strong opinons and arguments over the flatter voting system of eswiki. The other three languages didn't have much of either, and just kind of trusted whoever was writing the rules.
- Recommendation made that the processes used to define and change policy be adjusted to better reflect actual editor consensus rather than consensus among the technocratically empowered few. No hard suggestion included.
Overall I think it's a useful study, if not for Wikipedia itself then certainly for the history and sociology of internet communities. I'm not mad about it having been funded. ~2026-52337-89 (talk) 23:56, 28 September 2026 (UTC)Reply
The promise (when they were asking for $42K) was "identifying gaps in policy that are necessary for the needs of all English Wikipedians." Please name a gap in policy that was identified, and how we can fix that gap.
And again I ask, Was this the sort of thing that the donation banners describe your contributions as funding? See The Huge Fight Behind Those Pop-Up Fundraising Banners on Wikipedia in Slate --Guy Macon (talk) 02:14, 29 September 2026 (UTC)Reply
Thinking on how to resolve this, perhaps that's the solution - we establish a consensus that donations banners run on enwiki must accurately detail where the money is going. If something is beyond the scope of what donation banners said the WMF is spending money on, then either the WMF needs to cut the program, or not run banners on enwiki. BilledMammal (talk) 02:19, 29 September 2026 (UTC)Reply
@Guy Macon, I don't think expecting research to "succeed" in doing something 100% of the time is a reasonable expectation to have. Research can and does "fail" even with the best effort. Doing research on policies on Wikipedia is within Wikimedia Foundation's overarching programmatic goals to better the community to which donators donate to. There was a certain amount of funding allocated to the researchers, they tried to gain some useful information, they failed in the overall goal but found some interesting findings about internet communities which got published in a (what I assume) is a well respected journal. Sohom (talk) 02:55, 29 September 2026 (UTC)Reply
No need to ping me. I have a watchlist and know how to use it.
Useful? Sure. $42,356 USD of the donor's contributions useful? Not so much. A lot of those donors are living in poverty and sacrifice to give because they were told that this was needed to keep Wikipedia free.
There are many, many researchers who have published papers on Wikipedia in well respected journals. I question why only a handful of insiders get funding for doing that from the WMF. --Guy Macon (talk) 03:06, 29 September 2026 (UTC)Reply
seful? Sure. $42,356 USD of the donor's contributions useful?, that's how research works??? You provide a chunk of funding so that a group of researchers can try different methodologies and techniques and build prototypes to see if any work. If anything, 42K is fairly cheap in the research world for the length of time the project took (2022-2025) actually since if you think about it, a single grad student's yearly tuition and stipend in CS (being paid barely enough to sustain themselves, while also being one of the most funding rich areas) amounts to roughly that number.
I question why only a handful of insiders get funding for doing that from the WMF.. Guy Macon, I would suggest looking at the page for the dissemination of research funds which provides full transparency into who is allowed to put in proposals, how much money WMF is willing to give for research projects, the proposals submitted and rejected, the folks who decide who to fund (who are well respected researchers within the Wikimedia Research community). I don't think characterizing this as "insiders get funding" is a correct characterization. Sohom (talk) 03:22, 29 September 2026 (UTC)Reply
As I have said many times, I think that everyone at the WMF -- from engineers to the people who decide who and what to to fund, are doing a great job at what they are trying to do. I have zero doubt that you are really looking hard at who you decide to pay to do research, and by "insiders" I don't mean "your buddies" or anything implying corruption or malfeasance. I believe you are like the fine people back in the 1960s who were managing the war in Vietnam. They did seriously great work deciding whether to fund more jeeps or more tanks, whether the jungle boots could be made better, etc. Thousands of good decisions were being made by dedicated individuals doing a great job, but never once considering or even being allowed to consider whether we should be fighting a war in Vietnam at all.
Now consider the entire world population of people who have published research on Wikipedia. All of them. They all get paid by someone, usually their university. What percentage get WMF grants? Do an honest estimate. One in a thousand? What differentiates the two groups? One group knows how to craft a proposal that will get them funding in the usual ways academics get funding. The other has learned how to work the system to get WMF funding. How many of each group get free travel and lodging to go to Paris and spend time with Wikimedia staff and Wikipedia editors? How many in each group know the names of multiple WMF employees? How can you call the tiny minority that get grants anything other than insiders? --Guy Macon (talk) 06:47, 29 September 2026 (UTC)Reply
And again I ask, was this the sort of thing that the donation banners describe your contributions as funding? --Guy Macon (talk) 06:47, 29 September 2026 (UTC)Reply
There's lots of things the WMF is doing badly, but I don't think this particular funding decision is one of them. The article seems to be of good quality to me, and seems in line with the research questions in the application. TietoTeekkari (talk) 07:56, 29 September 2026 (UTC)Reply
And I fully agree, just as I agree that the military made some great decisions about funding jungle boots during the Vietnam war. They were great boots. I still have a pair and they look like new. And, just as I think that same military made those great boot decisions without even considering whether we should be there we should be fighting a war is Vietnam, I think the WMF did a fine job of picking which researcher to fund while never even considering whether we should be funding academic papers at all. And again I ask, was this the sort of thing that the donation banners describe your contributions as funding? Is there some special feature in Wikipedia's software that makes the preceding words invisible when I write them? --Guy Macon (talk) 13:50, 29 September 2026 (UTC)Reply
Not Tieto, but the last time I donated to Wikipedia I recall the banner saying something about "supporting and defending open knowledge projects." I'm probably not getting the phrasing exactly right. This research seems very relevant to Wikipedia and how our internal community and governance systems operate, along with how community-based open knowledge projects operate more broadly. ThadeusOfNazereth(he/him)Talk to Me! 19:17, 29 September 2026 (UTC)Reply
You mean this banner? The actual promise was "Please join the 2% of readers who give what they can to to help keep this valuable resource ad-free, up-to-date, and available for all". Nothing about supporting and defending open knowledge projects. The banner also talks about what you are doing when you visit Wikipedia, which I thought was a nice addition. Please note that something can be "very relevant to Wikipedia" without being what the donors were told their donations were funding.
You can see a bunch of fundraising banners here. I couldn't find any documentation as to which were shown when and to who. Perhaps someone else can find that information. --Guy Macon (talk) 23:45, 29 September 2026 (UTC)Reply
Guy, it sounds to me that you want the WMF to support Wikimedia and open knowledge projects but to do no research into how to support those projects, no research into how effective support for those projects is, whether support would be better directed towards maintaining existing (governance) structures or implementing different (governance) structures? If so, why do you think that? If not, please try explaining what I'm getting wrong. Thryduulf (talk) 19:47, 29 September 2026 (UTC)Reply
Not at all. The choice is not between "funding no research" and believing without evidence that this particular research actually ended up "maintaining/implementing governance structures". Can you name a single governance structure that was maintained or implemented by this research? Look at the results of this research again: The first two results are trivial and the third is just plain wrong.
In the banner at the donors -- many of whom live in poverty -- were not told that they were funding research, just as they were never told that they were funding trips to talk about Wikipedia in exotic vacation destinations. They were told that they were keeping Wikipedia ad-free, up-to-date, and available for all. I am fine with donors funding any research that has a reasonable chance of advancing those goals, but they should be told that some of their donations will be given to other individuals and organizations. --Guy Macon (talk) 23:45, 29 September 2026 (UTC)Reply
Research does not have to be successful to have been valuable, and if you only fund research that produces the results that you want it to then that's not research. Please explain how funding people and projects that write and maintain the projects, write and maintain the sources we use, research how the project can be improved, research how best to support the projects, etc. are not "keeping Wikipedia ad-free, up-to-date and available for all" either directly or indirectly? Thryduulf (talk) 00:04, 30 September 2026 (UTC)Reply
By those criteria I can't think of a single thing that the WMF spends money on that doesn't keep Wikipedia ad-free, up-to-date and available for all either directly or indirectly. I can't see any plausible way that not funding this sort of thing could possibly result in Wikipedia being ad supported, stop the volunteers from updating it, or make it unavailable, but you clearly do so I won't waste your time with further disagreement. --Guy Macon (talk) 00:18, 30 September 2026 (UTC)Reply
Can you name a single governance structure that was maintained or implemented by this research? To this among other research appears to be part of a body of research into policies that eventually led to the establishment of the NPOV Working Group and the subsequent implementation of the global baseline NPOV policy. Sohom (talk) 00:27, 30 September 2026 (UTC)Reply
Just to be clear: I think that Wikiresearch is in general a good thing, and funding some of this by WMF seems reasonable. The specifics of funding, methods, and transparency and ethics of the coordination of the research should, of course, be subject to transparent debate (apart from privacy issues). I'm not judging the validity of funding for this particular project. Boud (talk) 11:39, 30 September 2026 (UTC)Reply

I found meta:Baseline NPOV policy (a baseline NPOV policy for the Wikipedias which do not yet have an NPOV policy) to be a fine document. The question is whether it is worth the $42,356.25 USD ($88.43 per word) it cost the donors.
I would be interested in how many Wikipedias did not yet have an NPOV policy when it was created and how many have adopted it. If the donor is paying for "governance structures implemented by this research" it would seem reasonable to ask how many actual governance structures were implemented, not just a web page that seems like it would be useful to someone doing that.
I also noticed that it is only in English and French. One would think that some of that $42,356.25 USD would have been spent on translating it to the languages of the Wikipedias which do not yet have an NPOV policy. I'm just saying. --Guy Macon (talk) 16:14, 30 September 2026 (UTC)Reply
Couldn't small wikis just copy the NPOV policy from English, French or any other well-scrutinised Wikipedia, gratis, and alter any bits they deem inappropriate for their project? Certes (talk) 16:23, 30 September 2026 (UTC)Reply
@Certes, I would suggest looking at Serbian, Bosnian Wikipedia situations where despite copying the policy text over over they've had multiple significant issues surrounding the enforcement of said policies. Similar issues exists to smaller extent on other projects as well. Even on the rather bare-bones baseline policy, smaller communities have pushed back on the baseline policy due to concerns over the fact that it often constrains smaller language wikis in multiple ways in terms of prioritizing spoken history vs colonial written sources or preferring accounts in local language sources over the broader international consensus on a topic. Sohom (talk) 16:38, 30 September 2026 (UTC)Reply
Certes, the real problem in searching for worldwide solutions is that in my experience different language Wikis are effectively different planets with totally different climates. I sometimes look at the Italian, French, German and Spanish wikis, and their culture and customs are strikingly different, and range from those with a highly controlled system (eg German) to others with freewheeling on non-crucial topics. Their wiki cultures adre as different as the foods they eat. You can not give them a universal cookbook. Yesterday, all my dreams... (talk) 18:02, 30 September 2026 (UTC)Reply
@Guy Macon The policy hasn't yet gone into effect. It is still in proposal phase. Also re cost, 42K (and more actually) spent on building policy to avoiding incidents where Wikimedia project's credibility is called into question is imo money fairly well spent. Sohom (talk) 16:40, 30 September 2026 (UTC)Reply
That would indeed be money well spent if that's what the $42K bought. A bargain, I would say. The problem is that you have offered no examples where having One More Web Page On Meta actually avoided any incident where any Wikimedia project's credibility was called into question. Not even a plausible path where that might happen in the future. You seem to be saying that money spent on goals without any hint of a way to actually accomplish those goals is money well spent. I say that $88.43 per word is way above the going price for good intentions, which are currently so cheap that they pave roads with them. --Guy Macon (talk) 22:34, 30 September 2026 (UTC)Reply
The problem is that you have offered no examples where having One More Web Page On Meta actually avoided any incident where any Wikimedia project's credibility was called into question - The m:CheckUser policy and the m:UCoC would be prominent examples of the a "page on meta fixes problems" effect. Sohom (talk) 00:21, 1 October 2026 (UTC)Reply
Please accept my apologies in advance for saying this, but with phrasing like "Wikimedia Foundation's overarching programmatic goals to better the community" you could run for senate. The problem with their research results was similar. Too many words, too little substance. Sorry, but bluntness was needed here. Yesterday, all my dreams... (talk) 15:07, 30 September 2026 (UTC)Reply
Foundation's overall goal of making the wiki community better is what I meant. Ignore the word programmatic, it's a reference to the way WMF and other non-profits calculates it's budget, (as programmatic expenses and non programmatic expenses). Sohom (talk) 16:28, 30 September 2026 (UTC)Reply
Guy, I think $42,355.25 was wasted on the word salad they produced. The other $1 was useful, given that it made me laugh. Yesterday, all my dreams... (talk) 14:56, 30 September 2026 (UTC)Reply

In this edit, Guy Macon wrote, The first two results are trivial and the third is just plain wrong. I think it would be good if Textaural, who is an experienced enwiki editor, could respond here. The peer-reviewed research paper is technically a WP:RS, so usable in enwiki articles, but peer review does not guarantee correctness.

I agree that the third key result The status of rules [on these 5 Wikipedias is] established largely by individual decisions and rarely through collective decision making is extremely dubious for at least enwiki and frwiki. The strings !vote and not vote seem to be completely absent from the published paper. In a nutshell: If I'm the individual to first write the enwiki policy that 1+1=2 is a fact and nobody contests that, it's not due to my dictatorial authority; it's due to the WP:!VOTE decision-making algorithm and the nature of 1+1=2 being widely considered as reasonable, without anybody needing to argue about it (leaving aside modern foundations of mathematics). The whole point of not holding a vote on my hypothetical new policy is that getting 10,000 votes (not !votes) with a likely result of 99.9% in favour would be a huge waste of time. @Textaural: How do you justify your extraordinary conclusion that Within the English-language rules, the authority of rules was based on the editorial authority of one user, who may have found support through limited deliberation without having studied the role of WP:!VOTEs in the establishment of enwiki rules? !VOTEs are a form consensus decision-making, which is a form of collective decision-making. The authority of rules is not based on the editorial authority of one user; it's based (often) on the ability of one user to successfully describe and summarise the current consensus or predict the likely consensus.

Textaural: do you have any evidence for this claim that you and your authors assert in the paper? The Methods section of your paper says nothing about how you quantify the unexpressed non-objections to enwiki rules (a possible method would be to obtain permission to do a survey of Wikipedians to see how much they agree with existing rules and whether or not they have tried objecting to rules they disagree with; but you don't seem to have done that). Boud (talk) 11:07, 30 September 2026 (UTC)Reply

"Rules" !== policies to my understanding. Sohom (talk) 11:23, 30 September 2026 (UTC)Reply
The author gave examples of what they consider to be rules. Two of the examples are in English; Wikipedia:Proposed deletion (a policy) and Wikipedia:Disruptive editing (a behavioral guideline.). --Guy Macon (talk) 12:37, 30 September 2026 (UTC)Reply
Fair enough, my understanding is that the paper discusses the individual rules inside the policies instead of the policy as a whole but I can see it being taken in both ways Sohom (talk) 14:04, 30 September 2026 (UTC)Reply
I also noticed that the author uses https://artandfeminism.org/resources/research/unreliable-guidelines/ as a source. Artandfeminism.org appears to be funded by the WMF through WikiCed ( https://www.wikicred.org/). I can't find any record of how much the WMF gave artandfeminism.org to say "This research project identifies the ways that organizational values and processes around reliability and the reliable source guidelines are implicated in maintaining hierarchies and excluding marginalized knowledges and communities" and "Ultimately, the Reliable Source guidelines are an unreliable and incomplete guide to provide editors with a meaningful understanding of how to assess source reliability".
So what are we supposed to replace our reliable source guidelines with? "Triangulation". If anyone can understand what (cited by artandfeminism.org) is talking about please explain it to me. --Guy Macon (talk) 13:16, 30 September 2026 (UTC)Reply
I don't disagree that Art+Feminism is funded to a certain extent by the Foundation, but last I checked, it wasn't through WikiCred? WikiCred organizes meetups in the West Coast and builds tools for media reliability on Wikipedia, Art+Feminism advocates for better coverage of topics related to gender, feminism and art (among other things). They were one of the advocates for a global UCoC and and have long advocated for changing our guidelines to be more inclusive towards marginalized communities and better enforcement of civility policies. I don't think there is a correlation between WMF funding Art+Feminism, Art+Feminism being cited as community criticism of the reliability guidelines and WikiCred organizing conferences. Sohom (talk) 16:25, 30 September 2026 (UTC)Reply
Re: "last I checked, it wasn't through WikiCred?" you need to do a better job of checking. Look at Read the funding section. What we have here is an informal community of organizations united in the goal of writing proposals that will result in the WMF giving them grant money. If only the WMF would check to see if the goals in the proposal were actually met and, if not, find some other applicant to award the next round of grants to. --Guy Macon (talk) 23:00, 30 September 2026 (UTC)Reply
To quote the paper, Unreliable Guidelines, a project of Reading Together: Reliability and Multilingual Global Communities, is an inaugural Art+Feminism research report partially funded by WikiCred, which supports research, software projects and Wikimedia events about information reliability and credibility. WikiCred and Art+Feminism are together funding the paper with At+Feminism being the primary sponsor and WikiCred being a partial sponsor. The Funding section in particular say The expertise of the main researchers was partially funded by WikiCred and Art+Feminism. i.e. that this project was a collaboration between AF and WikiCred. Nowhere does it say that the Art+Feminism is funded by WikiCred. Sohom (talk) 00:27, 1 October 2026 (UTC)Reply
When I skimmed that paper yesterday, I also had several qualms with the methodology leading to those conclusions: off the top of my head, they ignored pages with a creation date before 2005; they only looked at edits which change the label (between essay, guideline, policy, etc.); and they only looked talk pages when explicitly mentioned in an edit summary.
To expound a little: if I'm reading it correctly, it would ignore (to list a few non-exhaustive examples), talk discussions not mentioned in the edit summary; discussions which affirm the status quo; changes which align content to community consensus without changing the "label" of a page; and implicit consensus where thousands of people have read a page and no one has found reason to change it. LittlePuppers (talk) 21:56, 30 September 2026 (UTC)Reply

Can I suggest separating, to the extent possible, the funding question from an evaluation of the paper? The WMF does not read the outcome of a project before funding the project, of course, so really it's a question of "should the WMF invest in research about Wikipedia governance" not "should the WMF invest in this paper". Yes, "was this worth it" is a fine question to ask, and one that the WMF answers internally, too, but any evaluation of the funding question should be based not on "I found one I'm skeptical of" but on a full accounting of related projects. I haven't engaged deeply with this paper yet (open in a tab in my "to read" window for some time, sadly), but the project sounds worth doing to me. We have precious few studies that examine Wikipedia governance anywhere but the English Wikipedia, and a direct comparison between several language editions is very potentially valuable. And NM&S is easily one of the top journals for non-computational Wikipedia research these days. That doesn't mean everything it publishes is perfect, but it has a high rejection rate, intense peer review process, and deep bench of Wiki-knowledgeable reviewers. That said, as with any Wikipedia research, it's also valuable when volunteers evaluate the content and correct the record, if appropriate. A quick search through the signpost archives doesn't turn up a match, so maybe it's a good fit for the next Research Report. — Rhododendrites talk \\ 16:16, 30 September 2026 (UTC)Reply

Here is a quick overview of highlights from the Wikimedia Foundation since our last issue on 15 September. Please help translate.


Starter Kit dashboard overview.

Highlights

  • Introducing the Wikipedia Starter Kit: The Foundation has launched a new Toolforge-hosted tool giving new and small language Wikipedia communities a guided pathway through essential onboarding tasks, activity indicators, and connections to the wider movement.
  • Apply for a scholarship to Wikimania 2027: Scholarship applications are now open to attend Wikimania 2027 in Santiago. To help with preparing submissions, orientation sessions in Spanish, Portuguese, and English will be held starting the week of October 5. The final deadline is October 31.
  • Reading Lists rolling out to all Wikipedias: Starting September 28, logged-in users on every Wikipedia will be able to save articles to a "read later" list using the bookmark icon, following successful rollouts on Arabic, Bengali, Chinese, Czech, English, French, Indonesian, and Vietnamese Wikipedias.
  • WikiCelebrate: WikiCelebrating Juan "BugWarp" Kulichevsky, who has contributed to Wikimedia projects since childhood and is passionate about sports photography and community building in Argentina and beyond.
  • The 2026 appointment cycle is open now through 31 October 2026 for the Ombuds Commission (OC) and the Case Review Committee (CRC). You can learn more about the committees, the roles, and how to apply on the Committee Appointments page on Meta-wiki.


Annual Goals Progress on Engage
See also: Growth · Product Safety and Integrity · Tech News · Language and Internationalization · The Wikipedia Library · list of movement events · Wikifunctions & Abstract Wikipedia

  • Community Wishlist 2027: The Community Wishlist Process is now updated, taking into account the feedback received. Look out for the call for wish submissions towards the end of October.
  • Wikimedia CEE Meeting 2026: From Sep 18-20, Wikimedians, open-knowledge researchers, and free-culture advocates from Central and Eastern Europe gathered for the first time in Romania. Take a look at the novelties and history behind CEE Meeting.
  • Increasing account creation: To convert more casual users into logged-in daily users, the evergreen account creation prompt is released into Beta / TestFlight for iOS and Android. Also, a new post-edit notice designed to encourage Temporary Account holders to register for a permanent account will now be released to all wikis. The new notice has proven to increase permanent account creation from 1.94% to 3.66%.
  • VisualEditor is now the default at English Wikipedia: this wiki has now adopted VisualEditor as the default editor on both desktop and mobile web, following a sitewide discussion that found overwhelming consensus to make the change.
  • Retesting image carousel: The image carousel experiment will be retested using three new versions of the design, updated based on community feedback. The team will assess results and determine with communities whether or not to proceed with the feature.
  • Wikifunctions: Your code can now call back to the Wikidata fetch functions. Wikifunctions have deployed an initial capability by which a code Implementation (in JavaScript or Python) can call certain other Wikifunctions Functions, get back the result, and have it available for further processing.
  • Latest experiments: A current experiment is testing whether offering a "generic guidance" path for Article Guidance that does not depend on a Wikidata match would increase the chances of more article creation attempts. See all live, upcoming, and completed experiments in Product & Technology.
  • Tech News: The latest highlights from Tech News weeks 38 and 39 include a Reader Growth experiment testing whether showing only part of an article at first on mobile, with a “Read more” button to reveal the rest, encourages people to keep reading. See also the 56 community-submitted tasks that were resolved over the last two weeks.
  • Suggestions for Transliterated Search: If you regularly use two different writing systems (such as Hindi and Latin alphabet), sometimes you accidentally type with the wrong one and need to switch to the right keyboard. Now, this gadget will be able to suggest query based on what you typed.
  • Peer-to-peer learning program: Let’s Connect, a peer-to-peer learning program is shifting to a community-organized model.
  • LLM Paste Check rolling out to Wikipedia: A new Edit Check will prompt editors who paste text likely copied from an external AI chatbot to consider whether it aligns with movement AI-use policies and decide whether to keep or remove it. It goes live as a default-on feature at English Wikipedia on Sep 24 and at all Wikipedias on Oct 8.


Annual Goals Progress on Enable
See also: Research newsletter · WikiLearn News · other newsletters on MediaWiki.org

  • Wikidata bulk downloads now free via Wikimedia Enterprise: New beta API endpoints let anyone download full Wikidata snapshots and pull hourly batch diffs at no cost, opening up bulk access that previously required a paid tier.
  • Accountability for Wikimedia production services: To ensure clear accountability for responding to incidents in Wikimedia production services, Wikimedia Foundation published the Service Catalog page. The page is the canonical source of service ownership for code deployed to the Wikimedia Foundation cluster. See also FAQ pages which contain more background details.
  • Datacenter switchover has been postponed: The equinox exercise revealed capacity issues, preventing the switch of all services to the other datacenter. The process has been paused, prioritizing investigating that issue, to ensure that we continue to be able to serve our users reliably. A new exercise will be scheduled. Meanwhile, all traffic and edits continue to work as usual.
  • Affiliates and grant funding: The conversations on the new proposed model, requirements and criteria continues until the end of next week with public calls on 1) Tuesday, September 29, 16:00-17:00 UTC on affiliate growth paths and Friday, October 2nd, 17:00-18:00 UTC on affiliate overlaps.


Annual Goals Progress on Reach
See also: Wikimedia Apps · Readers

  • AI literacy program: A new regional initiative in ASEAN was launched to support training and conversations around AI literacy, accountability, and Indigenous self-determination.
  • Testing Google preferred sources: An experiment is proposed on English Wikipedia to run a temporary (2-week) notice asking readers who come from Google to set Wikipedia as a “preferred source” in Google. We would see if this causes people who use Google to visit Wikipedia more. If your language wiki is interested, please reach out on the project's talk page.


Other Movement-curated newsletters & news
Diff blog · Goings-on · Planet Wikimedia · Signpost (en) · The Headword (en) · Kurier (de) · Actualités du Wiktionnaire (fr) · Regards sur l'actualité de la Wikimedia (fr) · Wikimag (fr) · Education · GLAM · Milestones · Wikidata · Central and Eastern Europe · other newsletters

MediaWiki message delivery 01:50, 30 September 2026 (UTC)Reply