What an AI reading Epistolæ, a corpus of medieval women’s letters, recovers that a graph of senders and recipients cannot.
Press enter or click to view image in full size
In 1201, in the county of Toulouse, 55 knights of Murel swore loyalty to Marie of Montpellier and her two young daughters. The document that records the oath survives as one letter in a large digital collection of medieval correspondence.
Turn that collection into a network the usual way and the letter almost disappears. The index gives one node for the knights and joins it to Marie and to her two daughters, two links and nothing else.
Now read the letter instead of cataloguing it. The sender is not one person. It is 55 men, each named in turn, Arnald Mascaron of Murel, Bernard Faber, Pontius Gerald, John Cavarada and the rest, and once the notary, the witnesses and the lords named only to fix the date are added, the letter names 62 people. A whole local hierarchy that the network had folded into the word “knights”.
Press enter or click to view image in full size
That gap, between two links and 62 people, is where this piece starts. An oath is an extreme case, a public act built to list everyone present. The real question is whether it is a freak or the visible end of something the graph does everywhere, and the only way to settle that is to measure.
Why a collection of women’s letters
The collection is Epistolæ, Medieval Women’s Letters, gathered and translated by Joan Ferrante, Professor Emerita at Columbia University. As I retrieved it, it holds just over two thousand letters to and from women, spanning the fourth to the thirteenth centuries.
I chose it on purpose, and the reason belongs at the front. Medieval women rarely held titled office. Their power was usually informal, exercised by asking, mediating and protecting rather than by signing in an official capacity. That kind of power lives in what a letter says, not in who is recorded as its sender. So a corpus of women’s letters is a sharp place to see whether a graph of senders and recipients is losing something.
One thing has to be said plainly, precisely because the choice was deliberate. This collection was assembled around women in the first place. That makes it a fine stage for showing what reading recovers, and a poor basis for any claim that compares women with men, since no neutral body of men’s letters sits beside it. I make no such claim. The finding is about the method, and the method holds for any historical letter collection. The women’s corpus is simply where the loss is easy to see.
The two networks
To measure the gap I built two versions of the network and compared them.
The first is the one most studies use. One node for every person who sends or receives a letter, one line for every letter that passes between two people. Nothing in the body is read. It comes straight from the index. Across the whole corpus that graph has 978 people and 1,187 links.
The second comes from reading. I had a language model go through the body of each letter and list, in a fixed form, every person it names, and in particular every act of intercession, meaning someone asking a powerful person to help or to harm a third party. Asking on behalf of another is the basic grammar of medieval favour, and it always involves at least three people, of whom at most two are the sender and the recipient.
One rule kept the exercise honest. The model only reads and lists. It never counts and it never draws a conclusion. Every figure below is then produced from its lists by ordinary, inspectable code. The reading is the model’s job. The arithmetic is not.
I ran this on a sample of 30 letters, 8 chosen by hand to calibrate the reading and 22 drawn by a seeded random generator, so the sample cannot be accused of being picked for its richness. The oath of Murel came up in the random draw.
Two kinds of absence
Across the 30 letters the metadata graph connects 54 people. The letters name 203. Before making anything of that ratio, two groups have to come out. Twelve of the 203 are figures from scripture or antiquity, Abraham, Moses, the Queen of Sheba, quoted as examples rather than acting in the story, and counting Moses as part of a medieval social network would only pad the result. Another 24 are not people at all but institutions and places, churches, castles, a cathedral chapter. What is left is 167 real people, close to three times the 54 the graph connects.
Press enter or click to view image in full size
Those 167 real people are missing from the graph in two different ways. About 31 of them the graph already has a node for. They write or receive letters elsewhere in the corpus, and here they are only mentioned. The node exists. What is missing is the link, the particular relationship this one letter records, such as a pope being petitioned or a king named in passing.
Get Carmine De Stefano’s stories in your inbox
Join Medium for free to get updates from this writer.
The other 136 appear nowhere in the corpus as a sender or a recipient, ever. These are missing nodes, people the map never held at all. They are the sworn knights, local scribes and notaries, bishops known only by their town, and a few larger figures who simply do not write here, such as Béla IV of Hungary and Frederick Barbarossa. In short, two thirds of the names the letters use are real people the graph never held.
Now the honest part, because the size of that effect is not spread evenly. It lives in the administrative documents. The oath of Murel alone supplies 59 of those 136 hidden people. Nearly half the sample, 14 letters of the 30, are charters, grants and confirmations, and those carry the longest lists of witnesses and kin. Read only the 15 plainly personal letters, setting the oath and the governance documents aside, and the graph still misses people, but far fewer, almost as many again as it connects. Reading roughly doubles the population, about two hidden people per letter rather than a crowd. So the loss is real everywhere and dramatic where the document is a public act. A graph drawn from the index feels the difference least exactly where a historian would look hardest.
The same person in more than one role
There is a second thing the metadata graph cannot show, and this one does not depend on the kind of document at all. Every line in that graph is the same undifferentiated tie. Yet the role a person plays is what matters, and roles live in the text. Across the 30 letters the reading found 11 acts of intercession, one person asking another to act for or against a third. When Eleanor of Aquitaine writes to Pope Celestine III to free her captured son Richard, the letter puts it without ceremony. “Our king is confined and on all sides anguish oppresses him.” That is a requester, a mediator and a beneficiary in a single line, and a graph of senders and recipients can hold only the first two of the three.
Across the sample the women appear in every position of that grammar. Some are asking a powerful person for a favour, like Eleanor or Berengaria of Navarre. Some are the power being asked, like Blanche of Castile or the empress Theophanu. Some are acted for or against, as when a pope rules against Marie of Champagne. Some act in their own legal authority by issuing a charter, which 7 women do here, and a charter has no real recipient, so the network drops the act altogether.
The sharpest point is that the same woman plays different roles in different letters. Marguerite of Provence petitions the pope in one letter and is the person his ruling protects in another. Elizabeth of the Cumans is the power a king comes to ask, and separately a landowner endowing a religious order on her own authority. A single fixed node cannot be both the one who asks and the one who is asked. The text says she was both. The graph has to choose.
The names are the easy part
A fair objection is that finding names needs no language model. It does not. A short script that grabs every capitalised phrase finds plenty, and a trained name recogniser would do better still. But neither can tell who is asking whom for what. Turning a heap of names into “this person asked that person to free a prisoner” is the reading, and the reading is where the roles come from. The names are the easy part. The roles are the point, and the roles are what a model recovers and a pattern match cannot.
What I did to keep it honest
The point of the exercise was to measure a gap, so a few limits belong in plain sight rather than in the small print.
This is a demonstration on 30 letters, not a statistic for the whole corpus. It sizes the effect and shows the method works. Running it across the full collection is the next step, and the exact ratios will move when it does.
Nearly half the sample is governance documents, which is why the count above is given by kind of document and not as one headline number. Charters and oaths are where the graph loses most, and folding them together with personal letters into a single multiplier would overstate the everyday case.
Of the 30 letters, 8 were read a second time and checked name by name against the source, which is how I know the model is not inventing people or quietly dropping them. The other 22 rest on that calibration.
The model only extracted. Every number here, the 54, the 203, the split into 31 and 136, the 62 people in the oath, was computed by plain code from the model’s lists. The one thing quoted from a letter above, Eleanor’s line, is Ferrante’s translation of the source, not a sentence the model wrote.
One name among those real people is really a figure of speech. The Virgin Mary is a correspondent elsewhere in the collection, so she counts as a node the graph already holds, yet here she is invoked as an example, not acting in the story. Leaving her out changes 167 to 166 and moves nothing else, so she is left in and flagged rather than quietly dropped.
The name matching is deliberately strict. It refuses to merge, say, Henry I of Champagne with Henry I of the Franks, who are different men, and as a side effect it also misses a few genuine matches, such as one spelling of Philip II Augustus. So the count of people already in the graph is a floor and the count of missing people is, if anything, slightly high. That makes the hidden population a conservative figure, not an inflated one.
One caveat about reproducibility. The arithmetic and the extracted lists are in the repository and can be rerun and inspected line by line. The reading itself was done inside the session rather than through a script that anyone can point at a model, so it is published as data to check, not yet as a button to press.
The point
A network drawn from a catalogue is not a photograph of a social world. It is the world’s index, and an index leaves out the contents by design. On this small sample, reading the contents recovered a population the map never held, largest in the administrative documents and smaller but real in the personal letters, and recovered a distinction the map cannot hold at all, the difference between the one who asks and the one who is asked.
None of this means the historians were wrong. They have always read the letters, one at a time. It means that when a corpus is compressed into a graph so it can be studied at scale, it is worth knowing exactly what the compression discards.
What is new is that the reading can now be done by a machine. Not a scan for capitalised names but a reading that follows who asks whom for what, carried across a whole archive rather than a single shelf. An agentic system that reads each letter, records its people and its acts, and leaves every step open to inspection makes close reading something that scales without ceasing to be reading. For historiography this is a rare kind of opportunity. A corpus too large to study except as an index can be read back into its full population and its roles, with the working shown so that every claim can be checked against the source. The archives that had to be flattened to be handled at all can now be met as what they always were, dense records of people acting on one another.
The method, the code and the full line by line output are in the repository. The corpus is Epistolæ, Medieval Women’s Letters, directed by Joan Ferrante at Columbia University, at epistolae.ctl.columbia.edu.