One lost source and the story changes
Assumes When two of them are first attested together and A record is not a proof.
When two of them are first attested together builds a bracket out of the record and finds it firm — the year at which all five lineages of practice are jointly attested is 1838, one year from where the same claims are popularly dated together. It ends on the observation that a joint date inherits its latest member and therefore rests entirely on that one entry.
The obvious next question is how much weight any single entry can carry, and the record answers it, because every entry states how many independent surviving sources stand behind it.
The numbers are small.
Five claims with one witness
The record carries fifteen entries and the source counts run from one to four. Five of the fifteen rest on exactly one surviving document:
- Paper reaching Japan, on the Nihon Shoki.
- Recreational folding in Japan, on Saikaku’s poem of 1680.
- The thousand cranes, on the Hiden Senbazuru Orikata of 1797.
- One fold solving a cubic, on Beloch’s paper of 1936 — which runs the other way, being attested fifty-five years before it is popularly dated, and is fifty years in the wrong language for exactly that reason.
- The traditional one-cut star, on a secondary account of 1873 — a trick a century older than the description and two centuries older than the theorem behind it.
A third of the record has a single witness, and none of it is an inference — every entry names something that survives. Take one document away from each and those five claims are not weakened — they are removed: nothing survives that attests them at all, and the entries become inferences rather than attestations.
That is a different kind of dependency from the usual one. A measurement with an uncertainty gets worse when a data point goes; a claim with one source ceases to be a claim.
What their removal does to the record’s headline
The record’s most-quoted statistic is the overrun: how far ahead of its earliest evidence a popular date runs. Eight of the fifteen claims run ahead at all — five of the remaining seven run the other way, being attested before they are popularly dated — and the field’s whole finding is that the eight run ahead a great deal.
Over the whole record the mean overrun is 357 years. Set the single-witness claims aside and it is 193 — a fall of nearly half.
The median over the whole record is 201.5 years. Set the same claims aside and it is 184.5 — a fall of seventeen years, which is nothing.
A statistic that halves and one that does not move is the signature of a concentrated effect. A mean is pulled by its extremes and a median is not, so a mean that collapses while a median holds says the extremes were doing the work.
They were. The two largest overruns in the record are recreational folding at nine hundred and eighty years and the thousand cranes at eight hundred and ninety-seven, and each of them rests on exactly one surviving document. Take those two documents away and the field’s headline phenomenon is a mild one.
What that does and does not mean
Three readings are available and they need separating, because the strong one is not supported and the weak one is not the point.
Not supported: the two claims are wrong. Saikaku’s poem is a printed source of known date and the Senbazuru Orikata survives. Nothing here casts doubt on either; what is at stake is what would be known if they did not.
Not supported: the record overstates the phenomenon. The overruns are what they are. A mean computed over the surviving record is a correct statement about the surviving record.
Supported: the shape of the field’s evidence is an accident of what survived. If two particular documents had burned, the record would show eight claims overrunning by a mean of under two hundred years and the field’s characteristic finding — that its dates run centuries ahead of anything that attests them — would look like a modest correction rather than a phenomenon.
That is the reading worth having, and it is uncomfortable in a specific way. A field’s most striking result is resting on two documents, and it is not resting on them because they are decisive but because they are the only ones there are.
What the two documents are
The whole finding rests on two particular objects, so they are worth naming rather than leaving as rows.
Saikaku’s poem of 1680 mentions folded butterflies. It is a printed source of known date and it is the earliest thing that attests recreational paper folding in Japan at all. The practice is popularly dated to around 700, so the overrun is nine hundred and eighty years — the largest in the record.
The Hiden Senbazuru Orikata of 1797 is the earliest surviving book of recreational paper folding, and its connected cranes are cut from one sheet rather than folded from separate ones. The thousand cranes are popularly dated to around 900, so the overrun is eight hundred and ninety-seven years.
Both are printed, both are of known date, and both are the only surviving thing for their claim. And they are not marginal documents in this field — the second is the single most-cited object in the whole of paper folding’s history, so its being a single witness is not obscurity but the shape of what survives from the period.
That is the sharpest form of the rung. The two largest gaps in the record are carried by the field’s best-known document and by a poem, and if either were lost the field’s characteristic finding would be half of what it is.
Why the concentration should have been expected
The result reads as a surprise and it is arguably the null, which is worth arguing because it changes how much weight to put on it.
An overrun is the distance between when a thing is popularly dated and when it was first written down. Both halves push the same way for a practice that grew: the popular date drifts earlier because a tradition’s account of itself gets older in the telling, and the earliest source is late because a growing practice is not written down until it is already ordinary.
The two effects multiply for the oldest claims and barely act on recent ones. A syllabus published in 1838 has neither: nobody has had time to make it older and it was written down deliberately.
So a record containing a few ancient practice claims and many modern result claims should show exactly this — a handful of enormous overruns and a body of small ones. The distribution is skewed by construction, and a mean over it was never the right statistic.
Which turns the rung’s finding into a methodological one rather than a discovery about the past. The field quotes a mean because a mean is what one quotes, and the quantity it summarises is not the kind that has a meaningful mean. A record mixing claims of different ages and different kinds has no typical overrun, and the median is a better summary precisely because it ignores the two rows that carry the story.
The kinds of source, and where the fragility is
The record grades sources as well as counting them, and the two gradings do not line up in the obvious way.
Of the fifteen entries, two rest on artefacts, two on manuscripts, nine on printed books and two on secondary accounts. None is an inference.
The single-witness entries are spread across those grades — one manuscript, three printed books and one secondary — so fragility is not the same as weakness of kind. A printed book of known date is a good source and a claim resting on exactly one of them is a fragile claim, and the two properties are independent.
That is worth stating because a reader grading the record by source kind would rank the printed entries well and miss the fragility entirely. The record carries both numbers and the field’s own summaries use the kind rather than the count, which is the omission this rung is an instance of: the quantity that is easiest to describe is not the quantity that decides how much the conclusion depends on any one thing.
A leave-one-out is the right instrument
The construction here has a name in other fields and it is worth borrowing deliberately, because it is the instrument this field lacks.
A leave-one-out asks what a conclusion becomes when each data point is removed in turn. It is standard where the data are measurements and it is almost never done where they are documents, presumably because a document does not feel like a data point.
It should. A record is a sample of what survived, the survival was not random, and a statistic over it inherits the sampling. Asking what each entry contributes to the headline is exactly the question a statistician would ask of any small sample, and this record is fifteen rows.
Doing it here takes one line and produces the finding above. Doing it on a larger historical record would take longer and would presumably produce the same shape of answer more often than not, because the distribution of what survives from a period is famously uneven and the surviving items are not a random draw.
The one thing that makes it awkward is that the removal is not hypothetical in the usual sense. Removing a measurement asks what the conclusion would be with less data; removing a document asks what it would be if a particular fire had gone slightly differently, which is a counterfactual with a date and a place attached to it.
What a second column would settle
The rung’s own limitation is that a source count is a single number, and it is worth saying exactly what a second number would buy, because the answer is more than it sounds.
The record states, for each claim, how many independent sources survive. It does not state when the second one is, and that is the field the fragility question actually turns on.
Consider two claims each with three sources. In the first they are 1600, 1610 and 1650, so losing the earliest moves the attestation by ten years and the claim barely notices. In the second they are 1600, 1780 and 1830, so losing the earliest moves the attestation by a hundred and eighty and the claim’s whole character changes — it stops being a seventeenth-century attestation and becomes an eighteenth-century one.
Those two are identical in this record. The count says how many losses a claim survives and says nothing about what surviving costs it, and the second is the quantity a leave-one-out ought to be measuring.
Adding a second-earliest year to each entry would turn the block removal here into a proper one-at-a-time: remove each earliest source in turn, recompute the overrun from whatever is next, and read off how much each document is worth to the record’s headline. Five of the fifteen would go to nothing, as they do here; the other ten would move by an amount nobody currently knows.
That is a small piece of data collection with a definite output, and it is the thing this ladder now owes. It would also make the joint-date bracket the rung below this one builds checkable in the same way, since a bracket resting on a claim whose second source is a century later is a different bracket from one resting on a claim whose sources cluster.
Which theorem was checked and how
The source counts go through the record’s own checker, which refuses an entry whose count is not a positive whole number — so a fragility computed over an entry with no count would not be computed at all.
The mean must fall by more than thirty per cent when the single-witness claims are set aside, or the collapse the essay is about did not happen and the figure refuses.
The median must move by less than a fifth of the mean, which is the other half of the finding: an effect spread through the record would move both, and the figure would say so rather than being drawn and misread.
And the two largest overruns must both be single-witness. That is the concrete form of the concentration claim, and it is asserted rather than pointed at — a record in which the extremes were well-attested would make this rung’s argument false and would refuse to draw.
Where the model stops
One loss per claim. The arithmetic removes the single-witness entries as a block, which is the effect of one loss falling on each of them. Losing two sources from a four-source claim is a different question and is not asked.
Independence is asserted by the record rather than checked. Two sources counted as independent might both descend from a third that has not survived, in which case the claim is more fragile than the count says, and nothing here can detect it.
The overrun is popular date minus earliest source year, and the popular date is a stated quantity rather than a measurement of what anybody believes. Every statistic here inherits that.
And the counterfactual is a counterfactual. No document was lost; the arithmetic says what the record would look like if some had been, and that is a statement about the record’s dependence rather than about the past.
What the picture cannot show
The bars show how many sources survive and cannot show how many did not. A claim resting on one document today might have rested on twenty in 1800, and the record has no way to carry that — so a single-witness entry is a statement about now.
Nor can it show the direction a loss would push a claim. Removing a claim’s only source removes the claim; it does not make the practice younger, and a reader who took the reduced mean as a better estimate of anything would be making exactly the error this rung is warning about.
The most conspicuous absence is the claims that are not in the record at all. A practice attested by a document that burned is a practice this field does not know about, and the leave-one-out is a lower bound on how much the survival matters — because it can only remove what survived, and cannot restore what did not.
The idealisation, named
The record is a set of claims, each with one popular date, one earliest source with a year and a kind, and a count of how many independent sources survive. Everything here is arithmetic on those four fields.
The strongest simplification is that a claim has one earliest source and a count. A claim with four sources spread over three centuries and one with four clustered in a decade have the same two numbers here and are in quite different positions — the first would survive losing its earliest and keep a source two hundred years later, and the second would barely move.
That matters for this rung’s conclusion in a stated direction. The record’s fields cannot distinguish a robust multi-source claim from a fragile one, so the four entries with three or four sources are treated as safe and might not be — the same shape of unstated hypothesis this collection has imported once before. The five single-witness entries are unambiguous, which is why the arithmetic is built on them.
And the record itself is fifteen rows chosen by whoever built it. It carries the claims this collection’s essays argue about, which is a selection with a purpose rather than a survey — so every statistic here is a statistic about a working list, and the essay’s claim is about that list’s dependence on its own weakest rows.
Where the ladder goes next
This ladder has taken three rungs to get from a story about a merge to an arithmetic about survival, and the next thing it owes is a field rather than a computation.
The record needs a second source column. Every entry states how many independent sources survive and none states what they are or when they are, so the leave-one-out here is the only fragility question the data can answer. Recording the second-earliest source for each claim would make the difference between a robust multi-source entry and a fragile one visible, and it would turn the counterfactual above from a block removal into a genuine one-at-a-time.
Sideways from here, the concentration finding belongs beside the record’s own account of what it cannot establish. That essay says a date is not derivable and that this field’s figures draw the shape of the evidence rather than the evidence; this rung finds that the shape has two rows in it doing most of the work, which is the sharpest thing the shape can be asked.
The habit worth carrying is one line long and it applies well outside this subject. When a mean moves and a median does not, the finding is a few rows rather than a tendency — so the next question is which rows, and the one after that is how much anybody should trust them.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- Nothing here is as old as it sounds attestation · documentary record · hiden senbazuru orikata · primary source · senbazuru
- A file has no paper documentary record · primary source
- Found by people not folding paper documentary record · independent discovery
- Nearly every cutting fails at one crane hiden senbazuru orikata · senbazuru
- The prediction held at eight and broke at ten hiden senbazuru orikata · senbazuru
- Which cranes can stay joined hiden senbazuru orikata · senbazuru
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AttestationDocumentary recordHiden senbazuru orikataIndependent discoveryPrimary sourceSenbazuru