Who found it, and when

The interval is wider than the number

The field's most-quoted statistic — how far ahead of its evidence a popular date runs — is the mean of eight positive gaps, and it is quoted as three hundred and fifty-seven years. Resampling the fifteen entries puts ninety per cent of its weight between a hundred and fifty and five hundred and ninety-one. The interval is wider than the number, the middle gap's interval is a fifth as wide, and the difference is the same two documents the leave-one-out found.

Assumes One lost source and the story changes and When two of them are first attested together.

One lost source and the story changes takes the five single-witness entries away and watches the field’s most-quoted statistic fall from 357 years to 193, while the median hardly moves. Its conclusion was that the effect is two documents rather than a tendency.

That is a statement about which entries the number rests on. It is not a statement about how much the number is worth, and the second question has a cheap answer that nobody has computed: resample the record and see what comes back.

The interval is wider than the numberHow far ahead of its evidence a popular date runs, recomputed under twenty thousand resamplings of the record. The statistic is quoted as a single number; the resamplings put ninety per cent of its weight across a range four times as wide as its own distance from zero would suggest.200400600800100002004006008001000120014001600mean years ahead of the evidenceresamplingsquoted: 3575%: 15095%: 591spread 13320,000 resamplings of the fifteen entries · five per cent of them fall below 150 and five per cent above 591
Fig. 1 How far ahead of its evidence a popular date runs, recomputed under twenty thousand resamplings of the fifteen entries. The quoted value is 357; five per cent of the resamplings fall below 150 and five per cent above 591.

Eight numbers, one statistic

The statistic is the mean of the gaps that are positive — the claims popularly dated earlier than anything surviving for them. There are eight of those among fifteen entries, and they run from a single year to nine hundred and eighty.

A mean of eight numbers spread over three orders of magnitude does not have three significant figures in it. Resampling says by how much: drawing fifteen entries from the fifteen with replacement, twenty thousand times, and recomputing each time gives a distribution with a spread of 133 years and a ninety per cent interval running from 150 to 591.

So the honest statement of the field’s headline figure is that it lies somewhere between a century and a half and six centuries, and 357 is a point in that range rather than a measurement of anything. The interval is wider than the number.

The middle gap is a different kind of quantity

The same resampling on the median tells the opposite story, and the contrast is what makes both of them useful.

The middle gap barely movesThe middle gap of the whole record under twenty thousand resamplings of the fifteen entries. It stays within a narrow band, because a median is decided by the entries in the middle and the record has many of those.200400600800020004000600080001000012000middle gap, yearsresamplingsquoted: 15%: -1095%: 110spread 8120,000 resamplings of the fifteen entries · the middle gap stays near where it was
Fig. 2 The middle gap of the whole record under the same twenty thousand resamplings. It sits at one year and its ninety per cent interval runs from −10 to 110.

The median of all fifteen gaps is one year, and its ninety per cent interval runs from −10 to 110. That is still wide in absolute terms, and it is a fifth the width of the mean’s interval on a quantity five hundred times smaller.

The reason is structural. A median is decided by whichever entries happen to sit in the middle, and this record has several entries close together near zero; a resampling that loses one of them finds another. A mean of positive gaps is decided by the largest entries, and this record has exactly two very large ones; a resampling that loses them both — which happens often enough to matter — produces a quite different number.

That is the same finding as the leave-one-out, arriving as a property of the estimator rather than as a fact about two documents. A mean is sensitive to its tail and this record is mostly tail. A record is not a proof says the field’s figures draw the shape of the evidence rather than the evidence; this says that even the shape has an error bar, and that the error bar depends on which summary is taken.

Which entries the number is made of

Which entries the number is made ofThe statistic recomputed fifteen times, each without one entry. Removing the two largest gaps takes it down by nearly a third; removing an entry dated close to its own evidence pushes it up, because that entry was pulling a mean of large numbers toward zero.leave one out, fifteen timesthe statistic is the mean of the gaps that are positive, so an entry leaves the mean when its gap doeswithout this entrythe statistic becomesmoves byits own gaprecreational-folding268-89980senbazuru280-77897ceremonial-wrapping351-6400paper-china35700paper-europe3570-94beloch3570-55yoshimura3570-18vertex-conditions3570-10yoshizawa-notation3570-7miura-ori3570-25pajarita366+9293paper-japan392+35110one-cut-star394+3797fold-and-cut397+4076froebel408+511the whole record gives 357; nothing here is a resampling, only the fifteen ways of leaving one out
Fig. 3 The statistic recomputed fifteen times, each without one entry: 268 without recreational folding, 280 without the senbazuru, and up to 408 without the Froebel entry. Six entries change it by nothing at all.

Leaving out one entry at a time makes the concentration explicit. Removing recreational folding takes the statistic from 357 to 268; removing the senbazuru takes it to 280. No other removal moves it by more than ten years downward.

Six of the fifteen move it by nothing whatever, because their gaps are negative — those claims are dated later than their earliest evidence and never entered a mean of positive gaps.

And five entries raise it when removed, by as much as fifty-one years. Those are the claims dated close to their own evidence: the Froebel entry with a gap of one year, the paper-in-Japan entry with a hundred and ten, the two cutting entries with seventy-six and ninety-seven. They are in the mean, they are small, and they are dragging it down. A record with more well-attested claims in it would report a smaller overrun, and not because the overrun had changed.

A mean that cannot report accuracy

There is something odd about the statistic itself, and the resampling makes it visible rather than causing it.

The quantity averages the gaps that are positive. Seven of the fifteen entries have negative gaps — claims popularly dated later than their earliest surviving evidence, which is the ordinary situation for a twentieth-century result whose publication is known and whose popular date lags it — and those seven are excluded by the definition.

So the statistic is a mean over a subset selected for being above zero, and a mean of positive numbers is positive. It cannot come out small however accurate the field becomes. Adding a hundred perfectly dated claims to the record would add a hundred entries with gaps near zero, half of them just above and half just below; the half just above would join the mean and pull it down a little, and the half just below would be discarded. The statistic would fall slowly and would never approach zero, because zero is its floor by construction.

That is a selection effect rather than an error, and the repair is to say what is being averaged. Among the claims dated ahead of their evidence, the overrun averages three and a half centuries is a true sentence. A popular date in this field runs three and a half centuries ahead of its evidence is not, and the difference is the seven entries the definition throws away.

When two of them are first attested togetherEach pair of lineages, with the earliest year at which both of them are attested by something that survives — which is simply the later of their two sources. The individual dates are argued about by centuries; this one is not, because it inherits the better-attested half of each pair rather than the worse. The whole set is jointly attested only from the last of them, and that date is the one the field's story about a merge is actually anchored to.the year the record can put two lineages in the same worldthe later of two surviving sources, which is what a joint claim rests onPaper is made in Chinawith paper reaches japan720Paper is made in Chinawith paper is folded for amusement 1680Paper reaches Japanwith paper is folded for amusement 1680Paper is made in Chinawith the thousand cranes1797Paper reaches Japanwith the thousand cranes1797Paper is folded for amusement with the thousand cranes1797Paper is made in Chinawith a five-pointed star from one s1873Paper reaches Japanwith a five-pointed star from one s1873Paper is folded for amusement with a five-pointed star from one s1873The thousand craneswith a five-pointed star from one s1873all 5 are jointly attested only from 1873 — every earlier joint claim is a claim about at most 4 of themeach row is the later of two surviving sources · A five-pointed star from one straight cut at 1873 is what the whole set waits for
Fig. 4 Five claims with the year at which each pair of them is jointly attested. A pairwise date is a maximum of two, so it inherits the better-attested half of each pair — a construction that behaves quite differently from a mean under resampling.

The two entries, named

Since two documents are carrying the statistic, it is worth saying which and why they are the way they are.

Recreational folding is dated popularly to about 700 and its earliest surviving source is 1680 — a gap of nine hundred and eighty years. The senbazuru is dated to about 900 with a source of 1797, a gap of eight hundred and ninety-seven. Both rest on exactly one surviving source, and both are practices rather than results: things people did, which leave documents only when somebody decides to write them down.

Two traditions and a merge is where the shape of that comes from. Ceremonial wrapping, recreational folding and the kindergarten syllabus are separate lineages, and the popular dates for the first two are inherited from a story that treats them as one continuous tradition with the oldest date in any of them attached to all three. The paper had to arrive first prices the one genuinely common ancestor and finds it a much later constraint than the popular dates assume.

So the two large gaps are not measurement noise. They are the two places where a popular date was set by a narrative rather than by a document, and averaging them with six ordinary entries produces a number that describes neither.

What the resampling is asking

A resampling is not a substitute for knowing the record, and it is worth saying what question it answers.

It treats the fifteen entries as a sample from a larger population of claims that might have been recorded, and asks how much the statistic would move if a different fifteen had been. That is an artificial question here: the record is not a random sample of anything, and there is no population of claims from which somebody drew these.

What makes it worth asking anyway is that the alternative is worse. The alternative is quoting 357 with no interval at all, which claims a precision the arithmetic cannot support from eight numbers however they were chosen. A resampling gives a lower bound on the uncertainty: even if the fifteen entries were a perfect sample, the statistic would still be this unstable, so the true uncertainty is at least this wide and probably wider.

What is repeated, against what survivesEach bar runs from the date a claim is generally given to the year of the oldest surviving source that attests it. Nearly every bar points forward, which means the claim is older in the telling than in the record; the two that point backwards are the cases where the practice was published long before anybody proved it.Paper reaches JapanFolded paper is used ceremonially in Japan400 yrPaper is folded for amusement in Japan980 yrThe thousand cranes897 yrThe pajarita is folded in Spain293 yrPaper folding is taught as geometryA five-pointed star from one straight cutAny straight-line drawing, from one straight cutyear of the source500100015002000the date generally giventhe oldest source that says somedian overrun 201.5 years
Fig. 5 The eight claims dated earlier than their own earliest surviving source, which are the eight the statistic is a mean of. Two of the bars are more than twice as long as any of the others.

Looking at the eight bars directly is the plainest version of the argument. Two of them run to nine centuries and the other six average a hundred and sixty-two years. A single number summarising that set is not summarising anything, and both halves of it are interesting for different reasons.

The kindergarten entry, which pulls the other way

One entry deserves separate mention because it is the extreme case of the pull-down effect and it is the field’s best-documented claim.

The Froebel entry has a gap of one year: popularly dated 1837, earliest surviving source 1838, four independent sources. Removing it raises the statistic by fifty-one years — more than any other single removal in either direction. The best-attested entry in the record is the one doing most to make the field look accurate.

The kindergarten was a geometry class follows that lineage forward, and its documentary position is no accident: a syllabus invented for schools in nineteenth-century Germany produces printed material immediately and in quantity. A practice that was passed hand to hand for centuries does not.

So the record has two populations in it as far as this statistic is concerned — claims about institutions, which are documented from the moment they exist, and claims about practices, which are documented when somebody happens to write them down — and the statistic averages across the two. The interval is the arithmetic noticing.

What a record with an interval would say instead

If the field’s figure has to be quoted, the interval says how. Three statements are available and they are not equivalent.

“A popular date typically runs about a year ahead of its evidence” is the median over the whole record, interval −10 to 110. It is the statement with the least in it and it is nearly exact.

“Among the claims dated ahead of their evidence, the overrun averages three and a half centuries, and the average is dominated by two entries” is the mean with its provenance attached. It is what one lost source and the story changes established and it is the honest use of the number.

“Two claims in this field are dated nine centuries before anything surviving” is the two entries themselves, with no averaging at all. It is the strongest statement the record supports and it does not need a statistic.

The third is the one that survives every test here. A statistic whose interval is wider than itself is usually a statistic doing worse than the data it is made of, and the data in this case is short enough to be read.

One lost source and the story changesEvery claim in the record with the number of independent surviving sources behind it, and what a single loss would leave. Several rest on one witness, so one fire would take them out of the record entirely — and the field's most-quoted statistic, how far ahead of its evidence a popular date runs, falls by nearly half when those rows are set aside — while the median hardly moves, because the effect is not a tendency in the record but two particular documents. The surviving sources are not thereby wrong; what is at stake is that the shape of the evidence is partly an accident of what burned.how many witnesses each claim has4 of 8 have one, and a loss removes them from the record rather than weakening themclaimsurviving sourcesafter one lossPaper reaches Japan1nothing attests itFolded paper is used ceremonially in Japan43 leftPaper is folded for amusement in Japan1nothing attests itThe thousand cranes1nothing attests itThe pajarita is folded in Spain21 leftPaper folding is taught as geometry43 leftA five-pointed star from one straight cut1nothing attests itAny straight-line drawing, from one straight cut21 leftby kind of source: artefact 1 · manuscript 2 (1 single) · printed 4 (2 single) · secondary 1 (1 single)the average overrun falls from 357 years to 193 without them, and the median from 201.5 to 184.5 — the effect is two rows, not a tendencywitnesses counted from the record itself · Paper is folded for amusement in Japan and The thousand cranes are the two largest overruns and have one document each
Fig. 6 The same eight claims with their surviving source counts, showing which of them rest on a single document. The two largest gaps are both single-witness entries.

Two intervals, read together

Putting the mean’s interval and the median’s beside each other gives a reading of the record that neither gives alone.

The median says the typical claim in this field is dated within a decade or so of its earliest evidence, and says it firmly: the interval is −10 to 110, and a good half of the entries sit within a few years of zero. By that measure the field is in reasonable order.

The mean over the ahead-claims says that the exceptions are enormous, and says it loosely: somewhere between a century and a half and six centuries, with two entries carrying it. By that measure the field is badly out.

Both are true and they are about different entries. A field described by its median is a field of well-dated results with a handful of legends attached; a field described by its mean of overruns is a field of legends with some results attached. The record supports the first description and the second is what gets quoted.

That is the practical value of computing an interval at all. It does not make the number better; it makes visible which of two available stories the record is actually telling, and the width of each interval says which story is safe to tell.

What the resampling cannot show

Every limitation here is about what a resampling is, and the first is the largest.

It cannot correct for the record being a record rather than a sample. The entries were chosen because they are the claims this subject argues about, so a resampling asks what would happen if the subject argued about a different fifteen of the same kind — which is a question with no clear meaning. The interval is an estimate of the estimator’s own instability, not of historical uncertainty.

It cannot see uncertainty in a date. Every entry carries one earliest year as though it were exact, and a source dated “the late sixteenth century” enters as a number. Whatever that costs the statistic is invisible to a resampling of the entries, since the resampling moves entries around and never moves a date.

And it cannot detect a missing entry. If there are claims the record does not hold, no amount of resampling the ones it does hold will find them. That is a different question with a different method and it is not this one.

What the model assumes

The gaps are as the record states them, with the popular date and the earliest surviving source each one number. A record is not a proof is where the limits of that are set out and none of them is repaired here.

A resample is fifteen draws with replacement, which is the ordinary bootstrap and treats the entries as exchangeable. They are not — an entry about paper and an entry about a book are different kinds of thing — but any weighting would be an assumption about the field rather than a fact about it.

Twenty thousand draws is enough that the quantiles are stable to a year or two, and the seed is fixed, so the same interval comes back every time this is computed.

And the statistic is the mean over positive gaps, which is the field’s own definition. A mean over all fifteen would be a different number with a different interval and would answer a different question.

How the numbers were checked

The interval is required to be wider than the statistic, which is the essay’s title and is checked on the computed quantiles rather than declared. A record whose statistic was well determined would fail that check and the figure would say so.

The median’s interval is required to be narrower than two hundred years, so the contrast between the two summaries is verified rather than described.

The leave-one-out is required to swing by more than a hundred years across the fifteen, and to contain at least one entry that raises the statistic when removed. The second is the counter-intuitive half and it would be easy to state wrongly.

And the resampling is seeded, so every number quoted here is reproducible rather than a draw somebody happened to get.

Still open: what the source counts could be asked

The record carries one column this essay has not used, and it is the one that would answer a different question entirely.

Every entry states how many independent sources survive for it — one, two, three or four — and the distribution of that column has never been read. Five entries rest on one source, five on two, three on three and two on four. Those counts are the raw material for a question the field cannot otherwise approach: how many claims left no surviving source at all, and therefore how much of this subject’s past is not in the record in any form.

There is a standard estimator for exactly that shape of data, it reads the low end of the distribution, and applying it to fifteen entries is either informative or a demonstration of why it is not. Which of those it is can be settled by computing it and looking at the interval, which is the same move this essay made on the headline statistic and got a useful answer from.

Sideways from here, the interval belongs beside the joint date. When two of them are first attested together finds a quantity that gets firmer as its inputs get shakier, because a maximum inherits its best member. That is the one statistic in this neighbourhood that a resampling should treat kindly, and nobody has checked whether it does.

The habit worth carrying is a question to ask of any summary of a short record. Recompute it without each entry in turn, and resample it, before quoting a third digit. The first tells which entries it is made of and the second tells how much it is worth, and between them they usually say that the entries are more interesting than the summary.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AttestationDocumentary recordExpected-valueIdentifiabilityIndependent discoveryPrimary source