Some discoveries would make it worse
Assumes The interval is wider than the number and A question the record is too small to answer.
One lost source and the story changes asks what the record would say if a document were taken away, and finds two entries carrying the field’s most-quoted number. The question it leaves is the one a historian actually faces, which is not what happens when a document is lost but what happens when one is found.
That question can be asked without inventing a single date. For each claim, suppose a source turned up two centuries earlier than the earliest now surviving, and recompute. The answers are not the ones the question invites.
Five discoveries that make the field look worse
The statistic is the mean of the gaps that are positive — the claims dated earlier than anything surviving for them. A discovery moves a claim’s earliest source back, which enlarges that claim’s gap if it was already ahead, and which can push a claim out of the mean entirely if its evidence overtakes its popular date.
A claim leaving the mean takes its own number with it, and if that number was small the mean rises. The kindergarten entry has a gap of one year; a source two centuries earlier gives it a gap of minus a hundred and ninety-nine, it drops out, and the mean of the remaining seven is 408 against 357. The fold-and-cut entry adds 40, the one-cut star 37, paper in Japan 35, and the pajarita 9.
So five of the fifteen possible discoveries would make the field’s summary of its own accuracy worse. None of them would make the field less accurate. In every one of those five cases the discovery brings the evidence closer to the claim, which is unambiguously a gain in what is known and unambiguously a loss in what the statistic reports.
That is a property of the statistic and not of history, and it is the clearest demonstration available that the number is measuring the wrong thing. A summary that gets worse when the evidence improves is a summary whose definition is doing something other than what it says.
Four discoveries that help, and by twenty-five years each
The four claims with very large gaps behave as expected and modestly. Ceremonial wrapping, recreational folding, the senbazuru and the pajarita each have gaps large enough that two centuries does not exhaust them, so each one’s gap falls by two hundred, and the mean of eight falls by two hundred over eight — twenty-five years, every time.
That arithmetic is worth noticing. The best a single discovery can do to a mean of eight is one eighth of whatever it moves, so a document two centuries early is worth twenty-five years off a figure of 357 — seven per cent. The statistic is almost completely insensitive to any one discovery, which is the same concentration the leave-one-out found, read forwards instead of backwards.
Six discoveries that change nothing
The third behaviour is the quietest and it is worth a paragraph because it is half the record.
Six claims change the statistic by exactly nothing when a document is found for them: paper in China, paper in Europe, Beloch’s paper, the Yoshimura pattern, the vertex conditions, the Miura fold and the notation entry. Every one of them already has a gap of zero or less — the claim is dated at or after its earliest evidence — so it was never in the mean, and pushing its evidence further back cannot bring it in.
That is half the record having no influence on the field’s headline number in either direction, which is a strong statement about what the number describes. It is a statistic about the pre-modern practices and nothing else; every twentieth-century result in the record is invisible to it.
A record is not a proof makes the point that the field’s figures draw the shape of the evidence rather than the evidence. This adds a sharper version: one of those figures draws the shape of eight entries, and which eight is decided by a sign test that six entries fail for the best possible reason.
No single document reaches 250
Putting the two behaviours together gives a question with a sharp answer: what would a newly found source have to be dated for the statistic to come down to 250?
Two claims can do it alone. A source for recreational folding dated 825 or earlier brings the figure to 250, and one for the senbazuru dated 942 or earlier. Both would be extraordinary finds — documents eight or nine centuries older than anything now known for those practices.
No document for any of the other thirteen reaches it. The other thirteen cannot at any date. That is not because their gaps are small; it is because of the exclusion. Push a claim’s evidence back far enough and its gap goes negative, the claim leaves the mean, and the mean of what remains is larger than 250 however early the document is. The statistic has a floor for each claim, and for thirteen of the fifteen that floor is above the target.
A reader may check that against the leverage table: the four claims that help at all move it by twenty-five years each, and 357 minus 25 is 332. So the field’s most-quoted figure cannot be moved appreciably by any single discovery, and can be moved at all only by discoveries about the two entries it is already known to rest on. It is a number about two documents pretending to be a number about a discipline.
Two discoveries together
The rows are one-at-a-time hypotheticals and they do not add, so it is worth asking what the best pair does.
Taking every pair of claims and moving both two centuries earlier, the best pair is ceremonial wrapping together with recreational folding, which brings the statistic to 307 — fifty years off, for two extraordinary documents. That is less than the sum of their individual effects would suggest, and the reason is that the second discovery is a smaller fraction of a mean that has already fallen.
The general shape is that the statistic resists. It is a mean whose two largest terms are nine hundred years, so any repair has to work on those terms, and working on them two hundred years at a time moves an eighth of two hundred each. Getting from 357 to 200 requires either a discovery of a thousand years or a change to the definition, and only one of those is available.
And there is a blunter observation available once the interval is beside the leverage. The best single discovery moves the statistic by 25 years and the statistic’s own ninety per cent interval is 441 years wide. Any one discovery is well inside the noise of the quantity it is supposed to improve, which means no individual find could ever be detected in this number at all.
What would move it
Three things would, and none of them is a discovery.
Changing the definition. A mean over all fifteen gaps rather than over the positive ones is 176 rather than 357, and it responds properly to evidence: every discovery lowers it, none raises it, and no claim is excluded for being accurately dated. It is a worse headline and a better statistic.
Changing the popular dates. The paper had to arrive first prices the one constraint all three lineages share, and it is later than two of the popular dates it would have to precede. The gaps are the distance between a popular date and a document, and the popular dates are the half nobody treats as variable. Two traditions and a merge shows where the two big ones come from — a story that treats three lineages as one — and correcting that story would close both large gaps without a single new document.
Adding entries. The kindergarten was a geometry class follows the one lineage whose documentation is uncomplicated, and it is the sort of entry a larger record would be made of. A record of sixty well-attested claims would put the two outliers among many, and the mean would report the field rather than the outliers. That is the same remedy a question the record is too small to answer arrives at from the other direction, and it is the only one that improves every statistic at once.
Of those three, only the first is available to somebody wanting a better number today, and it costs nothing: the mean over all fifteen gaps is already computable and already in the record. What stops it being used is that 176 is a less arresting figure than 357, which is a reason about writing rather than about evidence.
What a discovery is worth, priced properly
If the statistic responds badly, it is fair to ask what a document should be priced against instead, and the record supports one answer immediately.
The joint attestation date is the year by which every claim in a set is on the record, and it is the later of its members — so a discovery for the latest-attested claim moves it and a discovery for any other moves it by nothing. When two of them are first attested together builds that quantity for pairs and finds it firmer than either half, because a maximum inherits its best member.
Priced that way, the leverage table looks completely different. The claim worth finding a document for is whichever one is currently holding the joint date up, and that is a single identifiable entry rather than a mean’s largest contributor. It also has the property the mean lacks: a discovery can only ever move it earlier, so no find makes the field look worse.
That does not make the joint date a better summary of the field’s accuracy — it is not measuring accuracy at all — but it does make it a better target for effort. A quantity that improves monotonically with evidence is a quantity worth chasing; one that can go either way is not a target at all.
What the exercise cannot show
Each row asks a hypothetical and the hypotheticals are not equally plausible.
It cannot say which discoveries are possible. A source for the senbazuru dated 942 would require a documentary situation nobody has any reason to expect; a source for the Miura fold two centuries earlier is impossible on its face, since the fold was published in 1970. The arithmetic treats all fifteen rows alike and a historian would not.
It cannot handle a discovery that changes a popular date. A newly found document might well move the claim as well as the evidence — a source showing a practice is older than thought raises the popular date too — and the computation holds the popular date fixed, which is the assumption that makes it arithmetic rather than history.
It also treats the record’s inclusion rule as fixed. A question the record is too small to answer shows how much a curated list of disputes differs from a sample, and every row here inherits that.
And it cannot price a discovery’s real value. What a document is worth is what it says, and the only thing measured here is what it does to one summary statistic of a record of fifteen entries. A document that settled who folded what, and when, would be worth a great deal and might move this number by nothing.
The three behaviours, and why they are three
It is worth naming the structure that produces the table, because it is not special to this record.
A mean over a selected subset has three regimes for any perturbation of its inputs. An input deep inside the selection moves the mean in proportion to how much it moves and to one over the count. An input outside the selection moves it not at all. And an input near the boundary can cross it, which changes the count as well as the sum, and the sign of that change depends on whether the departing value was above or below the mean.
All three appear here. Four claims are deep inside — gaps of hundreds of years against a boundary at zero — and each is worth twenty-five. Six are outside and worth nothing. Five are near the boundary and cross it, every one of them from below the mean, so every one of them raises it.
That last fact is not a coincidence either. A claim near the exclusion boundary has a gap near zero, and a gap near zero is necessarily below a mean of 357. So every boundary-crossing in this record must raise the statistic, and a record shaped like this one could not produce a counter-example.
That is what makes the finding general rather than a quirk of fifteen entries. Any mean over the positive part of a distribution will be raised by improvements to its smallest positive members, whatever the field, and the improvements are exactly the cases where something was nearly right to begin with.
What the model assumes
Two centuries is the unit of discovery in the leverage figure, chosen because it is large enough to move several claims across the exclusion boundary and small enough to be conceivable. A different span gives different numbers and the same three behaviours.
The popular date does not move. This is the assumption doing most of the work and it is the least defensible; it is what makes the question answerable from the record alone.
One discovery at a time. The rows are independent hypotheticals, so they do not combine: two discoveries together give a number this table does not contain, and the best pair is not the two best rows.
And the record is otherwise unchanged. No entry is added, none is reclassified, and the source counts play no part in the arithmetic at all.
How the numbers were checked
At least one discovery must raise the statistic, which is the essay’s central and counter-intuitive claim, and the check is on the computed values rather than on the reasoning.
No single discovery may move it by more than a quarter, which is the insensitivity claim, and it is checked against the best row rather than the typical one.
The target figure must be reachable by some claims and not by others. A table where every row said “never” or every row gave a year would be a table with nothing in it, and both are refused.
And every row’s recomputation uses the full record with exactly one entry’s year replaced, so a bug that dropped an entry would change the base figure and be visible in the footer.
Still open: the record’s second date column
The thing most wanted here is the thing one lost source and the story changes asked for and nobody has built: a second source year for every claim.
The record states how many independent sources survive and the year of the earliest. It does not say when the second is. With that column, the leave-one-out becomes a genuine one-at-a-time removal rather than a block removal — take away the earliest source and the claim falls back to its second rather than leaving the record — and the leverage figure becomes an interpolation between two known dates rather than a hypothetical two centuries deep.
Every question asked of the record here would be sharper with that column and none of them needs it to be asked. That is the honest summary: the record as it stands supports an interval, a demonstration that an unseen-class estimate fails, and a leverage table, and it supports none of them with any precision.
Sideways from here, the leverage table is a research priority list read backwards. The two claims that could move the number are the two the field would most like documents for anyway, which is reassuring; the five whose documents would raise it are the ones where a discovery would be a genuine gain and would look like a loss, and knowing that in advance is worth something to whoever has to report it.
The habit worth carrying is about summaries that exclude. Ask what leaves the statistic when the data improves. A definition that drops cases for being accurate will reward inaccuracy, and the reward is invisible until somebody computes what a good outcome would do to the number.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- Nothing here is as old as it sounds attestation · documentary record · primary source
- A file has no paper documentary record · primary source
- Found by people not folding paper documentary record · independent discovery
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AttestationDocumentary recordExpected-valueIdentifiabilityIndependent discoveryPrimary source