Data Quality
This catalogue is large enough that nobody can check it by reading it. So it
checks itself: every figure below is recomputed from the database each time
the site is built, and none of it is maintained by hand.
Recomputed 2026-08-28
Scale
What the catalogue holds.
10,878
Works named
Distinct codes. A code is a work: a tablet, a prayer within one, an excerpt that circulates on its own.
63,807
Texts held
Rows — one work in one language, from one source.
524
Languages
Counting every language in which we hold at least one text.
27
Books catalogued
Collections with a known item order.
Findability
A work nobody can search for is, in practice, unnamed. This is the measure that matters most, and the one hardest to move.
4,424
Works in one language only
Usually an original with no translation yet. Not an error — no language is guaranteed to exist for any work — but it is the ceiling on who can read it.
952
Works in no widely-read language
Reachable only by someone who reads the original.
962
Works given an English first line
Works holding no English text at all, for which the catalogue's own English opening is surfaced in search — so the work can be found and identified even where no translation exists. Counted over whole works only: attaching a tablet's opening to an excerpt of it would misstate what the excerpt is. Overlaps the row above without being a subset of it.
Identity
Every text must carry a name. A missing name is the one state the catalogue does not allow.
0
Texts with no code
Must be zero. A text with no name cannot gather its translations.
6
Identity pending (TMP)
One work, identity not yet established. Constitutive but deferred — an honest state, not a defect.
66
Identity absent (X)
Identity could not be established with the tools used. Reopened whenever better tools arrive.
0
Malformed codes
Codes that do not parse. Should be zero.
Integrity
Every reference should resolve. These are the checks that catch a rename that swept some tables but not all.
0
Orphaned prayer-book entries
Structure rows pointing at a code that holds no text.
0
Orphaned relations
Link-outs pointing nowhere.
0
Texts that are empty
A row that names a work but holds no text for it, so it gathers nothing and gives a reader nothing.
0
Texts with leaked markup
Scraper residue, not markup as such. Stored HTML is normal here — 34,000 rows carry paragraph tags that render through Markdown — so this looks only for a stray <div wrapper or double-escaped markup, neither of which belongs in a text.
Linguistic signals
These are proxies, not proofs. Each correlates with a real class of error, each has honest exceptions, and none should be acted on without reading the rows. They exist because a corpus this size in this many languages cannot be checked by reading it.
11
Texts not in their language's script
A row labelled Persian written in Latin letters, or the reverse. Usually a mislabelled language or a transliteration filed as the original. Only languages with an unambiguous script are tested, and only when one script dominates the text.
28
One code and language, two unrelated texts
The same work claimed twice in one language by texts that share almost no words. Holding several renderings of a prayer in a language is normal and not counted here — a second translation still carries the same names and epithets. Texts under forty words are also left out, because at that length a rubric like <em>to be recited at noon</em> outweighs the prayer itself.
73
One item labelled as two languages
Byte-identical text carrying two different language labels, so one of them is wrong — invisible to a script test when the languages share an alphabet, as with Polish filed as Vietnamese or Slovak as Swedish. Most of the current count is a single fault: 61 paragraphs of the Words of Paradise whose Spanish, French and Portuguese rows hold the English text verbatim, so a reader asking for Spanish is served English. Persian/Arabic pairs are excluded by design — there the label records which section of the source a row came from — as are variant labels of one language, which are two editions rather than a mislabelling.
404
One text under two different works
Identical text filed under two codes. Sometimes a genuine duplicate from two sources, sometimes a work that has been split in two by mistake. Worth reading either way.
How many languages a work reaches
A code gathers every translation of one work under one name — that is what
makes a corpus in 500+ languages navigable at all. This is the shape of
what it gathers.
| Languages | Works |
|---|
| 1 language | 4,424 |
| 2 | 3,205 |
| 3–5 | 1,287 |
| 6–20 | 1,121 |
| more than 20 | 841 |
Coverage by book
Every book here has a known item order taken from the catalogue itself, so
an item we hold no text for is visible rather than simply absent.
Knowing what is missing is the precondition for filling it.
| Book | Items | With text | Coverage |
|---|
| iqan | 287 | 287 |
100% |
| esw | 270 | 270 |
100% |
| gems | 246 | 128 |
52% |
| swab | 230 | 229 |
100% |
| aqdas | 190 | 190 |
100% |
| pm | 183 | 183 |
100% |
| gleanings | 166 | 166 |
100% |
| hidden-words | 156 | 156 |
100% |
| pup | 140 | 133 |
95% |
| saq | 84 | 84 |
100% |
| light | 77 | 74 |
96% |
| swb | 75 | 67 |
89% |
| ridvan | 73 | 73 |
100% |
| memorials | 70 | 0 |
0% |
| paristalks | 58 | 57 |
98% |
| days | 44 | 44 |
100% |
| gpb | 27 | 27 |
100% |
| apab | 24 | 24 |
100% |
| apbh | 23 | 20 |
87% |
| tablets | 16 | 13 |
81% |
| divineplan | 14 | 5 |
36% |
| summons | 10 | 8 |
80% |
| call | 7 | 5 |
71% |
| tabernacle | 5 | 2 |
40% |
| wt | 3 | 0 |
0% |
| sdc | 1 | 0 |
0% |
| tn | 1 | 1 |
100% |
What cannot be counted
Everything above is structural: a reference resolves or it does not, a
count matches or it does not. Those checks are exact, and they run on
every build. But this is a corpus of language, and its most important
property — is this text the work its code says it is? — cannot be
settled by counting.
Some of it yields to calibration. A translation runs to a fairly
predictable length against its original, so a text far outside that band
is worth a look — though only as a flag, never a verdict: one text
flagged this way turned out to have a rubric inflating its length, and was
perfectly sound. The rest needs reading, and reading across languages
that few people share.
Where such a question cannot be settled, the catalogue says so rather than
guessing: the work keeps a name marked identity pending until
evidence arrives. Some of those have waited a long time. That is the
honest state for a question nobody has yet been able to answer, and it is
why the figures on this page describe what the catalogue can prove about
itself — not everything that is true of it.
What these numbers are not
Several figures that look like faults are the catalogue being honest. A
work held in only one language is normal — no language is guaranteed to
exist for any work, and a code with only translations is a name waiting
for its original rather than a gap in the evidence. A code marked
identity pending is a work that has been brought into being while
the question of which work it is stays open; that is a truthful state, and
preferable to a confident guess. The figures worth acting on are the ones
marked in the lists above.