Where Christ accumulates
The whole question in one moving picture. Every verse of the Latter-day Saint standard works runs along the bottom in canonical order; the height is the running count of verses that refer to Jesus Christ. Slope is the local rate — a 45° climb means nearly every verse, a flat stretch means none. Press play, drag anywhere on the plot to scrub, and switch the definition to watch a canon's measured Christ content change by a factor of eleven.
Running count
Standalone versions: five volumes · Gospels apart · static PDF · built by scripts/45_cumulative_series.R
What the corpus shows
Six findings, heaviest first. Each links to the analysis that produced it; the identifiers in italics point to the families below.
- The measured Christ content of a canon moves by a factor of eleven depending on the rule. The Doctrine and Covenants is 89.1% Christ-mentioning under the broad rule and 7.7% under the explicit one, because almost all of it is Christ speaking in first person rather than being named. The Book of Mormon moves 49.4 → 10.7, the Gospels 71.0 → 21.8, the Old Testament barely at all (0.76 → 0.73). No single number answers the question this project asks; the honest answer is a pair of numbers and the rule that separates them. B1, B3, E1
- The New Testament is two objects, and the difference is discourse. Split at the Gospels, the halves come apart: 71.0% against 35.0%. But they name Jesus at almost the same rate — 24.7% of Gospel verses and 22.0% of the rest. The entire gap is that he speaks in the Gospels and does not in the epistles: 47.4% of Gospel verses sit inside a span of his speech, against 0.7% after them. Any figure quoting one New Testament number is averaging a narrative and a correspondence. C1, E1
- A translator's parentheses can double a canon's measured Jesus content. Hilali & Khan's Qur'an codes at 2.73% against a cluster of 0.91–1.09% for nine other English translations. Strip parenthetical glosses and 8% of their hits survive, against 66–79% for the rest. The excess is translator commentary printed inline — “'Iesa (Jesus)” — not Qur'anic content. C2
- Three fifths of the Hebrew Bible's Christ signal is supplied by the New Testament. 101 of 162 explicit hits in the JPS Tanakh are coded only because a New Testament writer applies that passage to Jesus. Strip the list and the rate falls from 0.71% to 0.27%. On the Hebrew Bible the measure is legible as a measure of Christian reading practice, not of the text. F1
- Twenty-four English Bibles agree on the aggregate and disagree about the verses. Edition explains 0.10% of the variance in the verse-level indicator, which reads as stability. But of the 5,586 verses any of the 24 full Protestant Bibles codes, only 2,313 are coded by all of them: 58.6% are disputed — 58.3% in the New Testament, 62.5% in the Old. Agreement at the level a study reports coexists with disagreement at the level it measures. C1
- The two canons that match on rate use opposite vocabularies.
The Book of Mormon and the New Testament sit at 49.3% and 52.1% under the broad rule and at
10.7% and 18.9% under the strict one. Fifty-five per cent of the Book of Mormon's explicit hits fire
on the generic
lord_titleagainst 22% in the New Testament, where half fire on the name itself. Equal salience, different Christology. B2, D1
Six things the evidence would not support. They matter as much as the findings, because they bound what the corpus can be used to argue.
| The claim | What the evidence says | ID |
|---|---|---|
| Christ density rises toward the end of a book | No volume shows a significant positive slope on relative position; the one significant slope is negative, on a base rate under 1% | A2 |
| Measured density has drifted across four centuries | Fitted slope indistinguishable from zero with or without family fixed effects, and with one observation per family | C4 |
| John's rate is driven by the Johannine self-designations | John loses 60.4% of its coded verses to discourse inference and 28.0% to the bare name; the Johannine category accounts for under 2% | D3 |
| The per-verse denominator biases the ranking | Spearman correlation between the per-unit and per-thousand-word orderings is 0.92 | H4 |
| The Book of Mormon's chapter-level shape resembles Restoration scripture more than the New Testament's | Indistinguishable from the New Testament (Kolmogorov–Smirnov p = 0.074); the Doctrine and Covenants is a different object (p = 7×10−34) | A4 |
| Canons differ in vocabulary diversity beyond what length explains | Raw entropy tracks book length; rarefied to a common depth the differences largely disappear | D2 |
Across all twenty-two analyses: 14 claims supported, 7 mixed, 6 not supported, 1 underpowered. Every one is stated and marked in the families below.
Seven false positives, found by writing the worked examples. Building the examples section below meant reading one coded verse of each kind from every book of scripture, and that surfaced coding errors no aggregate would have shown. In the deuterocanon “Jesus” and “Jesu” render Jeshua the high priest or Joshua son of Nun, not Christ — fifteen verses across 1 Esdras, Sirach, 1 Maccabees and 2 Esdras 7:37. “Propitiation” is Levitical in Wisdom and Sirach. Tobit's “only begotten” are literal only children and Sirach's “son of the most High” is the reader who cares for orphans. 1 Enoch 60:10's “thou son of man” addresses Enoch, exactly as in Ezekiel. And in the hadith, “'Isa” is an ordinary Arabic name: nineteen reports were coded on narrators such as 'Isa bin Hafs in the chain of transmission.
All are now excluded, and the effect is concentrated: the Apocrypha/deuterocanon falls from 0.94% to 0.23%, which moves it from just above the Old Testament to clearly below it. An earlier version of this page reported the higher figure.
One correction carried through the analysis. The first version of the B2 decomposition filed every verse containing lord_title as “the Lord alone”, including verses that say “the Lord Jesus Christ” and so name him outright. The audit sample surfaced it — the first four verses drawn were 1 Thessalonians 4:1, 1 Timothy 6:14, Acts 19:5 and James 1:1. Every B2 figure reported here comes from the corrected version.
The corpus in two pictures
Before the question-by-question results, the two exhibits that show the whole shape of the thing. The Bible enters as four baselines, not one. The New Testament is split into the Gospels and everything after them, because the pooled figure describes neither: the Gospels code at 71.0% [69.5–72.4] and the rest of the New Testament at 35.0% [33.6–36.5]. Alongside them sit the Old Testament at 0.76% and the Apocrypha/deuterocanon — scripture for Catholic, Orthodox and Ethiopian Orthodox readers, and neither Old Testament nor New — at 0.23% [0.13–0.39].


ASalience and concentration
Concentration and density are close to the same thing, which limits what concentration adds.
The claimConcentration across chapters carries information about a canon beyond its overall rate.
Gini over chapter-level counts is almost mechanically determined by the rate: the Old Testament books all sit at Gini values of 0.93–0.98 — coefficients, not rates — because nearly every chapter is at zero, and the least concentrated books are simply the most saturated — 1 Peter (0.10), Ephesians (0.12), John (0.18). The one thing that does survive the confound: within the high-rate volumes, the Doctrine and Covenants is the most concentrated, with 36.6% of its references in the top decile of sections against 29.4% for the New Testament and 25.5% for the Book of Mormon.





A1_fig_lorenz.pdf · A1_fig_gini_scatter.pdf · fig_lds_books_ranked.pdf · fig_lds_books_ridge.pdf · fig_lds_chapters_ranked.pdf · A1_tab_concentration.csv
Books do not save Christ for the end.
The claimChrist density rises toward the end of a book.
Binomial models of chapter rate on relative position within book, fitted separately by volume: New Testament +0.12 (p = 0.13), Book of Mormon +0.10 (p = 0.25), Pearl of Great Price −0.16 (p = 0.40). The only significant slope is the Old Testament's, and it is negative (−0.64, p = 0.014) on a base rate under 1%, which is a handful of verses moving a coefficient. There is no compositional regularity here to find.

Jesus in the Qur'an is a Medinan subject, by a factor of eight.
The claimQur'anic Jesus material concentrates in the Medinan surahs.
Pooled across all ten English translations: Medinan surahs 3.49%, Meccan 0.42% — an 8.4-fold difference, robust in every individual translation. Al-Ma'idah, Ali 'Imran, An-Nisa and As-Saf carry most of it. The received account in the literature is that Jesus material belongs to the period of engagement with Christian interlocutors; that account survives measurement.

A3_fig_surahs.pdf · A3_tab_surahs.csv · A3_tab_surah_models.txt
The Book of Mormon matches the New Testament, not Restoration scripture, in chapter-level shape.
The claimThe Book of Mormon's chapter-level distribution resembles Restoration scripture more closely than the New Testament's.
Kolmogorov–Smirnov on chapter-rate distributions: Book of Mormon against New Testament D = 0.116, p = 0.074 — not distinguishable. Against the Doctrine and Covenants, D = 0.690, p = 7×10−34. Wasserstein distances tell the same story: 3.7 points to the New Testament, 41.4 to the D&C, 50.4 to the Old Testament. The two texts that match on the mean also match on the whole distribution.


A4_fig_ecdf.pdf · fig_bom_vs_bible.pdf · A4_tab_shape_tests.csv
BDefinitional sensitivity
Book rankings are robust to most definitional choices and not to the Jehovah one.
The claimBook rankings are robust to the choice among defensible codings.
Kendall rank correlation between book orderings: broad against strict τ = 0.774; broad against explicit-only 0.840; broad against broad-minus-"the Lord" 0.925. But broad against the Latter-day Saint Jehovah variant is τ = 0.699, and against strict it falls to 0.487. Everything except the doctrinal premise leaves the ordering roughly intact; the doctrinal premise reorders it.


B1_fig_spec_curve_volumes.pdf · B1_fig_bump.pdf · B1_tab_rank_stability.txt
“The Lord” is Christ 65% of the time in the New Testament; in Restoration scripture the answer depends on a doctrinal premise.
The claimWhere a verse is coded on the generic title alone, “the Lord” denotes Christ.
160 verses coded on the generic title alone, forty per volume, read in context, with exact binomial intervals throughout: New Testament 65% [48–79]; Book of Mormon 100% [91–100]; Doctrine and Covenants 95% [83–99]; Pearl of Great Price 97.5% [87–100].
The New Testament number is the only one that means what it appears to mean, because Paul routinely contrasts “God our Father” with “the Lord Jesus Christ” and the text itself settles the referent. In Restoration scripture the identification runs the other way: “the Lord” is Jehovah, and Jehovah is the premortal Christ, and the verses almost never distinguish Father from Son on their own. The Restoration estimates therefore depend on the Latter-day Saint Jehovah/Christ premise rather than on the surface wording alone — and that is the finding.
Corrected components, as a share of each volume's verses:
| Volume | Named or titled | “The Lord” alone | Christ speaking | Pronoun |
|---|---|---|---|---|
| New Testament | 18.9% | 4.4% | 16.3% | 12.6% |
| Book of Mormon | 10.7% | 16.5% | 14.9% | 7.3% |
| Doctrine and Covenants | 7.7% | 14.7% | 65.9% | 0.9% |
| Pearl of Great Price | 11.3% | 22.4% | 24.3% | 6.0% |
| Old Testament | 0.7% | — | 0.0% | 0.0% |


B2_fig_components.pdf · B2_fig_audit.pdf · B2_tab_audit_result.csv · B2_audit_lord_title_coded.csv
One premise multiplies the KJV Old Testament by thirty-two.
The claimIdentifying Jehovah with the premortal Christ materially changes the Hebrew Bible's rate.
0.76% under the headline coding, 24.74% once Jehovah is identified with the premortal Christ. This is the KJV Old Testament as the Latter-day Saint edition prints it, not the source-language Hebrew Bible. No word of the text changes; only the interpretive premise does. Book by book the shift is largest in Deuteronomy, Leviticus and the Former Prophets — the books with the densest covenant-name usage and the fewest messianic titles, which is to say the books where the premise is doing all the work.

Fragility ranks the canons differently from density.
The claimCanons differ systematically in how much of their rate survives a stricter rule.
Broad to strict, as a retained share: the New Testament keeps 36% of its rate (52.1 → 18.9), the Book of Mormon 22%, the Pearl of Great Price 18%, and the Doctrine and Covenants 9% (89.1 → 7.7). At the other end the low-rate corpora barely move at all, because what little they have is named outright: the Old Testament keeps 97% of its rate and the Apocrypha/deuterocanon almost all of its 0.23%. The ordering by fragility is close to the ordering by how much of the work is first-person revelation.

CTranslation and edition effects
The aggregate is stable; the verses are not.
The claimMeasured Jesus content is stable across English translations of the same canon.
On a balanced panel of 31,037 verses across 24 full Protestant Bibles, edition accounts for 0.10% of the variance in the broad indicator and 0.03% in the strict one. By that measure translation barely matters — but the pooled figure is an artefact of pooling. Fitted inside the Gospels alone, edition accounts for 2.33%, more than twenty times as much, and the residual rises from 9.2% to 24.3%. In the rest of the New Testament it is 0.13%.
The same compression shows in the spread. Across the 24 editions the whole-Bible rate has a standard deviation of 1.08 points; the Gospels alone have 7.36, ranging from 47.3% to 79.3% on the same 3,779 verses, while the rest of the New Testament has 1.74. The instability is concentrated exactly where discourse inference does the work — which is also the tier the validation found least reliable. Averaging a 0.8% Old Testament, a 71% Gospels and a 35% remainder into one 13% number does not describe any of them.
But of the 5,586 verses that at least one edition codes, only 2,313 are coded by all twenty-four. 58.6% are disputed — 58.3% within the New Testament and 62.5% within the Old, so this is not a testament-specific problem. The editions agree on the total and disagree on the membership, which is exactly the pattern that makes a content measure look more reliable than it is: any study reporting a single translation's verse list is reporting one draw from a distribution it never sees.

C1_fig_editions.pdf · C1_tab_variance_partition.csv · C1_tab_disagreement.csv
Ninety-two per cent of one Qur'an translation's Jesus content is its own parentheses.
The claimDifferences between Qur'an translations reflect translator apparatus rather than the text.
Hilali & Khan code at 2.73% against a cluster of 0.91–1.09% for nine other English translations. Strip material inside parentheses and brackets and 8% of their 141 explicit hits survive. “Surviving” here means the hit is still present after all parenthetical and bracketed material is removed. The comparable figures are 79% (Khattab), 76% (Rodwell), 73% (Sale, Abdel Haleem), 71% (Yusuf Ali, Pickthall), 66% (Asad). Their translation prints “'Iesa (Jesus)” and “(Jews and Christians)” inline as commentary.
Shakir is the case to be careful with: 24% survival on a rate no different from the cluster, which suggests grammatical supplement rather than commentary and needs the hand classification listed as the robustness check.

C2_fig_deglossing.pdf · tab_quran_editions.csv · C2_tab_deglossing_all.csv
Sacred-name editions encode translator judgments about “the Lord”.
The claimSacred-name editions change measured Christology, not merely orthography.
The New Heart English Bible codes at 13.48%; its Jehovah Edition, identical but for the divine name, at 13.01% — 148 verses lost, 90 of them on the string “the Lord”. The Restored Name KJV loses 1,056 verses against the King James, printing יהוה where the KJV prints “the Lord”.
Split by testament, the losses are almost entirely New Testament: 1,052 of the RNKJV's 1,056 and all 678 of the Messianic Edition's, against 4 and 0 in the Old. That is the opposite of what the substitution's name suggests, and it is the informative part. These committees are not merely restoring the covenant name where the Hebrew has it; they are reassigning roughly a quarter of the New Testament's “Lord” verses away from Jesus, in a corpus where the underlying Greek reads kyrios. On the New Testament alone the RNKJV drops from 52.11% to 39.40%.
This is not a lexicon bug. These editions encode a translator's judgment, verse by verse, about whether the underlying word is the covenant name or the title of Jesus — which is the judgment the B2 audit had to make by hand. They are a ready-made instrument for that question and should be used as one, and the two sources roughly agree: the audit judged 65% of New Testament “the Lord” verses to denote Christ, and the RNKJV's translators leave about three quarters of them in place.
No drift across four centuries.
The claimMeasured Christ density has drifted across four centuries of English translation.
Twenty-four complete Protestant Bibles, 1599 to 2022, fitted separately by testament. No slope reaches significance in either: Old Testament p = 0.31 raw, 0.72 with family fixed effects, 0.43 one-per-family; New Testament p = 0.11, 0.71 and 0.065. The closest thing to a trend is the last of those — a New Testament slope of 0.025 points a year with one observation per family — and on seven families that is not evidence of anything.
The apparent spread is family, not time: ten of the twenty-four descend from the King James and cluster tightly, while Geneva 1599 (38.7% in the New Testament) and the Berean Standard Bible 2022 (55.8%) bracket them without any monotone trend between.

DLexical theology
Five vocabulary clusters, and they are not the five traditions.
The claimEach canon has a distinctive vocabulary profile, and the Book of Mormon's groups with Restoration scripture.
Correspondence analysis on the categories that triggered each coded verse, 74% of inertia in two dimensions. Ward clustering gives: {Old Testament, Tanakh}, {Qur'an, Bukhari, Muslim}, {Book of Mormon, D&C, New Testament aggregate, Pauline epistles}, {the four Gospels, Pearl of Great Price}, and {Revelation} alone.
The interesting split is inside Christianity. The Gospels separate from the epistles because narrative uses the bare name while epistles use titles — and Restoration scripture lands on the epistolary side. So the claim is only half right: the Book of Mormon groups with an epistolary New Testament and D&C cluster, not uniquely with Restoration scripture. The D&C is nearby, but the strongest comparative anchor is the New Testament, and both nearest neighbours point there: Book of Mormon → New Testament, D&C → New Testament, Revelation → D&C. This is the vocabulary counterpart of A4's distributional result.

D1_fig_correspondence.pdf · D1_tab_clustering.txt · D1_tab_composition.csv
Vocabulary variety is a property of length, until you rarefy it away.
The claimCanons differ in vocabulary diversity beyond what their length explains.
Raw entropy tracks book length almost perfectly. Rarefied to a common 200 coded hits, the spread narrows sharply and the surviving differences are small. The measure is worth reporting only in rarefied form, and even then it discriminates less than composition does — which vocabulary a canon uses separates canons; how many kinds it uses barely does.

John is not distinctive for the Johannine self-designations.
The claimJohn's high rate is driven by the Johannine self-designations.
Leave-one-category-out attribution. John would lose 60.4% of its coded verses to the removal of discourse inference and 28.0% to the bare name. The johannine category — the light, the bread, the vine, the way — does not appear in its top four and accounts for under 2%. The received explanation for the highest-rate book in the Bible is not supported by this measure; John is high because Jesus talks for most of it.
Revelation is the case where the received explanation holds: lamb_title is its second-largest contributor at 13.6%. The D&C loses 74.9% to discourse inference, the Book of Mormon 44.8% to discourse and 33.4% to "the Lord".

Book-level genre coding supports the expected title pattern; the chapter-level test is still open.
The claimWhich title a text uses varies with its literary genre.
Standardised residuals from independence over the King James Bible: apocalyptic writing over-uses lamb_title and apocalyptic; the gospels over-use the bare name and son_of_man; the epistles over-use lord_title and sonship; prophecy over-uses holy_one and the NT-applied passages. The pattern is what a reader would predict, which is mild evidence the categories are cutting at real joints.
Genre here is assigned at book level, so Daniel counts entirely as apocalyptic and Zechariah entirely as prophecy. Chapter-level coding, blind to the outcome and with a two-coder reliability sample, remains the version worth doing.

ENaming versus discourse
Three corners, and every corpus sits somewhere real on the triangle.
The claimCanons differ in whether Jesus appears as a named object, a speaking subject, or a pronoun.
The Tanakh, the Old Testament, the Qur'an and Sahih al-Bukhari sit almost exactly on the named outright corner: in those corpora coded references are lexically explicit rather than inferred from discourse. The Doctrine and Covenants sits on the Christ speaking corner. The Gospels spread across the middle, Mark closest to speech and Matthew to naming. Acts, Romans, Hebrews and 1 Enoch run up the pronoun edge — books that talk about him in the third person without naming him often.

The Doctrine and Covenants swings 33 points on one parameter.
The claimThe Doctrine and Covenants rate is sensitive to the speech-span parameter.
Sweeping the speech-span decay from 5 verses to 60, re-running the whole classifier at each value: D&C 62.8% → 96.1%. The other volumes barely move — Book of Mormon 41.6 → 53.6, New Testament 45.9 → 54.1, Old Testament flat at 0.8. There is no plateau in the D&C curve, which means no value of the parameter is empirically privileged and the published 89.1% is a choice rather than an estimate.
Any paper using the D&C figure has to report this curve. The hand-coding fix — 150 D&C verses coded for "is Christ the speaker here", choosing the window that maximises agreement — is the way to turn the choice into an estimate, and it has not been done.

The error signature is clear; the sample is too small to model it.
The claimContextual errors are predicted by competing-referent density and distance from the antecedent.
Contextual precision is 78.8% [61.1–91.0] against 100% [89.4–100] for explicit hits, and every error but one in the validation sample was a false positive from a pronoun chain reaching one verse too far. But 33 contextual hits cannot support a model of correctness on referent density and distance. This needs the 500-verse sample listed in the agenda, and it has not been drawn.
FInterpretive frames
Three fifths of the Hebrew Bible's Christ signal is the New Testament reading backwards.
The claimMuch of the Hebrew Bible's Christ signal comes from passages the New Testament applies to Jesus.
Of the coded verses in the JPS Tanakh, 61.6% are coded only because the New Testament applies that passage to Jesus; in the KJV Old Testament, 57.7%. Strip the list and the rate falls from 0.71% to 0.27% (Tanakh) and 0.76% to 0.32% (Old Testament).
The appropriation flows are concentrated: Isaiah supplies the most claimed verses, overwhelmingly through the fourth Servant Song, and Matthew and Hebrews do the most claiming. The measure, on the Hebrew Bible, is legible as a measure of Christian reading practice rather than of the text — and that is a more interesting thing to have measured.

F1_fig_appropriation.pdf · F1_tab_provenance.csv · F1_tab_strip_override.txt
The frame matters, at a tenth of the magnitude first reported.
The claimThe interpretive frame materially changes measured rates outside the Christian canon.
Every corpus re-classified under three regimes. This is a counterfactual stress test, not a competing estimate: the no-frame numbers are what the coder would say if the rule were removed, and they are wrong on their face. Removing frames entirely: Qur'an 1.01 → 2.20%, Guru Granth Sahib 0.002 → 0.046%, Bhagavad Gita 0 → 0.14%, everything else unchanged. Tightening to name-only barely moves anything, because the published frame already excludes almost everything a name-only rule would.
The 27 Guru Granth Sahib verses the frame suppresses are exactly the ones it should: "He alone is our Savior", "Redeemer of sinners", "the true light" — Waheguru in every case. The Qur'an's larger movement is mostly downstream contextual propagation from five savior hits, so it is a false-positive cascade rather than a vocabulary question.

GCross-tradition Jesus
Islamic scripture uses the matronymic; Christian scripture almost never does.
The claimIslamic and Christian scripture name Jesus with systematically different vocabulary.
Restricted to the categories admissible under every frame, so the comparison is like for like. Both corpora lead with the bare name — the Qur'an 47%, the Gospels 90% — so the difference is not that one names him and the other titles him. It is what comes second. The Qur'an's vocabulary is 30.6% matronymic, “son of Mary”; the Gospels are 0.1%. The mirror image also holds: “of Nazareth” is 2.2% of the Gospels' vocabulary and 0.0% of the Qur'an's. One tradition identifies him by his mother, the other by his town. The pattern holds across all ten English translations, so it is the Qur'an's and not a translator's.
An earlier version of this analysis reported the headline as “Christian scripture names him by his office”. That was an artefact of a single nazarene category conflating the birthplace epithet with the matronymic. The lexicon now separates of_nazareth from son_of_mary, the whole corpus was reclassified, and the claim is narrower and correct: the contrast is matronymic usage, not office versus name.

Much of the hadith “Jesus” signal is about Christians, not about Jesus.
The claimA large share of the hadith signal is reference to Christians rather than to Jesus.
Share of each collection's coded reports whose only ground is the word “Christian” or “Nasara”: Muwatta Malik 76%, Jami' at-Tirmidhi 63%, Sunan an-Nasa'i 63%, Sunan Abu Dawud 47%, Sunan Ibn Majah 46%, Sahih al-Bukhari 46%, Sahih Muslim 30%. Excluding the category roughly halves most collections' rates — Bukhari 1.65 → 0.96%, Malik 1.38 → 0.33%. Hadith Qudsi is left out of this comparison: 40 reports and one coded hit will not support a share.
Whether a reference to a religious community counts as a reference to its founder is a genuine analytic choice. It should be stated, not defaulted.

The Son of Man title lives entirely in the Book of Parables; most other coded material does too.
The claim1 Enoch's Son of Man material concentrates in the Book of Parables.
Two claims, and only the narrower one takes “entirely”. All 14 hits on the son_of_man category fall in the Book of Parables, chapters 37–71 — the section whose relation to the Gospel title is the whole scholarly question. Of the book's 39 coded Christ-figure verses more broadly, 33 sit there; the remaining six are scattered across the Watchers, the Dreams and the Epistle.
By section, as a share of verses: Book of the Watchers 0.9%, Book of Parables 12.4%, Astronomical Book 0.0%, Book of Dreams 1.3%, Epistle of Enoch 1.0%. The measure localises the title without being told where to look.

The negative controls hold: nothing missed in 280 verses.
The claimThe non-Abrahamic canons are true negatives rather than coverage failures.
280 text units drawn at random from the six corpora that code at or near zero — the Bhagavad Gita, Dhammapada, Tao Te Ching, Analects, Kojiki, and 80 sampled from the Guru Granth Sahib — screened for any English string that could render a reference to Jesus, and the one flagged case read in full. Zero missed references, giving a 95% one-sided exact upper bound on the false-negative rate of 1.31% across the six.
The single flagged verse is Guru Granth Sahib, Ang 1216: "What can any poor mortal do to someone who has the Lord as his Savior and Protector?" — Waheguru, correctly not coded, and a clean illustration of what the frame is for.
HValidation
Exact when it sees a name, fallible when it infers.
The claimExplicit and contextual coding have materially different accuracy.
Explicit tier 100% [89.4–100] on 33 verses; contextual tier 78.8% [61.1–91.0] on 33; verses coded as no reference 97.4% [90.9–99.7] on 77. A pooled precision of 88% averages two instruments with different properties; the tiers should be reported separately.

A pooled coder-error stress test widens every interval past most of the rankings.
The claimPropagating coder error materially widens the interval around every rate.
Rogan–Gladen with Beta posteriors on sensitivity and specificity from the validation counts. The credible intervals are wide enough that adjacent works in the ranking overlap freely, and the width is driven by the 143-verse validation sample rather than by ambiguity in the texts: a reason to expand the validation, not to distrust the corpus.
Two cautions, because the correction is cruder than it looks. It applies a single pooled sensitivity and specificity, when H1 has just shown the explicit and contextual tiers have very different error profiles — so this is a stress test, not a better estimate. And every low-rate corpus whose observed rate falls below the pooled false-positive rate is clamped to zero: the Old Testament, the Tanakh, the Apocrypha at 0.23%, most hadith collections. That is an artefact of the estimator, not evidence that those corpora contain no references to Jesus.

The denominator does not change the answer.
The claimThe per-verse denominator biases the ranking against canons with long verses.
Spearman correlation between the per-verse and per-thousand-word orderings is 0.23. Verse length varies a great deal — hadith reports run many times longer than Proverbs — but not in a way that is correlated with Christ density, so the per-verse measure is not systematically favouring terse books. One fewer thing to worry about.

The three ways a verse gets coded
Every coded verse qualifies on one of three grounds. Below is one worked example of each, from each book of scripture in the corpus, with the matched string marked. Where a work has no instance of a mechanism that is stated rather than filled in — the Old Testament never has Christ as the speaker under this coding, and that absence is part of the result.
The Gospels (KJV)Protestant; Catholic; Orthodox; Ethiopian Orthodox; Latter-day Saint
New Testament beyond the Gospels (KJV)Protestant; Catholic; Orthodox; Ethiopian Orthodox; Latter-day Saint
Old Testament (KJV)Judaism; Protestant; Catholic; Orthodox; Ethiopian Orthodox; Latter-day Saint
Apocrypha / Deuterocanon (KJVA)Catholic; Orthodox; Ethiopian Orthodox
Book of MormonLatter-day Saint; Community of Christ
Doctrine and CovenantsLatter-day Saint
Pearl of Great PriceLatter-day Saint
1 EnochEthiopian Orthodox
Tanakh (JPS 1917)Judaism
MishnahRabbinic Judaism
Qur'an (Yusuf Ali)Islam
Sahih al-BukhariSunni Islam
Sahih MuslimSunni Islam
Sri Guru Granth SahibSikhism
Whose canon
Twenty-two works, and no two communities receive the same set. Every book-level table in output/analysis/ already carries a canons column, so in most cases no join is needed. output/00_CANON_MEMBERSHIP.csv maps all 938 books in the corpus to the communities that hold them as scripture.
| Books | Received as scripture by | n |
|---|---|---|
| Genesis – Malachi (the Hebrew Bible) | Judaism, Protestant, Catholic, Orthodox, Ethiopian Orthodox, Latter-day Saint | 39 |
| Matthew – Revelation | Protestant, Catholic, Orthodox, Ethiopian Orthodox, Latter-day Saint | 27 |
| Tobit, Judith, Wisdom, Sirach, Baruch, 1–2 Maccabees, and the additions to Esther and Daniel | Catholic, Orthodox, Ethiopian Orthodox — not Protestant or Jewish. Coded rate 0.23%, below the Old Testament's | 12 |
| 1 Esdras, Prayer of Manasses, Psalm 151, 3 Maccabees | Orthodox, Ethiopian Orthodox | 4 |
| 2 Esdras | Slavonic Orthodox only | 1 |
| 1 Nephi – Moroni (the Book of Mormon) | Latter-day Saint, Community of Christ | 15 |
| Doctrine and Covenants; Moses, Abraham, Joseph Smith—Matthew, Joseph Smith—History, the Articles of Faith | Latter-day Saint | 6 |
| 1 Enoch | Ethiopian Orthodox Tewahedo | 1 |
| The Qur'an | Islam — the same text for Sunni and Shia | 114 |
| Hadith Qudsi | Sacred hadith: religiously authoritative reports, not Qur'an-level scripture. Their authority and canonicity differ by community | 1 |
| Bukhari, Muslim, Abu Dawud, Tirmidhi, Nasa'i, Ibn Majah; Muwatta Malik; Nawawi's Forty | Sunni Islam — the Shia Four Books are not in this corpus | 381 |
| The Mishnah | Rabbinic Judaism | 62 |
| Bhagavad Gita | Hinduism, as smriti rather than shruti | 18 |
| Dhammapada | Theravada Buddhism, in the Khuddaka Nikaya | 26 |
| Sri Guru Granth Sahib | Sikhism | 1 |
| Tao Te Ching · Analects · Kojiki | Taoism · Confucianism · Shinto | 88 |
Where the lists genuinely differ, they are marked rather than smoothed. “Orthodox” here means the Greek reception; the Slavonic tradition additionally receives 2 Esdras, and 4 Maccabees sits in a Greek appendix rather than the canon proper. Catholic Bibles print the Letter of Jeremiah as Baruch 6 and the additions to Esther and Daniel inside those books rather than separately. The Ethiopian Orthodox Tewahedo canon is broader than any other and its enumeration is contested in the scholarly literature, so only the books actually present in this corpus are marked for it. Community of Christ receives the Book of Mormon and a Doctrine and Covenants, but its D&C differs in content from the Latter-day Saint one, so it is marked on the Book of Mormon alone.
Still open
What the corpus cannot yet answer, stated plainly. Two analyses were not attempted, two were run in a weaker form than intended, and two carry robustness work that should be done before any of this is published.
| Question | Why not | What it needs |
|---|---|---|
| Weaker form. D4 — titles by genre | Run at book level, not chapter level | ~1,600 chapters coded blind, two coders on an overlap |
| Weaker form. E3 — where anaphora breaks | 33 contextual hits is too few to model | 500-verse validation sample weighted to contextual hits |
| Not attempted. F3 — estimate the frame | No embeddings in the pipeline | Corpus-specific embeddings; canons above 5,000 verses only |
| Not attempted. H3 — LLM as a second coder | No three-way agreement study run | 2,000-verse three-way agreement study |
| Robustness. C2 | Shakir's 24% survival is unexplained | 50 stripped hits per translation classified supplement vs. commentary |
| Done since. G1 | nazarene conflated birthplace with matronymic | Split into of_nazareth and son_of_mary; the corpus was reclassified and G1 rewritten |
What this does to the first paper. The design holds, and two of its three planks got stronger. The de-glossing result is decisive. The frame result is smaller in magnitude than the design assumed and should be reported at its true size, with the Guru Granth Sahib verses quoted rather than counted. The scope result is unchanged.
The change is that the paper now has a fourth plank and it is the best one: twenty-four translations agree on the aggregate and disagree about 58.6% of the individual verses. That is the cleanest available demonstration that agreement at the level a study reports can coexist with disagreement at the level the study actually measures.