Method
2 209counted records, of 2 490 kept
Version of August 2026 · 8 source labels · 2 582 records harvested, 2 490 kept, 2 209 counted
Origenality is a bibliographic map of the scholarship on Origen of Alexandria. It gathers records from open catalogues, files them under a controlled vocabulary of themes, works of Origen and approaches, and draws the result so that a reader sees where the indexed field is thick and where it is thin. The current public build merges 2 582 source records into 2 490 work clusters; its manifest preserves both counts and the 92 collapsed duplicates. The subject headings and containers of 1 173 of those records, from seven of the source labels, were read again from their own catalogues on 13 September 2026.
This page says how the corpus is assembled, what is published and under what terms, and what the density figure does and does not measure.
0. Where this comes from
Crouzel gave the field its inventory. This continues the task, not the method.
From 1971 to 1996 Henri Crouzel published the Bibliographie critique d'Origène in three volumes: the first in 1971, a supplement in 1982 reaching back to 1489, and a second in 1996 covering as far as 1992. He compiled it by hand, over three decades, and he annotated it: each entry carried his judgement of what the study was worth. It remains the instrument the field works from, and the closing round table of Origeniana XIV (Boston, August 2026) named it as the thing that needs bringing up to date.
Its living continuation exists: the annual Origen bibliography published in Adamantius, the journal of the Italian research group GIROTA, since 1995. That bibliography is one of the sources this project federates.
What is claimed here is a continuity of purpose, and it is worth being exact about the limit. Crouzel judged; this map counts. It says how many studies sit in a neighbourhood of the field, never whether they are good, and never whether a project is original. A density is not a verdict, and no machine reading of titles and abstracts is a substitute for reading the studies. What is taken up is the older task of giving the field an inventory it can work from, with the means of 2026 and with the honesty of saying that the counting stops where the judging would begin.
The retrospective span, 1489 to 1992, is Crouzel's and is not yet here. The three volumes are not digitised in a form this project can read; until they are, every figure on this site is a figure about the catalogues named in section 1, not about the literature as a whole.
1. The corpus
Each source label is named and linked to its record system.
The corpus is built by asking catalogues about a person, not about a word. The authority record for Origen (GND 118590235, and the equivalents the national libraries hold) is a number a cataloguer attached to a book after reading it. Asking by that number rather than by the string Origen is what keeps the Spanish orígenes, the Italian origini and the French origines out: the name of the Father collides with the commonest word in the Romance languages for beginnings, and no amount of pattern matching separates them reliably.
Eight source labels contribute 2 582 catalogue records about Origen:
- Index Theologicus (IxTheo, Tübingen), hydrated from K10plus: 1 409
- Gnomon Bibliographische Datenbank (Eichstätt), served in MARC by B3Kat: 354
- K10plus, the German union catalogue, beyond the theology slice: 350
- Sudoc (ABES): 212 · B3Kat: 130
- Deutsche Nationalbibliothek: 51 · Library of Congress: 50 · Bibliothèque nationale de France: 26
The work-level merge collapses 92 duplicate records before the graph is built. It preserves all source identifiers and every field conflict. 84 clusters merge several records, and 19 of those join records from more than one source label. The 21 clusters whose source tags disagree are marked for review. Editions and translations of Origen's own works are set aside: they are primary sources, and counting them among the studies would say the field is larger than it is. Every record keeps the catalogue number it was harvested with, and every entry in the Explorer links back to the record it came from. Nothing on this site replaces a catalogue; it points at one.
How each record was classified, and how well
A record enters a figure only once it has been placed against the controlled vocabulary of § 3: whether Origen is its main object, which themes it touches, which of his works. For the records added in August 2026 that placement was made twice, independently, by two passes that could not see each other, and the disagreements were then arbitrated one by one.
The two passes agreed on relevance for 1 149 records of 1 274 (90.2%). Of the 125 disagreements, 118 were one notch apart. Agreement is highest where it matters most for the counts (core, 0.93) and lowest at the boundary between partial and marginal (0.72), which is where a study that devotes a chapter to Origen meets one that devotes a paragraph.
That figure covers relevance only, the axis that decides whether a record is counted. Agreement on themes, works and approaches has not been measured, and nothing here should be read as covering them.
386 records (30%) are flagged for review. The cause is structural and worth stating plainly: 96% of the records added carry no abstract, so the judgement rests on a title, its authors, its language and its subject headings. That is enough for core (12% flagged) and not enough for marginal, where 65% are flagged. Those flags call for a table of contents or the volume itself, not for another pass of the same method.
The primary layer, set aside and served
The same harvest returns the texts themselves: 1 441 editions, translations and manuscript witnesses of Origen's own works, from a Carolingian codex of the homilies on Kings to a translation of last year. None of them is counted in any figure on this site, and that is the point: a density of studies that counted the editions would say the field is larger than it is.
They are published all the same, as their own layer: data/primary-layer.jsonl,
with a summary the command line serves as origenality primary. Ten records predate the eleventh century, 206 the seventeenth. That is not a bibliography
of scholarship; it is the transmission of Origen, and it deserved better than being
discarded as a by-product.
Datings before 1450 are the holding library's, and approximate. One implausible date was set aside rather than guessed at, and 102 records carry no date at all: none was invented for them.
What is not here yet
Additional enrichment harvests exist in the working data: OpenAlex, Crossref, Semantic Scholar, Isidore, theses.fr, Dialnet, the Italian SBN, the Adamantius repertorio, the BIBP of Laval. No record of theirs is on this map, and no figure on this site counts one. They were gathered by keyword, their relevance filtering is not done, and a figure drawn from an unsorted pile would be worse than no figure.
They serve one purpose here, and it does not touch the counts. Where one of them summarises a publication a catalogue also holds, that summary is attached to the record and credited to the database that wrote it (§ 2). It adds prose to a record already in the corpus; it adds no record, and it moves no figure.
A count from catalogues measures those catalogues. IxTheo indexes German and European theology closely, the Gnomon database indexes classical scholarship, and subject headings are richer on recent records everywhere. All of that pushes the counts, and it is visible in the Observatory. The one span no catalogue covers well is the older literature: that is Crouzel's, and § 0 says why it is not here.
2. Rights, and what is published
Attribution on every summary, and removal for the asking.
An author, a title, a year, a journal, an identifier: bibliographic metadata are facts, and each catalogue behind this map publishes them under the terms the Credits page gives source by source. They are what this site republishes: title, authors, year, language, container, publisher, DOI, ISBN, subject headings, format. A field is published where the record carries it, and nothing stands in its place where it does not. 884 of the 2 490 work clusters name a publisher, 569 an ISBN, 165 a DOI.
Abstracts are shown as well. A reader who sees only a title cannot tell whether a study answers the question they came with, and a bibliographic map that withholds the one paragraph that would settle it serves nobody. So every summary in the Explorer is displayed, and every summary displayed names the database that wrote it and links to the record there. A summary written for this project says so instead of naming a database.
Removal on request. A publisher, a partner database or an author who asks for their summaries to be taken down gets it, without argument and without a negotiation. One message to [email protected] naming the database or the imprint is enough; the summaries come off the site and out of the next data release. The partner databases were told in writing on 15 August 2026, before any of this was published. The procedure is repeated on the Credits page, next to the licence of each source.
What the source declared about its own rights is kept with each record, internally, so that a request can be honoured in one pass rather than by hand. It is a trace, not a gate: it decides nothing about what is shown.
Coverage. The source catalogues describe rather than summarise: 177 of the 2 490 records carry a summary of their own. Where another database summarises the same publication, that summary is joined to the record by catalogue number, by DOI, or by title and year, and 190 records are covered that way. In all, 367 records of 2 490 (14.7 %) show a summary. The other 2 123 have no displayed summary.
Who wrote them. How a summary reached this file is one question; who wrote it is another, and the second is the one the credit answers. 187 of the 367 were written by a catalogue the record was harvested from and 180 by another database, whichever route they took to get here. The full breakdown, counted from the published file rather than typed here:
| Database that wrote the summary | Summaries |
|---|---|
| Index Theologicus (IxTheo / K10plus) | 141 |
| OpenAlex | 91 |
| GIROTA / Adamantius | 76 |
| ISIDORE (Huma-Num) | 6 |
| Crossref | 5 |
| Semantic Scholar | 2 |
| K10plus (GBV / BSZ) | 23 |
| Sudoc, Agence bibliographique de l'Enseignement supérieur (ABES) | 17 |
| B3Kat (Bibliotheksverbund Bayern / KOBV) | 2 |
| Library of Congress | 2 |
| Bibliothèque nationale de France | 1 |
| Gnomon Bibliographische Datenbank (Universität Eichstätt) | 1 |
3. The semantic vocabulary
Closed lists with stable identifiers, not coordinates in a latent space.
Each record is described on four independent axes, each a versioned file of closed values:
| Axis | Size | What it holds |
|---|---|---|
| Works | 33 | the works of Origen by their usual Latin titles and sigla, plus the default value for a study that names none |
| Themes | 16 / 61 | sixteen domains, sixty-one leaves; only a leaf is ever written on a record, the domain is for aggregation and for the map |
| Approaches | 10 | philological, exegetical, doctrinal, historical, reception, comparative, methodological, edition, review, survey |
| Relevance | 4 | the four classes that decide whether a record counts in a density figure |
The theme axis was built by crossing two sources, and the file keeps the trace of both: the thirty-seven published sections of the Adamantius repertorio, which give the axis of author and milieu, and the subject headings of the harvest itself, which give the doctrinal and exegetical axis. Every leaf lists the headings that motivated it.
Labels are given in English, German, French and Italian. The German label is the catalogue heading itself wherever one exists, the Italian follows the wording of the repertorio, and any label supplied by the vocabulary rather than found in a source is marked as supplied. The Explorer shows the English label on the map and the other three on the hover card, so that a reader who knows the field under its German headings can still recognise the cluster.
How a record is tagged
The tagging program sends the metadata of one record against the four vocabularies and accepts only values that exist in them: the schema is derived from the vocabulary files, so it cannot drift from them. A value that falls outside is dropped at validation, the repair is written into the record, and the record is flagged for review. Records that fail to be read at all are kept in a rejects file with their cause; none disappears. Identifiers are deterministic, so two runs address the same objects, and each record carries the vocabulary version, the prompt version and a digest of the exact input submitted. Changing the vocabulary means a new version and a new wave, so that tags made under different instructions stay comparable. A correction inside one wave is appended rather than overwritten, and the file is compacted at the end of the wave, its superseded lines kept in a history beside it. One wave file has been damaged by a renumbering tool and rebuilt from the site asset, losing its confidence and justification fields; it is kept as an archive, and nothing reads it.
The hand-tagged reference set, and where it stands
The project holds itself to one figure: on a set of records tagged by hand, the program and the hand must agree at least nine times in ten on the question that decides every count: does this record enter the density or not? The set was fifty records for the second wave, drawn with a fixed seed and tagged against the same vocabulary and the same instructions. The tags this version draws come from that second wave, run over the whole working corpus and joined back to the catalogue records this map shows.
No current agreement figure is published on this page. The previous measurement, taken under instructions that have since been rewritten, came in under the bar. Six records decided it, and they fall on two boundaries the earlier instructions had left to judgement: an old printed edition of a text of Origen, and the line between a study where he holds an identifiable section and one where a cataloguer's classification is all there is. Both are now settled in writing, and the second is doubled by a rule in the validator that no answer can override.
The fifty have since been tagged again by hand under those settled instructions, and the figure will be published from a run under the same instructions and from nothing else. Comparing the new hand set against the second wave, the run whose tags this version draws, gives 48 of 50 on the density question, but that comparison moves the reference rather than the program, and it is not the measurement: it is recorded in the working notes, not printed as a validation. Republishing an earlier figure as a current validation would be reporting a test that was never run under the present instructions.
A known bias, and what became of it. The rule used in the first wave demanded that the metadata name Origen positively, and filed everything else as not about Origen. On a curated perimeter, where cataloguers have already attached each record to Origen's authority record, that rule was too severe: a study of original sin in the Fathers, or of ministry in the early Church, was filed out although Origen is one of its witnesses. Seven of the eight disagreements in the thirty-record pilot pointed that way. The instructions now floor the class at mentioned only for any record catalogued as being about him, and keep not about Origen for homonyms and for texts by Origen catalogued as texts about him. The harvest was classed again under them: the reservoir holds 1 record where it held 223, and mentioned only holds 280 where it held 9, retrieved by a search and counted in no figure. What remains of the bias runs the other way now: a record the catalogue attached to Origen is credited with a mention even where the metadata alone would not carry it.
4. The density figure
A compass toward thin ground, not a verdict on originality.
What it is. A count of records inside a named node of the vocabulary: 47 works in this perimeter means forty-seven records of this harvest are filed under that node. The node has a name, a path and a definition, so the figure can be checked. It always travels with four things: the source corpus, the wave that produced the tags, the relevance threshold applied, and the share of the node flagged for review. In the Explorer they are printed under the figure of every cluster, search and whole-corpus panel, with the number of those records that carry an abstract and the number without a year.
What it is not. It is not a measure of the originality of a project, and not a measure of quality. A thin area may be thin because nobody has looked, or because there is nothing to find, or because this catalogue does not index the language in which the work was done. The figure opens a question; it does not close one.
Which records count. Every figure on this site is a count of the same 2 209 records: those classed core or partial, where Origen is the subject or holds an identifiable section of the argument. Of the 2 490 records kept, 280 are classed mentioned only and 1 is held outside the count. Neither class enters a figure on any page, whether the bars of the Observatory, the language counts, the number on a question chip or the density of a cluster.
Retrieval is wider than counting, and stays wider. A record where Origen is mentioned only still answers a search and still answers the four questions: nothing is put aside, as the map says on every page. It is returned, listed and readable, and it is not added to the figure. Where a surface can return more than it counts, it says so in words and gives the second number: N further works are mentioned only and are listed below the count.
What the count leaves out, and does not hide. Every one of the 2 490 records carries a class. Twenty-eight did not: their cluster had been split or joined by the deduplication after the wave had run, or a mechanical pre-sort had set them aside, and the count held them outside every figure for want of a tag rather than guess one. They were sent back to the tagger in a later pass, and the build now refuses to publish while a record on display carries no class.
What it does not include. The size of a disc on the map is an academic weight: the mean of a citation percentile inside the work's own cohort (same decade, same document type, same language) and a structural percentile in the graph. That weight is drawn, and only drawn. It never filters, it never removes anything from a search, and it never enters the density. A record without a citation figure keeps the base size and reads no citation data rather than being pushed down.
The citation figure is unevenly available, and the unevenness runs along languages. Citation counts come from a single measurable source, and it does not cover the field evenly. Across the working corpus of 42 210 clusters, 35.0 % carry a count at all, and the share by language of publication is Spanish 88.9 %, German 39.8 %, English 35.9 %, Italian 24.4 %, French 9.6 %. A French article is not less cited than a Spanish one; it is less often indexed where citations are counted. This is why the weight is a cohort rank rather than a raw count (a record is ranked against others of its own decade, type and language, so that a thinly indexed language is not read as a thinly cited one), and why the weight never filters anything out of a result. Read across languages, the disc sizes still carry that bias, and no correction here removes it.
No reading or download figures appear anywhere on this site. No source holds them reliably, and simulating one would be an invention.
5. Nothing is thrown away
The map has reservoirs rather than a bin. Records classed as not about Origen have one; publications that bear on no single work of Origen have theirs. A reservoir is folded under the field, named and counted, and it is drawn on the map when the reader asks for it: the named clusters are what the field is for, and a reservoir left open would be the widest cloud on it. They remain searchable, and they remain in every export.
One rule for the counts, on this page, in the questions of the Explorer and in the Observatory alike: a figure counts the 2 209 records classed core or partial (§ 4), and nothing else. The reservoirs hold the rest, named and counted as reservoirs. Saying how much of a harvest cannot be placed is part of what a map of a field owes its reader.
6. Reproducibility
- The vocabulary is versioned and cited on every record.
- The instructions given to the classifier are versioned and can be printed word for word.
- The schema is derived from the vocabulary, not copied from it.
- The exact input submitted for each record is fingerprinted.
- The hand-tagged reference set fixes the human reference and is published with the run, agreement figure and disagreements alike.
- Rejects and review rates are published with any figure drawn from the tags.
- Every density figure states its perimeter.
A third party who runs the same vocabulary, the same instructions and the same schema on the same records can measure the difference record by record. That is the sense in which a classification of this kind can be checked at all.
7. Questions
Short answers, each standing on its own.
What does the density figure measure?
A density is a count of records inside a named node of the vocabulary, and nothing more. When the map says forty-seven works sit in a perimeter, forty-seven records of this harvest are filed under that node, out of the 2 209 the site counts. The node has a name, a path and a definition, so the figure can be checked against the same data. What a density does not measure is the originality of a project: thin ground may be unexplored, or exhausted, or written in a language this catalogue indexes poorly, and the map cannot tell those three apart. It opens a question rather than settling one. Academic weight, drawn as the size of a disc, is a separate thing again: a citation percentile inside a cohort of the same decade, type and language, which never filters a search and never enters a count.
Which sources does the map draw on?
Eight source labels, each queried through the authority record for Origen (GND 118590235) and its national equivalents: the Index Theologicus of Tübingen hydrated from K10plus, the Gnomon Bibliographische Datenbank served in MARC by B3Kat, K10plus beyond its theology slice, Sudoc, B3Kat, the Deutsche Nationalbibliothek, the Library of Congress and the Bibliothèque nationale de France. § 1 gives the number of records from each, and the Credits page gives each licence. Every entry in the Explorer links back to the record it came from. Nine further harvests exist in the working data: OpenAlex, Crossref, Semantic Scholar, Isidore, theses.fr, Dialnet, the Italian SBN, the Adamantius repertorio and the BIBP of Laval. These further harvests put no record on this map and move no figure, because their relevance filtering is not finished. They serve one purpose here: where one of them summarises a publication a catalogue also holds, that summary is attached to the record and credited to whoever wrote it. Subscription indexes cannot be harvested and are absent, so coverage before 1980 is thinner than the counts suggest.
How are the summaries credited, and how are they removed?
Every summary shown in the Explorer names the database that wrote it and links to the record there; a summary written for this project says so instead of naming a database. The regime is attribution and takedown. A publisher, a partner database or an author who wants their summaries off the site writes to [email protected], naming the database, the imprint or the article, and those summaries come off the site and out of the next data release. No justification is asked for and none is needed. The removal is scripted rather than done by hand, so it lands in one pass and cannot be half applied. The record keeps its bibliographic metadata and its link, so nothing is lost to the reader but the paragraph. The partner databases were written to on 15 August 2026, before any of this was published.
How should Origenality be cited?
In a footnote: Romain Girardi, Origenality: a bibliographic map of Origen studies,
2026, https://origenality.com. The machine-readable form is in CITATION.cff at
the root of the repository, which carries the version, the date of release and the
affiliation. Cite the underlying databases as well when you cite a record: each source has
its attribution template in ATTRIBUTION.md, and Semantic Scholar in particular
requires attribution under ODC-BY. The site is a pointer rather than a substitute, and a
record found here should be read in the catalogue that describes it. The code is under the
MIT licence; the data are published under attribution and takedown, source by source, as
DATA_POLICY.md sets out, and no single licence is asserted over the whole.