Indexes
The indexes are the registers of the corpus: the lexicon, the onomasticon, the prosopography, the seal catalogue. Each of them is a list of entities, and each entity gathers its own attestations across the texts.
They are not compiled separately from the edition. An index entry acquires an attestation when an editor assigns a token to it; the list is a view over those assignments, never a document to be maintained by hand. This has a practical consequence worth stating at the outset: you do not add a word to the lexicon by editing the lexicon, you add it by lemmatising an occurrence.
1. Headword or named entity? A distinction that decides everything
The corpus stores every written word as a token. What an index registers is a type: the abstract entity of which the token is an instance. Two questions look alike at this point, and they are not the same question:
Of which word is this form an instance? — and — Which name does this spelling render?
The first is answered by lemmatisation, the second by named-entity recognition. In computational-linguistic terms they belong to different tasks, and DAPCA keeps them on separate planes because conflating them corrupts both.
1.1 The dictionary headword
A headword — a lemma — is the citation form of a lexeme: the class of all inflected
forms that share one lexical meaning. i-ša-am, i-ša-am-mu and ta-aš-ta-ma are
realisations of the lexeme šâmu "to buy".
Assigning a lemma is a paradigmatic statement. It says: this form belongs to that paradigm, and therefore to that dictionary meaning. The entry has a morphology, a root, a language, a guideword; it has a place in a lexicon that could be printed as such.
1.2 The named entity
A named entity is not a lexeme. It is a referring expression. pil₂-su-DINGIR-KUR does not
mean anything in the way šâmu means "to buy" — it designates.
Normalising it as Pilsu-Dagān states that this spelling renders that name. It says nothing about who bore the name, and it does not place the form in an inflectional paradigm: it places it in an onomasticon. The four classes DAPCA uses — personal, divine, geographic, month name — are types of referent, the Assyriological counterpart of the categories of a named-entity recogniser.
1.3 Why the two planes never meet
A token — in a specific context — belongs to one plane or the other, never to both, and the plane is decided by the token's class, which is fixed in the transliteration by the marker prefixes:
| Class of the token | Plane | Register |
|---|---|---|
PN / PNF (personal, m./f.), DN (divine), GN (geographic), MN (month) |
named entity | Named entities |
no class, TOP (urban topography), WN (profession, role) |
lexical | Lexicon |
A named token can only receive a name form; a lexical token can only receive a lemma. This is enforced, not merely recommended: an assignment that would cross the line is refused, or detached with a warning.
Where the difficulty really lies: logograms
The hard cases in a cuneiform corpus are not ambiguities of the model but of the
script. KA₂ may be the noun bābu "gate" or the designation of a place; DUMU is
māru "son" and, at the same time, the pivot of the filiation formula.
DAPCA resolves this at the level of the token, text by text, rather than deciding once and for all about the sign. The same graphic sequence may therefore be lexical in one line and onomastic in another, which is exactly what the evidence supports.
The classes TOP and WN sit deliberately on the lexical side: an urban toponym
or a professional designation is an Akkadian term used referentially, not a proper
name.
1.4 And a third act: identifying the individual
Normalising a name is a judgement about writing and language. Deciding that this attestation of Pilsu-Dagān and that one, in another cuneiform tablet, are the same man is a judgement about history.
Two men called Pilsu-Dagān share one name form and are two distinct persons; one man whose name is spelled three ways has three spellings and one identity. Keeping the acts apart is what allows you to be certain about a reading while remaining undecided about its bearer. See §3.
2. The four registers
| Register | What it collects | The act that feeds it | Example |
|---|---|---|---|
| 1. Lexicon | lexemes, by citation form | lemmatisation | ana "to, for" |
| 2. Named entities | normalised forms of proper names | onomastic normalisation | Pilsu-Dagān |
| 3. Historical Persons | historical individuals | prosopographic identification | Pilsu-Dagān son of Baʿlu-kabar |
| 4. Seal catalogue | seal objects | cataloguing, then attestation | seal A45 |
Genealogy is not a fifth register but a layer derived from the third (Historical Persons), and
the Global lemmatizer is the tool that feeds the first two (Lexicon and Named entities).
3. A historical person is not a marked filiation
Two operations look adjacent and are conceptually distinct. One is done on a tablet, in Mark genealogies; the other maintains a register.
| Mark genealogies | Historical Persons | |
|---|---|---|
| What it asserts | this text writes that the name in line 4 is the son of the name in line 5 | these attestations, in different texts, are one and the same man |
| What it attaches to | two occurrences on the same tablet | an individual, through their attestations |
| Who supports it | the text | the editor, on prosopographic grounds |
| If it is wrong | the text has been misread | the identification is mistaken |
| Does it hold outside the text | no | yes, by definition |
Mark genealogies records evidence; a historical person is an interpretation. A tablet may state a filiation between two men neither of whom has been identified — the statement stands, and remains useful, without any person existing.
How evidence becomes a family tree is described in Genealogy. It is not automatic, and the step that is usually missing is not the marking of more filiations.
4. Adding entries and linking them to what exists
The single most important thing to understand about editorial work here:
Creating an entry and linking it to an occurrence are two separate operations
They are done in different places, and either can exist without the other. An entry with no attestations exists but does not appear in the lists, which show attested material by default. An occurrence with no entry stays in the lemmatizer, waiting.
| Register | Where the entry comes from | Where it is linked to occurrences |
|---|---|---|
| Lexicon | not created here — the entries come from eBL and, forthcoming, LAD (details) | Global lemmatizer, Lemmata tab |
| Named entities | created here, with New Named Entity | Global lemmatizer, the four name tabs |
| Historical Persons | created here, with New Person | the tablet's Prosopography view |
| Seal catalogue | created here, with New Seal | the tablet's Seals mode |
The lexicon is the exception, and the reason is worth stating. The other three registers describe material that only this project holds: the names as this corpus spells them, the individuals who appear in it, the seals impressed on its tablets. Nobody else can supply them. An Akkadian dictionary is not in that position — it exists, it is maintained by projects devoted to it, and DAPCA aligns to it rather than duplicating it.
4.1 The three ways a link comes into being
- At import, automatically. Loading a text can re-attach its tokens to entries that already exist, each within its own plane, never creating any (Adding a new tablet). It is a starting point, not a result: where a spelling points to several entries the importer takes the most frequently attested one and flags the word for review.
- In bulk, in the lemmatizer. You select the occurrences that share one interpretation and assign the same entry to all of them at once (/lemmatizer/).
- One at a time, from the entity's own page or from the tablet — a prosopographic
attestation, a seal attestation. A user — with the correct role (administrator or advanced user) —
can create and/or edit the link between a token and its normalized form from the
/tablet/<id>/viewpage (/tablet/5/view) after clicking on the single transliterated word.
4.2 Nothing propagates backwards
Warning
Correcting the transliteration of a token after it has been lemmatised does not update the entry, and it does not warn the entry. When a token is split, the lemma follows the token you designate — but prosopographic attestations, kinship relations and annotations do not (for instance Edit tablet).
The registers record editorial statements. They are revised by editors, not recomputed.
5. Who sees what
| Register | To read | To edit |
|---|---|---|
| Lexicon | anyone | editorial accounts |
| Named entities | anyone | permission to edit named entities |
| Historical Persons, Genealogy | permission for prosopography | permission to edit prosopography |
| Seal catalogue | permission for seals | permission to edit seals |
| Global lemmatizer | editorial accounts | editorial accounts only (Advanced user, Administrator) |
Asymmetries
Two asymmetries to keep in mind. The lemmatizer depends on the type of account, not on a module permission: an account holding onomastic permissions but not editorial status cannot open it. And attestation counts everywhere are relative to the texts your account may see: a person or a name form attested only outside your perimeter may be missing from a list that another account finds populated.