Development
This section covers the technical side of building the corpus: the conventions and procedures that apply before and around the editorial work documented elsewhere in this wiki.
Two matters are covered so far, one for each form in which material enters the corpus — the text and its reproductions:
| Page | What it covers | Who needs it |
|---|---|---|
| Transliteration encoding | the diplomatic notation: line and surface syntax, sign relationships, breakage, semantic classifiers, permitted characters | anyone entering or correcting a text |
| Images and hand copies | what the system does to an uploaded file, how to prepare one beforehand, and with which tools | anyone preparing reproductions |
They have more in common than they appear to. In both cases the platform does something determinate with what you give it: the parser decomposes the transliteration into lines, tokens and sign readings; the upload rescales, re-encodes and renames the image file. Knowing in advance what will happen is what lets you prepare material so that nothing is lost on the way in.
Where the neighbouring documentation lives
The operations themselves are documented under The Tablet > Adding a new tablet — importing a text, correcting tokens, annotating — and the registers they feed under Indexes. This section is about the conventions those operations assume.
Both procedures require an editorial account. Reading these pages requires nothing: the notation is also the convention in which the corpus is displayed, and it is worth knowing in order to read a text critically.