Skip to content

Global lemmatizer

The lemmatizer is the tool that turns occurrences into attestations. It is where a token in a text acquires its lexical or onomastic entry, and therefore where the Lexicon and the Named entities register are actually built.

It depends on the type of account, not on a module permission

The lemmatizer is open to Advanced users and Administrators only. An account holding permission to edit named entities but without editorial status cannot open it — this is the most frequent reason for "I have the rights and it still refuses me".

Global lemmatizer

Global lemmatizer


1. Five tabs, five sets of tokens

Each tab is a filter on the class of the token, and each writes to its own target:

Tab Tokens it lists Assigns
Lemmata tokens with no class, plus TOP and WN a lexical entry
Personal Names (PN/PNF) PN, PNF a name form
Divine Names (DN) DN a name form
Geographic Names (GN) GN a name form
Month Names (MN) MN a name form

The tabs are not a preference: they are the two planes of the model made operational. A token appears in exactly one of them, and cannot be assigned an entry from the other side.

If a word is in the wrong tab, the transliteration is what needs fixing

The class comes from the marker prefixes written in the transliteration. A personal name that was never marked will show up among the ordinary words and will never reach the onomastic tabs. Correct the marker in Edit tablet, then come back.


2. The working cycle

  1. Search a written form. The query is normalised exactly as in the search engine: you may type sz for š, h for ḫ, digits for indices.
  2. Read the groups. Occurrences come back grouped by written form, across the whole corpus — every tablet in which that spelling occurs, not one text at a time. This is the point of the tool: an editorial decision taken once is applied to the corpus.
  3. Restrict by site if you are working on a single archive.
  4. Select the occurrences that share one interpretation. This is the editorial act, and it is deliberately manual — see §3.
  5. Choose the entry from the dictionary field and submit. Every selected occurrence receives it in a single operation.

By default the list shows only occurrences that have not yet been assigned. Tick include already lemmatized occurrences to see the assigned ones as well: that is how you revise an earlier decision, or reassign a form that was attached to the wrong entry.

Working cycle

Global lemmatizer


3. Three things the tool does not do

It does not create entries. The lemmatizer assigns existing ones, and what to do when the entry you need is missing depends on which plane you are working in:

  • a name form is created here, with New Named Entity — create it, then come back;
  • a lexical entry is not created at all. The lexicon comes from eBL, and forthcoming from LAD (why). Check the headword in eBL: if it is there and missing from DAPCA, it is an alignment problem; if it is missing there too, that is where it should be reported. Creating it locally is discouraged and, as things stand, not properly set up.

Grouping by written form is not analysis. Two occurrences of the same spelling may require different entries: the same graphic sequence can be a lexeme in one line and a proper name in another, and homographic lexemes are common. The grouping saves you the labour of finding the occurrences; deciding which of them belong together is your statement, which is why the selection is not pre-ticked.

It does not touch the text. Assigning an entry does not modify the transliteration, and correcting the transliteration afterwards does not update the assignment (nothing propagates backwards).


4. FK audit / cleanup

/lemmatizer/fk-auditAdministrators only.

A maintenance report over assignments that are inconsistent with the class of the token they sit on. Such rows exist for historical reasons: past defects in the token-splitting code, and automatic linking applied to tokens with no written form. The page lists them in three disjoint categories:

Category What it finds
Empty notation with an assignment tokens with no written form that nevertheless carry an entry — breaks and empty tokens, which cannot legitimately have one
Lexical lemma on a name token a token classified as a proper name carrying a lexical entry
Name form on a non-name token a token that is not classified as a name carrying an onomastic entry

Cleanup

Global lemmatizer

Read it as a report, not as a to-do list

The page opens as a read-only report, and this is deliberate: the mismatch categories mix genuine artefacts with legitimate editorial decisions. A logographic writing such as KA₂ classified as a geographic name but linked to the lexeme bābu, or DUMU linked to māru, is a considered choice, not a defect.

Each row links to the token in the editing view. Reviewing them one at a time is the normal procedure; applying a category in bulk is the exception, and it is irreversible from within the page.

A detached token is not lost: it keeps its class and reappears in the appropriate tab of the lemmatizer, to be reassigned.

Comments