Philology Meets Machine Learning: Digital Annotation and Semantic Modelling of Byzantine Greek Texts
Diplomatic transcriptions of medieval texts pose particular challenges for digital annotation: their fragmentary preservation, non-standard orthography, and heterogeneous linguistic features complicate processing and limit the reuse of existing tools. To address these issues, we developed a full annotation pipeline for unedited medieval Greek, including part-of-speech tagging, morphological analysis, and lemmatisation. A manually created gold standard…
