Welcome
Welcome to the Post-OCR Corrections Using Latin Annotations (POCULA) project. POCULA is a text data correction project led by Patrick J. Burns at NYU’s Institute for the Study of the Ancient World Library. The project addresses corruption in digitized Latin texts from large repositories by developing machine learning models to correct optical character recognition (OCR) errors.
POCULA aims to compile 10,000 Latin sentence pairs from scanned Latin texts available online, matching digitized documents to original scans where possible. The project trains interim correction models at 1,000, 2,500, and 5,000 sentence milestones.
In addition, the project uses corrected Latin text as the basis for building pedagogical resources like student readers.
The project uses the LatinCy pretrained NLP pipelines developed by Burns to support Latin text processing.