Post-OCR Corrections Using Latin Annotations (POCULA) is a text data correction project led by Patrick J. Burns at NYU’s Institute for the Study of the Ancient World Library. The project addresses corruption in digitized Latin texts from large repositories like Common Corpus by developing machine learning models to correct optical character recognition (OCR) errors. The project also used corrected Latin text as the basis for building pedagogical resources like student readers.

Goals

POCULA aims to:

  • Compile 10,000 Latin sentence pairs from Common Corpus texts
  • Match digitized documents to original scans where possible
  • Train interim correction models at 1,000, 2,500, and 5,000 sentence milestones
  • Enable correction of corrupted Latin text at scale to improve language model training

By correcting even 1% of Common Corpus’s 34 billion tokens, POCULA could yield approximately 340 million tokens—among the largest available digitized Latin collections—providing a foundation for improved Latin language models.

Team

  • Patrick J. Burns, Project Director
  • Patrick Liu, Contributor (Summer 2025)
  • Kai Leigh Harrison, Contributor (Spring 2026)
  • Kevin O’Neill, Contributor (Spring 2026)
  • Mihou Tatsumi, Contributor (Spring 2026)

Acknowledgements

POCULA is supported by the Institute for the Study of the Ancient World at New York University.