
A Part-of-Speech Tagger for Older Scots: A spaCy-Based Model Using the CLAWS7 Tagset
Abstract
This paper presents a part-of-speech (PoS) tagger developed for Older Scots. The tool was developed using spaCy machine learning models and trained on approximately 100,000 words of pre-tagged Older Scots literature. The tagger’s underlying spaCy model uses the CLAWS7 PoS tagset instead of the Universal Dependencies tagset typically used in spaCy. The overall accuracy of the tagger is 89.83%, with performance constrained by the limited availability of Older Scots material and by considerable spelling variation. Despite these limitations, the tagger represents an important step forward in the computational treatment of Older Scots and provides a practical tool for researchers working with historical Scottish linguistic material. The spaCy model and tagging script are available through the JOHD Dataverse repository.
© 2026 Beth Beattie, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.