
T’OMIM: A Morphologically Annotated Dataset of Parallel Passages in the Hebrew Bible
Abstract
Inner-biblical parallel passages have been catalogued by generations of Hebrew biblical scholars, but a machine-readable, morphologically annotated dataset of these parallels has not been available for computational research. T’OMIM (Tanakh Observable Matches of Intertextual Mimesis) addresses this gap. The dataset pairs 554 narrative parallels from the Chronicles synoptic tradition with 256 poetic parallels drawn from major studies of biblical parallelism, aligning each pair to the Biblia Hebraica Stuttgartensia Amstelodamensis (BHSA) dataset, maintained by the Eep Talstra Centre for Bible and Computer (ETCBC), at both verse and word granularity. The resulting four Apache Parquet files are released on Zenodo under a CC-BY-4.0 license, supporting research on semantic similarity, text reuse, and intertextual retrieval in Classical Hebrew.
© 2026 David M. Smiley, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.