Skip to main content
Have a personal or library account? Click to login
Reuse by Design: A Pivot-Based Architecture for the KIParla Corpus of Spoken Italian Cover

Reuse by Design: A Pivot-Based Architecture for the KIParla Corpus of Spoken Italian

Open Access
|Jun 2026

References

  1. Artstein, R., & Poesio, M. (2008). Inter-coder agreement for computational linguistics. Computational Linguistics, 34(4), 555596. https://direct.mit.edu/coli/article/34/4/555-596/1999
  2. Cerruti, M., & Ballarè, S. (2021). ParlaTO: corpus del parlato di torino. Bollettino dell’Atlante Linguistico Italiano (BALI), 44, 171196.
  3. Chrupała, G. (2023). Putting Natural in Natural Language Processing. In A. Rogers, J. Boyd-Graber, & N. Okazaki (Eds.), Findings of the Association for Computational Linguistics: ACL 2023 (pp. 78207827). Association for Computational Linguistics. 10.18653/v1/2023.findings-acl.495
  4. Dobrovoljc, K. (2022). Spoken Language Treebanks in Universal Dependencies: an Overview. In N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk, & S. Piperidis (Eds.), Proceedings of the Thirteenth Language Resources and Evaluation Conference (pp. 17981806). European Language Resources Association. https://aclanthology.org/2022.lrec-1.191/
  5. Gagliardi, G. (2018). Inter-Annotator Agreement in linguistica: una rassegna critica. In Proceedings of the Fifth Italian Conference on Computational Linguistics (CLiC-it 2018) (pp. 207213). CEUR Workshop Proceedings. https://aclanthology.org/2018.clicit-1.37/
  6. Goria, E., & Mauri, C. (2018). Il corpus KIParla: una nuova risorsa per lo studio dell’italiano parlato. In F. Masini & F. Tamburini (Eds.), CLUB Working Papers in Linguistics vol. 2 (pp. 96116). Alma Mater Studiorum Università di Bologna.
  7. Hedeland, H., & Schmidt, T. (2022). The TEI-based ISO Standard ‘Transcription of spoken language’as an Exchange Format within CLARIN and beyond. CLARIN Annual Conference (pp. 3445). https://ecp.ep.liu.se/index.php/clarin/article/view/415
  8. Institute of the Czech National Corpus (2025). NoSketch Engine. https://www.korpus.cz/noske. (Corpus query system).
  9. Jefferson, G. (2004). Glossary of transcript symbols with an introduction. In G. H. Lerner (Ed.), Pragmatics & Beyond New Series vol. 125 (pp. 1331). John Benjamins Publishing Company. https://benjamins.com/catalog/pbns.125.02jef
  10. Linell, P. (2019). The Written Language Bias (WLB) in linguistics 40 years after. Language Sciences, 76. 10.1016/j.langsci.2019.05.003
  11. MacWhinney, B. (2019). CHAT manual. TalkBank. Retrieved 2026-02-24, from https://talkbank.org/0info/manuals/CHAT.pdf
  12. Mauri, C., Ballarè, S., Goria, E., Cerruti, M., & Suriano, F. (2019). KIParla Corpus: A New Resource for Spoken Italian. In Proceedings of the Sixth Italian Conference on Computational Linguistics.
  13. Mauri, C., Ballarè, S., & Zucchini, E. (2024a). Modulo KIPasti. 10.60760/unibo/kipasti
  14. Mauri, C., Ballarè, S., & Zucchini, E. (2024b). Modulo ParlaBO. 10.60760/unibo/parlabo
  15. Mauri, C., Zucchini, E., Goria, E., & Bernasconi, B. (2025). Il progetto DiverSIta. Documentare la diversità nell’italiano parlato. CLUB – Circolo Linguistico dell’Università di Bologna. 10.6092/unibo/amsacta/8647
  16. Max Planck Institute For Psycholinguistics, The Language Archive (2025). ELAN (version 7.0) [computer software]. https://archive.mpi.nl/tla/elan
  17. Pannitto, L., Zucchini, E., Ballarè, S., Bosco, C., Mauri, C., & Sanguinetti, M. (2025). Introducing KIParla forest: seeds for a UD annotation of interactional syntax. In E. Hajičová & S. Kahane (Eds.), Proceedings of the Eighth International Conference on Dependency Linguistics (depling, Syntaxfest 2025) (pp. 5473). Association for Computational Linguistics. https://aclanthology.org/2025.depling-1.5/
  18. Steiner, I. (2017). A DevOps Manifesto for Speech Corpus Management. In Proceedings of ESSV 2017.
  19. Straka, M. (2018). UDPipe 2.0 Prototype at CoNLL 2018 UD Shared Task. In Proceedings of the CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies (pp. 197207). Association for Computational Linguistics. https://www.aclweb.org/anthology/K18-2020
  20. Waldon, B., & Schneider, N. (2025). A GitHub-based workflow for annotated resource development. In S. Peng & I. Rehbein (Eds.), Proceedings of the 19th Linguistic Annotation Workshop (LAW-XIX-2025) (pp. 326331). Association for Computational Linguistics. https://aclanthology.org/2025.law-1.27/
  21. Werthmann, A. (2025). From spoken language data to TEI-based ISO. Harmonizing language data: Standards for linguistic resources, 4, 145. 10.1515/9783112208212-007
  22. Wittenburg, P., Brugman, H., Russel, A., Klassmann, A., & Sloetjes, H. (2006). ELAN: a professional framework for multimodality research. In Proceedings of the fifth international conference on language resources and evaluation (LREC 2006). European Language Resources Association (ELRA). 10.63317/5pwa5zpssv4z
DOI: https://doi.org/10.5334/johd.527 | Journal eISSN: 2059-481X
Language: English
Page range: 81 - 81
Submitted on: Mar 1, 2026
Accepted on: Jun 10, 2026
Published on: Jun 24, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Ludovica Pannitto, Caterina Mauri, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.