Skip to main content
Have a personal or library account? Click to login
Reuse by Design: A Pivot-Based Architecture for the KIParla Corpus of Spoken Italian Cover

Reuse by Design: A Pivot-Based Architecture for the KIParla Corpus of Spoken Italian

Open Access
|Jun 2026

Figures & Tables

Figure 1

Current KIParla pipeline: data processing phases.

Figure 2

Screenshot of ELAN showing the alignment between the audio (.wav) file and the transcription units for file BOA1010 in KIP.

Figure 3

The same excerpt from conversation BOA1010 in three representations. Top: abridged EAF/XML source of BOA1010 (line numbers from the actual file). Time slots ts8–ts12 define intervals that are interleaved in time (ts10=7861 ms falls inside ts8–ts9=7677–9586 ms), yet the annotations that reference them — annotation a5 in tier BO039 (line 1218) and annotations a6–a7 in tier BO038 (lines 3337–3343) — are separated by over 2100 lines of XML, connected only through opaque time-slot identifiers. Bottom left: Jeffersonian linearization. Bottom right: orthographic linearization.

Table 1

Metadata and repository information for the deposited KIParla transcript datasets.

KIP transcripts
Repository locationDOI: https://doi.org/10.60760/unibo/kip Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2123
Repository nameCLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), hosted within the CLARIN infrastructure.
Object nameKIP-transcripts-v1.1.0.zip
Format names and versionsZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata files and documentation (README.md, CITATION.cff, LICENSE). Version: v1.1.0.
Creation datesData collection: 2016-01-01 to 2019-12-31 (recordings conducted between 2016 and 2019).
Dataset creatorsSilvia Ballarè (University of Bologna) – data revision and pseudonymization; Eugenio Goria (University of Turin) – corpus preparation, module coordination; Caterina Mauri (University of Bologna) – scientific coordination and project supervision.
LanguageItalian (standard Italian, with occasional regional varieties); English (used for documentation).
LicenseCreative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0).
Publication date2019-06-04.
ParlaTO transcripts
Repository locationDOI: https://doi.org/10.60760/unibo/parlato Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2125
Repository nameCLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), within the CLARIN infrastructure.
Object nameParlaTO-transcripts-v1.1.0.zip
Format names and versionsZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata and documentation files. Version: v1.1.0.
Creation datesData collection: 2018-01-01 to 2020-12-31 (semi-structured interviews conducted between 2018 and 2020).
Dataset creatorsSilvia Ballarè (University of Bologna) – module coordination and corpus preparation; Massimo Cerruti (University of Turin) – scientific coordination and project supervision.
LanguageItalian (spoken Italian, with regional variation typical of Turin and Piedmont); English (used for documentation).
LicenseCreative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0).
Publication date2020-06-01.
KIPasti transcripts
Repository locationDOI: https://doi.org/10.60760/unibo/kipasti Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2124
Repository nameCLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), within the CLARIN infrastructure.
Object nameKIPasti-transcript-v1.1.0.zip
Format names and versionsZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata and documentation files. Version: v1.1.0.
Creation datesData collection: 2020-01-01 to 2024-12-31 (mealtime conversations recorded between 2020 and 2024).
Dataset creatorsSilvia Ballarè (University of Bologna) – module coordination and corpus design; Caterina Mauri (University of Bologna) – module coordination and corpus design; Eleonora Zucchini (University of Bologna) – data revision and pseudonymization.
LanguageItalian (predominantly), with frequent occurrences of regional dialect varieties; English (used for documentation).
LicenseCreative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0).
Publication date2024-04-30.
ParlaBO transcripts
Repository locationDOI: https://doi.org/10.60760/unibo/parlabo Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2126
Repository nameCLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), within the CLARIN infrastructure.
Object nameParlaBO-transcripts-v1.1.0.zip
Format names and versionsZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata and documentation files (README.md, CITATION.cff, LICENSE). Version: v1.1.0.
Creation datesData collection: 2021-01-01 to 2024-12-31 (semi-structured interviews conducted between 2021 and 2024 in Bologna and its province).
Dataset creatorsSilvia Ballarè (University of Bologna) – module coordination and corpus preparation; Caterina Mauri (University of Bologna) – module coordination and corpus preparation; Eleonora Zucchini (University of Bologna) – data revision and pseudonymization.
LanguageItalian (spoken Italian, with regional variation typical of Bologna and Emilia-Romagna, including occasional dialectal segments); English (used for documentation).
LicenseCreative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0).
Publication date2024-10-01.
StraParlaBO and StraParlaTO transcripts
currently being deposited.
Table 2

Transcription conventions.

SYMBOLMEANINGTYPELEVEL
word,Weakly rising intonationCategoricalToken
word?Rising intonationCategoricalToken
word.Falling intonationCategoricalToken
wo:rdProlonged soundCategoricalCharacter
(.)Short pauseCategoricalToken
word=wordProsodically linked units
°word°Lower volumeCategoricalToken
WORDHigher volumeCategoricalToken
wor-Interrupted wordCategoricalToken
>word<Faster speechCategoricalSpan
<word>Slower speechCategoricalSpan
[word]Overlapping speechRelationalSpan
(word)Uncertain transcription (transcriber’s guess)CategoricalSpan
xxxUnintelligible sequence (one x per syllable)CategoricalToken
((laughs))Non-verbal behaviorCategoricalToken
Figure 4

Depiction of a case where one TU (A) participates in two distinct overlap events. Example extracted from Figure 2.

Figure 5

Depiction of a case where one TU (B) participates in two distinct overlap events. Example extracted from Figure 2.

Table 3

Description of the columns composing the TSV format. The file contains a header useful to interpret columns.

COLUMNDESCRIPTION
token_idUnique identifier of the token within the conversation. It consists of two integers separated by a hyphen (^\d+-\d+$). The first encodes the ID of the TU. It must be treated as an opaque identifier.
speakerEither a valid KIParla speaker code or ??? to signal missing participant metadata (!^[A-Z]+[0-9]+|\?3$ !). Parsed as a categorical string.
tu_idProgressive identifier of the TU (^\d+). Tokens sharing the same tu_id form a contiguous block and reconstruct one validated TU.
spanExact substring of the original Jefferson transcription corresponding to the token, including all symbols (e.g., colons, brackets, punctuation, variation markers). Must be treated as a literal string without normalization.
formNormalized orthographic representation of the token. Jefferson-specific symbols are removed and encoded in dedicated feature columns. Short pauses are encoded as [PAUSE]; unintelligible tokens are normalized to x. Parsed as a Unicode string representing the orthographic base form.
typeCategorical label describing the token class. Possible values include linguistic, nonverbalbehavior, shortpause, unknown, and error (reserved for cases requiring manual inspection). Parsed as a closed-set categorical variable.
jefferson_featsPipe-separated list of key–value pairs in the format Feature=Value. The field may be empty. Parsers should split on |, then on =. Possible values are described in Table 4. Field may be empty.
alignAlignment metadata, present only on the first and last token of each TU. It is encoded as a pipe-separated key–value pairs: AlignBegin=<float> and/or AlignEnd=<float>, expressed in seconds. Enables deterministic reconstruction of time-aligned units.
prolongationsComma-separated list of elongation encodings in the format <char_id>x<count>. char_id is the zero-based index in form; count is the number of consecutive colons in the original span. Example: 2x2,6x1. Field may be empty.
paceEncodes participation in fast or slow spans. Either empty or one of Fast=<start>-<end> or Slow=<start>-<end>, where indices are zero-based over form.
guessesComma-separated list of uncertain character ranges in the format <start>-<end>, using zero-based indices over form. Field may be empty.
overlapsComma-separated list of overlap encodings in the format <start>-<end>(<overlap_id>). Indices are zero-based over form; overlap_id is the progressive identifier of the temporal overlap event derived via the clique-based algorithm. If unresolved, the identifier is ?.
Figure 6

Verticalized format of conversation BOA1010.

Table 4

Feature-Value pairs for column jefferson_feats.

FEATUREVALUES
SpaceAfterNo when tokenization happens on apostrophe
ProsodicLinkYes when tokenization happens on the = symbol
IntonationFalling, Rising, WeaklyRising for . , ? respectively
InterruptedYes for tokens ending in -
TruncatedYes for tokens ending in '
VolumeHigh, Low for tokens in capitalized spans or spans enclosed in °…°
Figure 7

Towards a new bidirectional pipeline.

Figure 8

Conversation KPC001 after lemmatization and POS tagging. Jefferson features are not shown for reasons of space.

DOI: https://doi.org/10.5334/johd.527 | Journal eISSN: 2059-481X
Language: English
Page range: 81 - 81
Submitted on: Mar 1, 2026
Accepted on: Jun 10, 2026
Published on: Jun 24, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Ludovica Pannitto, Caterina Mauri, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.