
Figure 1
Current KIParla pipeline: data processing phases.

Figure 2
Screenshot of ELAN showing the alignment between the audio (.wav) file and the transcription units for file BOA1010 in KIP.

Figure 3
The same excerpt from conversation BOA1010 in three representations. Top: abridged EAF/XML source of BOA1010 (line numbers from the actual file). Time slots ts8–ts12 define intervals that are interleaved in time (ts10=7861 ms falls inside ts8–ts9=7677–9586 ms), yet the annotations that reference them — annotation a5 in tier BO039 (line 1218) and annotations a6–a7 in tier BO038 (lines 3337–3343) — are separated by over 2100 lines of XML, connected only through opaque time-slot identifiers. Bottom left: Jeffersonian linearization. Bottom right: orthographic linearization.
Table 1
Metadata and repository information for the deposited KIParla transcript datasets.
| KIP transcripts | |
| Repository location | DOI: https://doi.org/10.60760/unibo/kip Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2123 |
| Repository name | CLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), hosted within the CLARIN infrastructure. |
| Object name | KIP-transcripts-v1.1.0.zip |
| Format names and versions | ZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata files and documentation (README.md, CITATION.cff, LICENSE). Version: v1.1.0. |
| Creation dates | Data collection: 2016-01-01 to 2019-12-31 (recordings conducted between 2016 and 2019). |
| Dataset creators | Silvia Ballarè (University of Bologna) – data revision and pseudonymization; Eugenio Goria (University of Turin) – corpus preparation, module coordination; Caterina Mauri (University of Bologna) – scientific coordination and project supervision. |
| Language | Italian (standard Italian, with occasional regional varieties); English (used for documentation). |
| License | Creative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0). |
| Publication date | 2019-06-04. |
| ParlaTO transcripts | |
| Repository location | DOI: https://doi.org/10.60760/unibo/parlato Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2125 |
| Repository name | CLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), within the CLARIN infrastructure. |
| Object name | ParlaTO-transcripts-v1.1.0.zip |
| Format names and versions | ZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata and documentation files. Version: v1.1.0. |
| Creation dates | Data collection: 2018-01-01 to 2020-12-31 (semi-structured interviews conducted between 2018 and 2020). |
| Dataset creators | Silvia Ballarè (University of Bologna) – module coordination and corpus preparation; Massimo Cerruti (University of Turin) – scientific coordination and project supervision. |
| Language | Italian (spoken Italian, with regional variation typical of Turin and Piedmont); English (used for documentation). |
| License | Creative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0). |
| Publication date | 2020-06-01. |
| KIPasti transcripts | |
| Repository location | DOI: https://doi.org/10.60760/unibo/kipasti Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2124 |
| Repository name | CLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), within the CLARIN infrastructure. |
| Object name | KIPasti-transcript-v1.1.0.zip |
| Format names and versions | ZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata and documentation files. Version: v1.1.0. |
| Creation dates | Data collection: 2020-01-01 to 2024-12-31 (mealtime conversations recorded between 2020 and 2024). |
| Dataset creators | Silvia Ballarè (University of Bologna) – module coordination and corpus design; Caterina Mauri (University of Bologna) – module coordination and corpus design; Eleonora Zucchini (University of Bologna) – data revision and pseudonymization. |
| Language | Italian (predominantly), with frequent occurrences of regional dialect varieties; English (used for documentation). |
| License | Creative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0). |
| Publication date | 2024-04-30. |
| ParlaBO transcripts | |
| Repository location | DOI: https://doi.org/10.60760/unibo/parlabo Persistent identifier (Handle): https://hdl.handle.net/20.500.11752/OPEN-2126 |
| Repository name | CLARIN-IT Repository (ILC-CNR, Institute for Computational Linguistics “A. Zampolli”), within the CLARIN infrastructure. |
| Object name | ParlaBO-transcripts-v1.1.0.zip |
| Format names and versions | ZIP archive containing: .eaf (ELAN time-aligned XML transcription files); .txt (linearized Jefferson-style transcripts); .txt (linearized orthographic transcripts); .tsv (tokenized tab-separated files); metadata and documentation files (README.md, CITATION.cff, LICENSE). Version: v1.1.0. |
| Creation dates | Data collection: 2021-01-01 to 2024-12-31 (semi-structured interviews conducted between 2021 and 2024 in Bologna and its province). |
| Dataset creators | Silvia Ballarè (University of Bologna) – module coordination and corpus preparation; Caterina Mauri (University of Bologna) – module coordination and corpus preparation; Eleonora Zucchini (University of Bologna) – data revision and pseudonymization. |
| Language | Italian (spoken Italian, with regional variation typical of Bologna and Emilia-Romagna, including occasional dialectal segments); English (used for documentation). |
| License | Creative Commons Attribution–NonCommercial–ShareAlike 4.0 International (CC BY-NC-SA 4.0). |
| Publication date | 2024-10-01. |
| StraParlaBO and StraParlaTO transcripts | |
| – | currently being deposited. |
Table 2
Transcription conventions.
| SYMBOL | MEANING | TYPE | LEVEL |
|---|---|---|---|
| word, | Weakly rising intonation | Categorical | Token |
| word? | Rising intonation | Categorical | Token |
| word. | Falling intonation | Categorical | Token |
| wo:rd | Prolonged sound | Categorical | Character |
| (.) | Short pause | Categorical | Token |
| word=word | Prosodically linked units | ||
| °word° | Lower volume | Categorical | Token |
| WORD | Higher volume | Categorical | Token |
| wor- | Interrupted word | Categorical | Token |
| >word< | Faster speech | Categorical | Span |
| <word> | Slower speech | Categorical | Span |
| [word] | Overlapping speech | Relational | Span |
| (word) | Uncertain transcription (transcriber’s guess) | Categorical | Span |
| xxx | Unintelligible sequence (one x per syllable) | Categorical | Token |
| ((laughs)) | Non-verbal behavior | Categorical | Token |

Figure 4
Depiction of a case where one TU (A) participates in two distinct overlap events. Example extracted from Figure 2.

Figure 5
Depiction of a case where one TU (B) participates in two distinct overlap events. Example extracted from Figure 2.
Table 3
Description of the columns composing the TSV format. The file contains a header useful to interpret columns.
| COLUMN | DESCRIPTION |
|---|---|
| token_id | Unique identifier of the token within the conversation. It consists of two integers separated by a hyphen (^\d+-\d+$). The first encodes the ID of the TU. It must be treated as an opaque identifier. |
| speaker | Either a valid KIParla speaker code or ??? to signal missing participant metadata (!^[A-Z]+[0-9]+|\?3$ !). Parsed as a categorical string. |
| tu_id | Progressive identifier of the TU (^\d+). Tokens sharing the same tu_id form a contiguous block and reconstruct one validated TU. |
| span | Exact substring of the original Jefferson transcription corresponding to the token, including all symbols (e.g., colons, brackets, punctuation, variation markers). Must be treated as a literal string without normalization. |
| form | Normalized orthographic representation of the token. Jefferson-specific symbols are removed and encoded in dedicated feature columns. Short pauses are encoded as [PAUSE]; unintelligible tokens are normalized to x. Parsed as a Unicode string representing the orthographic base form. |
| type | Categorical label describing the token class. Possible values include linguistic, nonverbalbehavior, shortpause, unknown, and error (reserved for cases requiring manual inspection). Parsed as a closed-set categorical variable. |
| jefferson_feats | Pipe-separated list of key–value pairs in the format Feature=Value. The field may be empty. Parsers should split on |, then on =. Possible values are described in Table 4. Field may be empty. |
| align | Alignment metadata, present only on the first and last token of each TU. It is encoded as a pipe-separated key–value pairs: AlignBegin=<float> and/or AlignEnd=<float>, expressed in seconds. Enables deterministic reconstruction of time-aligned units. |
| prolongations | Comma-separated list of elongation encodings in the format <char_id>x<count>. char_id is the zero-based index in form; count is the number of consecutive colons in the original span. Example: 2x2,6x1. Field may be empty. |
| pace | Encodes participation in fast or slow spans. Either empty or one of Fast=<start>-<end> or Slow=<start>-<end>, where indices are zero-based over form. |
| guesses | Comma-separated list of uncertain character ranges in the format <start>-<end>, using zero-based indices over form. Field may be empty. |
| overlaps | Comma-separated list of overlap encodings in the format <start>-<end>(<overlap_id>). Indices are zero-based over form; overlap_id is the progressive identifier of the temporal overlap event derived via the clique-based algorithm. If unresolved, the identifier is ?. |

Figure 6
Verticalized format of conversation BOA1010.
Table 4
Feature-Value pairs for column jefferson_feats.
| FEATURE | VALUES |
|---|---|
| SpaceAfter | No when tokenization happens on apostrophe |
| ProsodicLink | Yes when tokenization happens on the = symbol |
| Intonation | Falling, Rising, WeaklyRising for . , ? respectively |
| Interrupted | Yes for tokens ending in - |
| Truncated | Yes for tokens ending in ' |
| Volume | High, Low for tokens in capitalized spans or spans enclosed in °…° |

Figure 7
Towards a new bidirectional pipeline.

Figure 8
Conversation KPC001 after lemmatization and POS tagging. Jefferson features are not shown for reasons of space.
