Skip to main content
Have a personal or library account? Click to login
Reviving Legacy WordNet-Like Resources: MariTerm and ItalWordNet Renewal Through Mutual Expansion and Plug-In Links Cover

Reviving Legacy WordNet-Like Resources: MariTerm and ItalWordNet Renewal Through Mutual Expansion and Plug-In Links

Open Access
|Jul 2026

Full Article

1 Context and Motivation

1.1 The Challenge of Legacy Lexical Resources

Lexical-semantic databases developed during the digital humanities’ formative decades represent invaluable repositories of expert linguistic knowledge. Yet, many face obsolescence as technological standards evolve, funding cycles end, and original development teams disperse. ItalWordNet (Roventini et al., 1998) and MariTerm (Marinelli & Spadoni, 2007), both created by CNR-ILC in the late 1990s and regularly expanded during the early 2000s, exemplify this situation. Despite containing rich semantic information, both resources existed in outdated XML schemas which are incompatible with contemporary standards. Moreover, they lacked systematic cross-referencing despite theoretical frameworks for it, and were not available through modern research infrastructures.

The urgency of addressing these legacy resources stems from multiple factors. First, the specialized domain knowledge they encode, particularly in MariTerm’s case, developed in collaboration with the Port of Livorno, cannot be easily reconstructed without substantial expert input and institutional partnerships. Second, both resources were developed with overlapping conceptual coverage but remained largely disconnected, representing duplicated effort that could be leveraged through systematic linking. Third, neither fully complied with FAIR principles (Wilkinson et al., 2016), limiting their discoverability, accessibility, interoperability, and reusability for contemporary research.

1.2 Opportunities in Resource Integration

Previous successful efforts in linking WordNet resources to specialized ontologies (Frontini et al., 2016; Bizzoni et al., 2014) demonstrated the value of cross-resource integration. However, these projects typically involved creating new resources aligned with existing standards rather than rehabilitating legacy databases. Our work addresses the distinct challenge of reviving existing resources while preserving their accumulated knowledge and simultaneously enhancing both through mutual expansion.

The maritime domain presented opportunities for this approach. Maritime terminology encompasses multiple specialized subdomains while maintaining connections to general vocabulary. MariTerm’s specialized depth combined with IWN’s broad coverage suggested that bidirectional enrichment could serve both specialized users (maritime professionals, translators, domain experts) and general-purpose applications requiring maritime terminology coverage.

1.3 Project Objectives

This work pursued three interconnected objectives, the first being resource renewal. We wanted to transform MariTerm from obsolete formats into FAIR-compliant databases suitable for long-term preservation and contemporary research infrastructure deployment. Our second aim was expansion, specifically by enriching IWN with maritime-specific semantic relations and concepts from MariTerm. At the same time, we strived to update MariTerm with systematic connections to general Italian vocabulary. Finally, our third objective was linking, i.e., establishing systematic cross-resource references enabling integrated consultation and exploitation of both databases.

Critically, we aimed to achieve these objectives through reuse-oriented methodologies that could inform similar legacy resource rehabilitation efforts, balancing automated processing efficiency with quality assurance through expert validation.

1.4 Related work

Merging and aligning lexical resources represents an active research area spanning multiple languages and domains. McCrae et al. (2012) developed the lemon model (later OntoLex-Lemon, McCrae et al., 2017) specifically to enable interoperability between lexical resources and ontologies, addressing the structural heterogeneity we encountered. Bond and Foster (2013) coordinated the Open Multilingual Wordnet initiative, linking wordnets across 26 languages through shared senses among Wiktionary and WordNet, demonstrating the value of systematic cross-resource connections.

Domain-specific integration efforts offer methodological precedents for our work. Frontini et al. (2016) linked the GeoNames ontology to WordNet, enriching general lexical resources with geographical domain knowledge through automated alignment and manual validation, an approach parallel to our maritime domain expansion. Bizzoni et al. (2014) created Ancient Greek WordNet by integrating multiple classical language resources, addressing challenges of historical language data comparable to our legacy resource rehabilitation. Navigli and Ponzetto (2012) linked Wikipedia pages to WordNet for the creation of BabelNet, demonstrating large-scale integration benefits despite substantial structural differences between resources.

Legacy resource renewal specifically has received less systematic attention. Most rehabilitation efforts focus on format conversion (for ItalWordNet, Bartolini, 2016) rather than content expansion during conversion. Our approach differs by prioritizing mutual enrichment before format migration, leveraging existing structures to preserve and enhance linguistic content that might otherwise be lost during conversion overhead.

Alignment methodologies vary substantially. Matuschek and Gurevych (2013) used a graph-based algorithm for cross-lingual word sense alignment, while Navigli (2009) employed graph-based methods for sense disambiguation across resources. Methods for measuring similarity currently include comparing neural network-based word embeddings (Macilenti et al., 2025; Zhang et al., 2018), with some of them being applied to terms within the same ontology (Duong et al., 2018). However, for the scope of our work, we experimented a TF-IDF (Term Frequency-Inverse Document Frequency) approach which, though simpler to other state-of-the-art methods, proved effective for Italian lexicographic definitions due to terminology redundancy of both databases, offering a lightweight alternative when distributional data is unavailable. The weighted relation similarity framework builds on the work by Tülü et al. (2019), who quantified semantic relation importance for graph-based WordNet operations.

This work contributes to legacy resource rehabilitation by demonstrating that mutual expansion can enhance both specialized and general resources simultaneously, creating value exceeding simple format modernization while establishing a foundation for eventual conversion to contemporary standards.

2 Description of the Dataset

Repository Location: https://hdl.handle.net/20.500.11752/OPEN-1034

Repository Name: ILC4CLARIN (CNR-ILC repository within CLARIN infrastructure)

Object Name: MariTerm 1.2.

Original Format Names and Versions: Custom XML schema based on EuroWordNet model, non-compliant with Global WordNet standards

Creation Dates: 2004–2007, updated 2025

Dataset Creators: Rita Marinelli (ILC-CNR), Giovanni Spadoni (S. Spadoni s.r.l. Shipping Agency)

Languages: Italian (lemmas, definitions, metadata); English (some definitions, InterLingual Index alignment references)

License: Creative Commons–NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)

Publication Date: 2025-01-09

2.1 More about the Resources

In addition to the description provided above, Table 1 provides a breakdown of the data contained in ItalWordNet (IWN) and MariTerm used in our work, prior to any reuse or redeposition.

Table 1

Content breakdown of IWN and MariTerm.

RESOURCESYNSETSLEMMASINTERNAL SEMANTIC RELATIONS
ItalWordNet49,12145,754131,830
MariTerm2,5053,2756,442

2.1.1 MariTerm

Developed as a WordNet-like lexical database mapping maritime terminology to domain-specific concepts, MariTerm was created through collaboration between CNR-ILC computational linguists and Port of Livorno domain experts. The original resource structure supported cross-resource linking through plug-in mechanisms, though these required population through integration work. The resource’s specialized knowledge encompassed two primary subdomains: technical/nautical terminology, including ship components, navigation equipment, maritime operations, nautical measurements, and maritime transport domain, with synsets dedicated to port operations, cargo handling, shipping logistics, and commercial classifications.

The resource’s value lay not merely in terminology enumeration but in expert-encoded semantic relations reflecting domain conceptual structures.

2.1.2 A note on ItalWordNet

ItalWordNet (Roventini et al., 1998) originated from the EuroWordNet project (Vossen, 1998) and the Italian national SI-TAL initiative, creating a lexical database organizing Italian words into synsets. The version used for this work uses an adapted schema created in the 1990’s by the ILC to allow for interoperability among resources.1 Its representation of content is adapted from EuroWordNet, with synsets contained within “WORD_MEANING” tags and aligned to Princeton WordNet via the InterLingual Index (ILI) rather than through direct mapping. Semantic relations included standard WordNet types (hyperonymy, hyponymy, meronymy, etc.) plus EuroWordNet-specific relations like cross-part-of-speech near-synonymy variants. Albeit the version of IWN described in this work is currently not available on ILC4CLARIN, it is available in the GitHub repository listed as Supplementary File 1, alongside the original MariTerm version sharing its XML schema.2

3 Reuse Process and Methodology

3.1 Strategic Approach: Mutual Enrichment

Rather than simply converting legacy formats to modern standards, we adopted a mutual enrichment strategy as the primary rehabilitation focus. IWN provided broad lexical coverage and integration with general Italian vocabulary, while MariTerm offered domain expertise and specialized semantic granularity. The expansion process flowed bidirectionally:

  • MariTerm → IWN: Adding missing maritime concepts, specialized semantic relations

  • IWN <> MariTerm: Systematic links to shared vocabulary, integration with broader lexical network

This approach transformed simple format conversion into knowledge integration, increasing both resources’ value beyond their original states.

3.2 Pipeline Architecture

3.2.1 Candidate Extraction and Structural Alignment

The process began with synset alignment based on lemma matching. We adopted a conservative matching strategy: synsets aligned only when their first-listed lemmas matched exactly. This prioritized precision over recall, reducing false positives from coincidental homonyms. From complete resource extracts, this yielded 1.157 preliminary candidate synset pairs, instances where both resources contained synsets beginning with identical lemmas. These candidates underwent further filtering through similarity scoring.

3.2.2 Similarity Scoring and Match Validation

We employed TF-IDF vectorization applied to synset and target word definitions, followed by cosine similarity comparison. While TF-IDF is typically applied to longer documents rather than short lexicographic definitions, we found it effective in this context for two reasons. First, the substantial redundancy of terminology across synset definitions, with both resources repeatedly referencing core concepts like “nave” (ship), “mare” (sea), “porto” (port), created sufficient term distribution variation for TF-IDF weighting to function meaningfully. Second, the computational efficiency of TF-IDF allowed us to process approximately 200.000 potential definition comparisons within practical time constraints, whereas more sophisticated semantic similarity measures (word embeddings, sentence transformers) would have required substantially greater computational resources. Besides, our approach allowed us to compare matching definitions even when these were provided in specialized and general language. By contrast, methods rewarding words pairs co-occurring often like Positive PMI (PPMI; as used by Melamud et al., 2015) would have not allowed for the same level of granularity.

As to the technical implementation, we used Python’s scikit-learn TfidfVectorizer with default parameters (lowercase conversion, tokenization on word boundaries, no stop word removal to preserve maritime domain terminology, L2 normalization). Each gloss was treated as a separate document, yielding sparse TF-IDF vectors that were compared using cosine similarity.

The scoring framework combined two components: direct definition similarity and weighted relation similarity. The former was obtained by comparing glosses of synsets sharing lemmas across resources, while the latter was obtained by comparing definitions of semantically related target synsets, multiplied by relation-specific weights. For the scope of this work, we only considered hyponymy, hyperonymy, near synonymy and their variants involving different part of speech.3 Weights reflected semantic relation importance based on Tülü et al. (2019), which are reported in Table 2.

Table 2

Breakdown of weights adopted for calculating synset similarity (from Tülü et al., 2019).

SEMANTIC RELATIONASSIGNED WEIGHT
Hyperonymy0.82
Hyponymy0.7
Near synonymy0.6
Other relations specific to EuroWordNet0.5

Combined scores rewarded shared attributes while penalizing discrepancies. With the help of this new system, we managed to identify a total of 1.267 matches. We also established a high-confidence threshold through iterative validation. Beginning with an initial threshold of 0.40 (cosine similarity for gloss comparison), we manually reviewed 100 randomly sampled candidate pairs. All pairs scoring above 0.43 showed correct alignment, while pairs between 0.30-0.43 required additional relation similarity support. We therefore set 0.43 as the primary threshold for gloss-based matching, accepting lower scores (minimum 0.13) when weighted relation similarity exceeded 0.09. This graduated threshold approach balanced avoiding false matches with capturing valid pairs with divergent but synonymous definitions).

3.2.3 IWN Expansion and Synset Creation

For the 747 validated matches, the expansion process performed four main operations:

  1. Adding missing relations: When MariTerm contained semantic relations (e.g., hyperonymy links) whose target synsets were absent from the corresponding ItalWordNet synset, we added these relations to ItalWordNet. This created 1.160 new relation instances connecting ItalWordNet synsets to maritime-specific concepts.

  2. Creating new synsets: When added relations pointed to concepts not yet present in ItalWordNet (e.g., specialized maritime equipment or nautical measurements), we created new synsets with fresh numerical identifiers. Definitions were inherited from MariTerm, and appropriate structural metadata (part of speech, lexical domain) was added. This generated 363 new ItalWordNet synsets, expanding maritime terminology coverage.

  3. Managing definitions through fallback mechanisms: The pipeline addressed missing definitions by searching a synset’s semantic relations in predefined priority order (near-synonyms first, then hyperonyms, then hyponyms), inheriting the first available gloss. For conflicting definitions, ItalWordNet glosses took precedence over MariTerm’s, except when ItalWordNet entries lacked definitions entirely.

  4. Ensuring reciprocal consistency: WordNet resources require bidirectional relations (if synset A has hyponymy to B, B must have hyperonymy to A). The pipeline automatically generated these inverse relations, maintaining structural coherence across both resources.

Out of 747 synsets selected for expansion, 397 IWN synsets directly updated (53.1% of validated matches). The remainder required no changes as they already contained equivalent relations, demonstrating partial pre-existing overlap.

3.2.4 Manual Validation and Correction

To check matches’ quality, we conducted a manual review of 400 updated synsets. This revealed 20 incorrectly aligned synsets requiring correction (5% error rate), ~100 new synsets with identifier conflicts requiring manual renumbering and 241 synsets with internal gloss inconsistencies, where definitions were present for target references but missing from main entries. A dedicated Python script resolved the gloss inconsistencies through systematic propagation. The manual corrections resulted in refined databases with improved internal consistency.

3.2.5 Cross-Resource Linking

Plug-in relations are cross-resource linking mechanisms designed in the EuroWordNet framework to connect synsets across different wordnets without merging them into a single database. Unlike internal semantic relations (hyperonymy, meronymy) that link synsets within the same resource, plug-in relations establish correspondences between equivalent or related concepts across separate lexical databases, enabling users to traverse from concepts in one resource to their counterparts in another. For example, a MariTerm synset for “veliero” (sailing ship) might have a plug-synonym relation linking to the general ItalWordNet synset for the same concept, allowing navigation from specialized maritime terminology to broader Italian vocabulary networks.

The plug-in relation mechanisms described in MariTerm’s original documentation were absent in the XML structure, thus requiring population through integration work. Consequently, nodes were automatically restored and populated for both WordNets during this phase via our Python toolkit. The most represented types of plug-in relations across both resources were:

  • Plug-synonym relations, connecting overlapping synsets with similar meanings across resources. The deposited MariTerm version now contains 759 links to synsets inside IWN, and 751 links in the opposite direction.

  • Plug-relation links: connecting semantic relations added to IWN back to source MariTerm synsets, with 843 relation-level links from IWN to MariTerm. Unfortunately, the process was unsuccessful for the opposite direction, with MariTerm shared relations being not mapped correctly to IWN.

4 Outcomes and Experience

4.1 Quantitative Results

Results of the expansion and linking process are illustrated in Table 3, Table 4 instead provides a summary of content inside the redeposited MariTerm available now on ILC4CLARIN.

Table 3

Summary of results.

DATA INVOLVEDCOUNT
Synset matches747
MariTerm>IWN expansion
IWN Word meanings updated397
New synsets created in IWN363
New relations in IWN1.160
IWN<>MariTerm Linking
Mapping nodes in IWN pointing to MariTerm content759
IWN synsets linked to MariTerm via plug-synonymy742
New IWN relations linked to MariTerm843
Mapping nodes in MariTerm pointing to IWN751
MariTerm synsets linked to IWN via plug-synonymy1.085
MariTerm relations linked to IWN0
Table 4

Summary of content of the deposited, updated MariTerm.

SYNSETSLEMMASSYNSET RELATIONS
2.5323.3216.442

We also provide more insight on the overall performance of our system. From the updated list of 1.267 candidates, this algorithm identified 747 high-confidence matches. Manual review of the matches revealed 20 incorrect alignments, 10 being false positives, and the other half being false negatives. We also examined a random sample of 104 out of the 520 rejected matches (accounting for 20% of synset couples that did not pass our filters). Since the sample yielded 11 valid matched missed by our thresholds, we estimated that our system has skipped around 55 correct synset matches. All of this suggests an F1-score of 0.954 for the whole similarity-based matching algorithm. However, discussion remains open regarding a particular category of synsets that were not matched, but that could be considered false negatives. Namely, our algorithm rejected 242 matches on shared synsets on named entities (harbour cities and maritime locations). In all odds, this was caused by the lack of gloss and semantic relations in the dedicated synsets both in MariTerm and IWN. Given that the mismatch happened due to structural issues related to data, we chose to consider these instances as true negatives. In that case, that would still achieve an F1 score of 82%. Table 5 provides a full breakdown of F-scores, with precision and recall for both cases.

Table 5

Summary of content precision, recall and F1 scores by scenario.

SCENARIOTRUE POSITIVESTRUE NEGATIVESFALSE POSITIVESFALSE NEGATIVESPRECISIONRECALLF1-SCORE
Not counting named entities73345514650.980.920.95
Counting named entities733206143140.980.700.82

4.2 Challenges

The expansion and linking process exhibited multiple challenges, which are all outlined in the following Sections.

4.2.1 Format Non-Compliance

Neither resource conformed to contemporary standards like Global WordNet Association schemas or OntoLex-Lemon, limiting interoperability with modern NLP tools and linked open data infrastructures. The decision to defer format conversion while prioritizing content enrichment reflected both time constraints and strategic considerations.

First, the conversion process from legacy EuroWordNet XML to OntoLex-Lemon RDF was beyond the scope of this work, a single internship project. The institutional priorities at CNR-ILC focused on content-level rehabilitation, such as rescuing and enhancing the linguistic knowledge encoded in these resources, rather than format migration, which could be undertaken as separate future work. Second, performing mutual expansion within the existing XML format allowed us to leverage existing parsing tools, validation scripts, and institutional knowledge of the schema, reducing development time and error risk. Third, restoring the plug-in relation mechanisms in the MariTerm XML structure provided a viable framework for cross-resource linking; recreating equivalent functionality in a new format would have required additional design and implementation work.

We recognize that full OntoLex-Lemon conversion would significantly enhance reusability for modern NLP pipelines and LOD integration. However, achieving FAIR compliance through structured metadata, persistent identifiers, and CLARIN repository deposition represents substantial progress from the resources’ state of neglect. The enhanced XML resources now serve as a stable foundation for future format migration, with enriched content that would have been lost if conversion had been prioritized over expansion. As Bartolini (2016) demonstrated for the original ItalWordNet, conversion to RDF-based formats is technically feasible and represents a clear next step for these expanded resources.

4.2.2 Structural Archaeology

Working with legacy data required extensive “archaeological” interpretation. In our case, documentation did not match implementation details: while MariTerm’s introductory paper (Marinelli and Spadoni, 2007) pointed to the presence of plug-in relation nodes, these were absent in the resource before any processing. Systematic population of these connections thus required the integration work described here, as well as careful interpretation of the intended functionality from original documentation.

4.2.3 Semantic Divergence

Despite conceptual overlap (e.g., both containing synsets for ‘navigare’—to sail/navigate), different numerical identifiers prevented direct alignment. Even when lemma and sense attributes matched, semantic relations often diverged, suggesting independent development paths. Figures 1, 2 and 3 illustrate the IWN synset for “ormeggio” (EN: moorage, place to moor), the first MariTerm equivalent detected by lemma and sense number matching (mooring rope) and the correct MariTerm synset match, obtained by our similarity framework (moorage, place to moor).

Figure 1

Synset for “ormeggio” (moorage, place of) – IWN.

Figure 2

Synset for “ormeggio” (mooring rope, match based on lemma and sense number) – MariTerm.

Figure 3

MariTerm equivalent for “ormeggio” (moorage, place of) - match from similarity calculation.

This reflected independent development: IWN evolved from European lexicographic traditions emphasizing general language; MariTerm developed through port authority collaboration emphasizing operational maritime contexts. Bridging these discrepancies required interpretation beyond simple structural matching. Moreover, both resources frequently included multiple synonymous lemmas within single synsets (listed under the ‘LITERAL’ tag in the ‘VARIANTS’), complicating automated matching and potentially causing redundant alignments.

4.2.4 Missing Metadata and Other Synsets Issues

As briefly mentioned in 3.2.3, we also had to account for definition inconsistencies in both WordNets. As a matter of fact, matched synset pairs frequently displayed different glosses or missing definitions. For example, ‘pagaia’ (paddle) definitions varied significantly, posing a serious issue as to the definition to use whenever creating new synsets from scratch, when missing:

  • IWN: “wide-bladed oar with which one rows without resting it on the gunwale”

  • MariTerm: “paddle with blade-like ends handled by holding it in the middle with both hands—used on canoes and other watercraft of the river or bathing type

In other cases, one resource provided definitions while the other had none (e.g., ‘azimut’/azimuth: MariTerm defined it as “in astronomical navigation, the arc of a circle between north and the vertical of the star itself” while IWN’s entry was empty), thus increasing the risk of losing crucial information for all our operations.

4.2.5 The Quality-Automation Tension

The 5% error rate in automated alignment (20 incorrect matches among 400 reviewed) demonstrates the inevitable tension between processing efficiency and accuracy. Higher similarity thresholds would reduce errors but also eliminate valid matches with unusual definitions. Our threshold settings prioritized recall for subsequent manual review over pure precision.

The 241 gloss inconsistencies (new synsets with definitions for relation targets but not main entries) revealed that seemingly simple operations like inheriting definition from relation target can create new inconsistencies when applied systematically. This argues for iterative validation cycles rather than single-pass automated processing.

4.2.6 Duplicate Synsets and Asymmetric Bidirectionality

A short inquiry before linking revealed that the version of IWN used for our work contained 803 synsets that were duplicates of other entries. In hindsight, this discovery might explain a series of asymmetries emerged during mapping and linking. As a matter of fact, out of the 747 matched synsets across both resources, our code mapped 759 synsets from IWN to its MariTerm counterparts, with 751 in the opposite direction. Albeit only involving around 1.6% and 0.5% of matched synsets, this calls for refining logic during the linking stage and checking which synsets are getting their plug-in node incorrectly populated.

Duplicates may also be at the core of inconsistent mapping of relations. As a matter of fact, some plug-in nodes inside MariTerm contain more than one plug-in synonym relation for a total of 1,085 plug-in synonymy relation. Ideally, each plug-in node should contain one such relation instead. The IWN version used in this work, instead, lists less plug-synonymy relations than matched synsets, which may be caused by some mapping nodes only containing other shared relations.

4.3 Benefits of Mutual Expansion

Despite challenges, the mutual enrichment approach delivered tangible benefits. Enhanced specialization depth is the first natural consequence, with IWN now including maritime-specific conceptual distinctions previously absent, thus reflecting both general and domain standards. A second benefit coming from our work is a better contextualization of terminology. MariTerm terms now systematically connect to general Italian vocabulary hierarchies, enabling users to situate specialized concepts within broader semantic networks. Thirdly, another significant benefit lies in the preservation of domain expertise embodied in MariTerm’s relation structure. This knowledge, encoded through collaboration with maritime professionals, cannot be automatically reconstructed from textual corpora and would be lost if MariTerm remained isolated and deprecated. Moving on to the mapping, the plug-in nodes now allow users to navigate from specialized maritime concepts to general Italian vocabulary and back, supporting both domain experts needing general context and general users encountering maritime terminology. Finally, MariTerm’s availability through CLARIN ensures long-term accessibility for Italian NLP research, multilingual lexicography, specialized translation, and domain-specific applications like automatic text classification of maritime documents.

4.4 FAIR Compliance and Redeposition

Following expansion and linking, we provide a brief self-assessment of FAIR compliance for the redeposited MariTerm:

Findability: Assignment of persistent identifiers (Handle system IDs) and deposition in CLARIN-accessible ILC4CLARIN repository with comprehensive metadata

Accessibility: Open licensing (CC BY-NC-SA 4.0) and provision through standard repository protocols

Interoperability: While full conversion to Global WordNet or OntoLex-Lemon formats remains future work, the enhanced XML includes explicit cross-resource references enabling programmatic traversal

Reusability: Detailed documentation of expansion methodology, provenance information for new relations, and clear indication of source resources for integrated content

5 Recommendations and Good Practices

5.1 Methodological Recommendations

Based on our experience, we propose several recommendations for similar legacy resource rehabilitation endeavours. First and foremost, when multiple related legacy resources exist, it is worth investigating opportunities for cross-resource expansion rather than treating each independently. Even resources with apparent overlap may encode complementary knowledge, where one’s weaknesses often correspond to another’s strengths: in our case, IWN’s gaps in maritime specialization were precisely the areas where MariTerm excelled, and vice versa.

Preserving source fidelity throughout processing is equally important. Original resource structures should be retained and transformations applied to copies, with all modifications documented explicitly. This enables later verification that no information was inadvertently lost, a concern particularly relevant when working with resources whose documentation does not fully reflect the actual data. Finally, it is worth designing for incremental improvement from the outset. Perfect integration is unattainable in a single pass, and asymmetric linking patterns, like those we observed, represent not failures but natural consequences of resource scope differences. Documenting known limitations clearly guides future enhancement while still delivering substantial value from partial integration. Along the same lines, full conversion to current standards (Global WordNet Association formats, OntoLex-Lemon) represents important future work that would maximize interoperability with modern NLP infrastructures. However, prioritizing content enrichment over format migration in this initial rehabilitation phase enabled us to leverage existing institutional expertise and XML-based tools while establishing FAIR compliance through repository deposition. Format conversion can be undertaken more efficiently once content stabilizes. We prioritised FAIR compliance and repository deposition while deferring complete format migration, balancing immediate accessibility with development flexibility.

5.2 Technical Good Practices

On the technical side, our experience offers several lessons worth generalising. TF-IDF performed well for scoring despite being designed for longer documents, largely because definition redundancy across synsets created sufficient term distribution variation. For resources with sparser or more unique definitions, more sophisticated semantic similarity measures, such as word embeddings or sentence transformers, may be preferable. Whatever metric is adopted, relation-specific weights proved crucial for accurate matching: deriving them from existing research (Tülü et al., 2019) provided empirical grounding, though manual assignment was still necessary for EuroWordNet-specific relations not covered in that work. The rationale behind such choices should be documented explicitly to support reproducibility.

Adopting fallback mechanisms for missing glosses was indispensable in our case but required careful design to prevent error propagation: inheritance from hyperonyms should take priority over hyponyms, and circular references must be explicitly avoided. More broadly, when legacy resources already include cross-linking structures, even if unpopulated, as was the case with MariTerm’s plug-in nodes, leveraging these existing frameworks is preferable to introducing new mechanisms. This preserves compatibility with the original resource documentation and paves the way for improving interoperability with existing tools in the future.

5.3 Organizational and Policy Recommendations

Two organizational lessons stand out from this project. The first concerns documentation as an active form of preservation. Detailed methodology descriptions serve a dual purpose: they enable reproducibility and create knowledge artefacts that outlast individual projects, allowing others to apply analogous approaches to different resource pairs. The second is institutional memory and sustainability. Legacy resources often languish precisely because original developers moved on and institutional knowledge faded with them. Our work was only possible because CNR-ILC maintained both resources and the expertise to interpret them. Establishing succession planning and thorough documentation practices from the outset of any resource development project is the most effective safeguard against future obsolescence.

5.4 Future-Oriented Considerations

Several directions remain open for both resources specifically and for the methodology more broadly. While our enhanced resources remain in XML, conversion to RDF and OntoLex-Lemon formats would enable full integration into Linked Open Data infrastructures. Quochi et al. (to appear) demonstrated a viable path for IWN through the Open Multilingual Wordnet; applying the same approach to the expanded version would extend international interoperability significantly. The maritime domain concepts identified in MariTerm could also be elaborated into a formal domain ontology following the methodology proposed by Niero (2006), whereby new synsets inherit hyperonymy from the general resource and hyponymy and synonymy from the specialised database.

Closer to the current work, the zero count for MariTerm relation links to IWN, the duplication of plug-synonyms inside of MariTerm’s mapping nodes and incorrect mappings remain open issues, addressable through systematic review of the code logic adopted for our work. Once implementation issues are fixed, operation can be expanded to other semantic relations such as meronymy and holonymy. Matches in MariTerm could even undergo a double validation process by repeating the process also future iterations of IWN, removing links to deprecated outdated terms. And beyond the maritime domain, our methodology could be transferable to other specialised terminological resources, paving the way to better, yet time-efficient management of general and specialised Italian lexicography.

All in all, our work demonstrates that legacy lexical resources need not to face obsolescence. Through strategic mutual expansion and cross-resource linking, we transformed two Italian lexical databases, representing decades of lexicographic work and irreplaceable domain expertise, into FAIR-compliant, interlinked assets now preserved in CLARIN infrastructure. The challenges encountered are not unique to our resources but characteristic of legacy data across the digital humanities. Our graduated automation approach, combining algorithmic efficiency with strategic manual validation, offers a replicable methodology for similar rehabilitation efforts. Rehabilitation is achievable and valuable, but requires realistic expectations about automation limitations, sustained investment in quality assurance, and tolerance for incremental progress. The prospect of letting accumulated scholarly knowledge decay into inaccessibility makes the effort worthwhile.

Supplementary File 1

Title

MT2IWN – MariTerm to ItalWordNet Integration Toolkit.

Description

Python toolkit implementing the complete lexical integration pipeline described in Section 3. Includes modular source code for candidate extraction, similarity scoring, and XML integration, command-line tools for reproducing all analyses and comprehensive documentation with usage examples.

The toolkit processes XML-encoded MariTerm and ItalWordNet resources to identify shared lexical entries, calculate semantic similarity across glosses and relations, filter candidates using multi-threshold criteria, and generate bidirectional plugin links.

Other

Available at: https://github.com/CoPhi/mt2iwn.

License: MIT.

DOI: https://doi.org/10.5281/zenodo.18788538

Notes

[1] MariTerm employed the same XML representation schema as this version of IWN, facilitating structural compatibility.

[2] Supplementary File 1 also contains every iteration of the original files through each step of our work to better document changes.

[3] Namely “has_xpos_hyponym”, “has_xpos_hyperonym”, “xpos_near_synonym”.

[4] Scripts for calculating F1-score were added to the GitHub repository after the first review round. The metric was calculated via the scikit-learn library.

Acknowledgements

The authors thank the Port of Livorno Authority for the original collaboration that enabled MariTerm’s creation, the H2IOSC Project (Humanities and Cultural Heritage Italian Open Science Cloud), and ILC4CLARIN (Certified CLARIN B-Center) for supporting the resource renewal and redeposition work.

Author Contributions

Lucia Galiero: Conceptualization, Data curation, Formal Analysis, Investigation, Methodology, Software, Validation, Writing – original draft.

Federico Boschetti: Conceptualization, Methodology, Software, Supervision.

Riccardo Del Gratta: Resources, Software, Supervision, Writing – review & editing.

Angelo Mario Del Grosso: Methodology, Resources, Supervision, Writing – review & editing.

Monica Monachini: Funding acquisition, Project administration, Resources, Supervision, Writing – review & editing.

DOI: https://doi.org/10.5334/johd.528 | Journal eISSN: 2059-481X
Language: English
Page range: 92 - 92
Submitted on: Feb 27, 2026
Accepted on: May 27, 2026
Published on: Jul 10, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Lucia Galiero, Federico Boschetti, Riccardo Del Gratta, Angelo Mario Del Grosso, Monica Monachini, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.