Skip to main content
Have a personal or library account? Click to login
The Third Option: integrating AI training‑data licensing into scholarly publishing workflows Cover

The Third Option: integrating AI training‑data licensing into scholarly publishing workflows

By:   
Open Access
|Jul 2026

Full Article

Introduction

The expansion of large language models has increased demand for high‑quality training corpora, including specialist scientific and professional text. Scholarly articles are attractive inputs because they are curated, structured and produced within an accountability system that includes peer review, editorial oversight and post‑publication correction. Publishers and scholarly societies face increasing pressure to clarify how scholarly works are used in model development.

Two coupled failure modes have become visible. First, agreements involving large collections of scholarly texts can become contested when authors and scholarly communities perceive insufficient transparency, weak participation or unclear value distribution (National Communication Association, 2024). Second, attempts to restore legitimacy after the fact by contacting authors of previously published works can be difficult to scale because they require identity resolution, outreach, response handling and long‑run record‑keeping. These co‑ordination and administrative burdens are transaction costs in the sense of transaction cost economics (Coase, 1937; Williamson, 2010). These tensions resemble social licence failures identified in other data‑intensive governance settings, where legality alone is insufficient to sustain trust‑based legitimacy (Carter et al., 2015; Muller et al., 2021; Shaw et al., 2020).

This article proposes a forward‑looking alternative. Instead of unilateral back‑catalogue action or retroactive outreach, journals can embed a standardized opt‑in choice at the contracting stage after editorial acceptance. The author labels this institutional design the ‘Third Option’: a post‑acceptance, opt‑in addendum in which authors authorize defined artificial intelligence (AI) training uses of the accepted article in exchange for an author‑facing benefit. The benefit is not conceptually restricted to article processing charge (APC) discounts. In diamond open access (OA) contexts, it can be implemented via fee‑neutral benefits, institutional rebates or community reinvestment commitments consistent with diamond sustainability models (Dufour et al., 2023).

The central claim is not that a new permission mechanism is always legally necessary, particularly for Creative Commons Attribution (CC BY) works. Rather, the Third Option is a governance and workflow design intended to reduce transaction costs, strengthen procedural legitimacy through editorial separation and support compliance‑oriented dataset products through provenance, documentation and rights hygiene. Publisher communications suggest that AI‑related content licensing is economically salient even as governance remains contested (Wiley, 2025). This article sets out the mechanism and a methods‑style evaluation protocol for future empirical testing, with key objections addressed in a separate section.

Insights’ readership includes librarians, publishers, learned societies and infrastructure providers working on scholarly communication. For these stakeholders, the practical problem is not only whether AI training‑data use is legally permitted, but how consent, transparency and benefit allocation can be operationalized at scale without undermining trust or equity. The Third Option is offered as a workflow‑level governance design and a testable evaluation protocol that journals, platforms or consortia can pilot.

Practical steps for journals (an editorial checklist) include:

  • collect opt in only after acceptance and keep consent status operationally separate from editorial decision‑making

  • use a plain‑language disclosure of scope, terms and withdrawal limits and avoid pre‑checked boxes

  • treat the addendum as a governance package including provenance metadata, rights hygiene, dataset documentation, versioning and periodic aggregate transparency reporting.

A worked example of these steps within a generic journal workflow can be found in the section titled ‘The Third Option model’.

Debate map: recurring positions and the institutional deadlock

Table 1 summarizes four common stances in debates about scholarly content and AI training.

Table 1

Debate map (summary of positions in current debates)

PositionSummary (core claim and implications)
Unilateral back‑catalogue licensingEfficient for rightsholders but can provoke backlash if communities perceive weak transparency, weak participation or unclear benefit allocation (National Communication Association, 2024).
Retroactive opt‑in outreachImproves procedural legitimacy by seeking author decisions, but it is costly and difficult to scale, particularly for large back catalogues (Coase, 1937; Williamson, 2010).
‘Open licences already solve it’When works are CC BY, permission exists for broad reuse, so additional licensing can be redundant. This stance often under‑specifies provenance, documentation and auditability for curated dataset supply (Bender & Friedman, 2018; Gebru et al., 2021).
‘Scraping and legal exceptions dominate’Developers can scrape public text or rely on text and data mining exceptions, reducing willingness to pay. This stance may underestimate compliance incentives and the value of curated datasets with documentation and assurances (European Union, 2019, 2024; Mitchell et al., 2019).

The deadlock persists because the sector is simultaneously negotiating legitimacy (who should have agency, and what constitutes fair process) and operational scale (how governance can be implemented at a low marginal cost per article). The deadlock is therefore institutional rather than merely legal: unilateral licensing can satisfy scale but weaken perceived legitimacy, while retroactive consent can strengthen legitimacy but often fails to scale. The Third Option addresses these jointly by relocating governance to the standard contracting stage and by treating documentation and provenance as part of the dataset product definition, not as optional add‑ons.

Definitions and scope

As detailed in Table 2, this article uses stipulated operational definitions (Layer A) to enable precise argumentation. Representative anchors are provided where they help link the proposal to the established literature.

Table 2

Definitions (Layer A operational definitions with representative anchors)

TermLayer A: operational definition (this article)Representative anchorsBoundary conditionsImplications
Third OptionA post‑acceptance, opt‑in addendum integrated into the publication agreement, authorizing defined AI training uses of the accepted article in exchange for an author‑facing benefit.Coase (1937); Williamson (2010)Consent requested only after acceptance; does not alter peer review.Converts consent into a routine workflow artefact rather than a retroactive campaign.
AI training useUse of article text (and optionally tables) as input to train, fine‑tune or evaluate machine‑learning models, distinct from the redistribution of the article as an access substitute.Bender and Friedman (2018); Mitchell et al. (2019)Does not resolve all downstream misuse risks.Enables permitted and prohibited use boundaries to be stated and audited.
Author‑facing benefitA tangible benefit offered for opting in that is monetary or fee‑neutral, depending on the publishing model.Dufour et al. (2023); Smith et al. (2021)Must be independent of acceptance decisions; design must address inducement risk.Supports uptake but can also threaten legitimacy without safeguards.
Social licenceStakeholder‑perceived legitimacy grounded in trust, transparency and perceived public value, beyond legal permission.Carter et al. (2015); Muller et al. (2021)Contextual and contested; not a legal standard.Motivates transparency reporting and procedural safeguards.
Provenance recordMachine‑readable metadata linking content to persistent identifiers, version, date and inclusion parameters, enabling traceability and audit.Buneman and Tan (2019)Provenance does not guarantee accuracy of content.Supports traceability, correction and audit narratives.
Withdrawal (future‑only)Exclusion of an article from future dataset releases after opt in is withdrawn, with disclosure that prior training may not be reversible.Buneman and Tan (2019); Xu et al. (2023)Cannot guarantee model untraining; cannot revoke open licences.Provides an operationally coherent agency mechanism tied to versioning.

The scope boundaries are as follows:

  1. This article does not offer legal advice and does not claim that any single mechanism is legally required across jurisdictions.

  2. The proposal is forward‑looking and does not solve governance for all legacy back‑catalogue content.

  3. This article does not assume developers will pay for licensed datasets under all market conditions; willingness to pay is treated as empirically testable.

  4. The Third Option does not override open licences. Where works are CC BY, it functions primarily as a governance, packaging and assurance mechanism for curated dataset supply.

The Third Option model

The Third Option is implemented as a post‑acceptance addendum to the publication agreement. Authors may be informed at submission that an addendum may be offered after acceptance, but the opt‑in decision is requested only after acceptance to reduce perceived influence on editorial outcomes. The addendum specifies the scope of included components (for example, full text and tables by default; third‑party content excluded unless cleared), permitted and prohibited uses (training versus redistribution or access substitution), provenance and documentation obligations, optional terms and renewal and withdrawal semantics, where the latter is defined as future‑only exclusion from subsequent dataset releases.

The implementation protocol (numbered steps) is as follows:

  1. After editorial acceptance, present the standard publication agreement and, separately, a Third Option addendum with a plain‑language disclosure of scope, permitted uses, prohibited uses, terms and withdrawal limits.

  2. Record affirmative opt in, storing consent records operationally separate from editorial systems (editorial independence).

  3. Generate a machine‑readable provenance record linked to the article identifier and inclusion parameters (Buneman & Tan, 2019).

  4. Package consenting works into a curated dataset release accompanied by dataset documentation and rights‑hygiene rules (Gebru et al., 2021).

  5. Provide dataset releases to licensees under contractual assurances addressing permitted uses, security controls, auditability and onward‑transfer constraints (Mitchell et al., 2019).

  6. Maintain dataset versioning. If withdrawal occurs, exclude the work from future dataset releases and notify licensees of the updated version status, while disclosing the limits of reversing past training (Buneman & Tan, 2019; Xu et al., 2023).

  7. Publish periodic aggregate transparency reporting on uptake and benefit allocation, including equity indicators, while protecting author privacy (Muller et al., 2021).

Worked example

To illustrate the workflow, consider a journal that, once an article is accepted, sends the author the standard publication agreement together with a separate Third Option addendum. The author may decline the addendum without affecting publication. If the author opts in, consent is recorded outside the editorial system, the accepted text and tables are included by default, third‑party material is excluded unless separately cleared and provenance metadata are attached for the next curated dataset release. If the author later withdraws, the article is excluded from subsequent releases, and the change is reflected in version documentation and licensee notices.

Decision table: benefit structures by publishing model

The Third Option is defined by workflow‑integrated opt‑in governance, not by any single benefit type. Table 3 lists feasible benefit structures and constraints.

Table 3

Decision table (publishing model and feasible author‑facing benefit structures)

Publishing modelExample benefit structuresConstraint to state explicitly
Subscription‑only journalDirect author honorarium; society membership reduction; institutional credit; community reinvestment commitment.Benefit must not be tied to acceptance; access terms remain unchanged.
Hybrid journal (subscription + optional OA with APC)APC discount or rebate; waiver top‑up; institutional rebate; fee‑neutral benefit.Needs‑based waivers should be independent of consent; monitor disparate uptake (Smith et al., 2021).
Diamond open access journal (no APC, often CC BY)Fee‑neutral benefits (author services); institutional rebates; conference discounts; community reinvestment commitments aligned with diamond funding strategies.Addendum cannot restrict CC BY rights; value must be framed as governance, provenance and assurance (Dufour et al., 2023).
Full open access journal (APC‑funded)APC discount or rebate; waiver top‑up; institutional rebate.Equity risks may be amplified and require monitoring (Ellers et al., 2017; Smith et al., 2021).

Why the addendum can be valuable even when permission exists

A recurring objection is that, for open‑licence content, additional licensing adds nothing. The Third Option is designed to make a narrower and more operational claim: that many stakeholders value a compliance‑oriented dataset product defined by provenance, documentation and rights hygiene, not merely by permission. Documentation artefacts such as data statements, dataset datasheets and model reporting norms are treated as governance tools in machine learning (Bender & Friedman, 2018; Gebru et al., 2021; Mitchell et al., 2019). Provenance and versioning support traceability and auditing (Buneman & Tan, 2019).

This rationale is strengthened in jurisdictions where training‑data governance is explicitly regulated or where legal exceptions are coupled to rights‑reservation mechanisms. In the European Union, for example, the DSM Directive includes text and data mining exceptions, one of which allows rightsholders to reserve their rights in an appropriate manner, while the EU AI Act imposes parallel transparency and copyright‑compliance obligations on providers of general‑purpose AI models (European Union, 2019, 2024; Margoni & Kretschmer, 2022; Ziaja, 2024).

Safeguards for legitimacy and equity

Because the Third Option links opt in to a benefit, it creates an inducement risk. This risk is credible in fee‑based publishing systems, where APCs and waivers can affect participation and geographic diversity (Ellers et al., 2017; Smith et al., 2021). The ethical literature distinguishes between acceptable incentives and ‘undue inducement’ that pressures individuals to consent against their preferences due to resource constraints (McGregor, 2005). Accordingly, as detailed in Table 4, a minimal safeguard set is part of the design itself, not an optional feature.

Table 4

Safeguards (risks, safeguards and implementation notes)

RiskRecommended safeguardImplementation note
Perceived editorial influenceCollect consent only post‑acceptance; consent status not visible to editors or reviewers.Separate editorial and licensing workflows with audit logs.
Undue inducementOffer equivalent routes and needs‑based waivers independent of consent.Declining the addendum must not foreclose viable publication options (McGregor, 2005; Smith et al., 2021).
Information asymmetryStandardized, plain‑language disclosure of scope, terms and withdrawal limits.Avoid pre‑checked boxes; allow time for consideration where feasible.
Irreversibility misunderstandingDisclose limits of reversing prior training; define withdrawal as future‑only.Tie withdrawal to dataset versioning and licensee notifications (Xu et al., 2023).
Third‑party rights leakageExclude third‑party content by default unless separately cleared.Require author attestation and implement checks where feasible.
Opaque value distributionAggregate transparency reporting on uptake and benefit allocation.Report at journal or publisher level while protecting author privacy (Muller et al., 2021).

Strategic dynamics: path dependence and cumulative advantage

The Third Option has strategic implications for journals and publishers. If AI‑assisted discovery, summarization and recommendation tools become a stable layer of scholarly infrastructure, inclusion in well‑documented training corpora may function as a competitive differentiator for some stakeholders. The mechanism resembles cumulative advantage dynamics in science, where early advantages can attract further adoption and reinforce advantage (Merton, 1968). Indexing and metrics histories also illustrate how infrastructure placement can shape competitiveness and perceived legitimacy (Garfield, 2006).

This article treats the analogy as a testable hypothesis rather than as a forecast.

Evaluating the path‑dependence claim requires explicit boundary conditions:

  1. Tools materially differentiate between provenance‑rich licensed corpora and scraped or uncurated text.

  2. Procurement and compliance processes value auditability and documentation sufficiently to sustain willingness to pay.

  3. Author submission decisions are influenced by AI‑mediated visibility in ways that are not fully substituted by open web availability.

Stakeholder incentive analysis

Authors face three salient dimensions: agency, cost and visibility. Agency matters because authors can regard AI training use as ethically salient even when legally permitted. A standardized post‑acceptance opt in makes the decision explicit and reduces surprise, supporting the social licence principles of transparency and trust (Muller et al., 2021). Cost structures differ by publishing model: APC rebates may be meaningful in APC contexts, while diamond open access implementations require fee‑neutral benefits or community reinvestment commitments (Dufour et al., 2023).

Visibility claims require careful treatment. Evidence of an open access citation advantage is mixed and field‑dependent, but large‑scale analyses and systematic reviews often find an average association between openness and citation outcomes (Langham‑Putrow et al., 2021; Piwowar et al., 2018). Any analogous ‘AI‑mediated discoverability’ effect should therefore be framed as a hypothesis for evaluation rather than a promised outcome.

Publishers and societies, in turn, benefit from lower transaction costs, as the Third Option locates governance at the standard contracting stage (Coase, 1937; Williamson, 2010). It also provides a legitimacy mechanism: editorial independence plus transparency reporting can reduce reputational risk relative to unilateral action (Carter et al., 2015).

Developers receive curated and documented data under the Third Option. Documentation and reporting norms are often treated as governance tools in machine learning, supporting auditing and accountability (Bender & Friedman, 2018; Gebru et al., 2021; Mitchell et al., 2019). Provenance and versioning further support traceability and correction processes (Buneman & Tan, 2019). Regulatory attention to training‑data transparency can increase demand for provenance‑rich datasets in some contexts (European Union, 2024).

Institutions and funders face compliance, cost and reputational considerations. In APC‑funded publishing, benefits could reduce effective compliance costs. Under diamond open access arrangements, revenue from dataset licensing could support community infrastructure if aligned with diamond funding models and disclosed transparently (Dufour et al., 2023).

Objections and replies

  • O1. ‘For CC BY works, permission already exists, so the Third Option is redundant.’

    Reply: For CC BY content, the addendum cannot restrict public reuse rights. In these cases, the Third Option is best understood in terms of governance and packaging: explicit author‑facing disclosure, standardized provenance, rights hygiene and dataset documentation for curated releases (Bender & Friedman, 2018; Gebru et al., 2021).

  • O2. ‘Developers will scrape anyway, so they will not pay.’

    Reply: Scraping is technically feasible but often yields uncertain provenance, incomplete rights hygiene and weaker auditability. Curated datasets offer a different product: documented composition, provenance and contractual assurances aligned to compliance expectations (Gebru et al., 2021; Mitchell et al., 2019). Whether these differences sustain payment is not assumed; it is an empirical question for evaluation.

  • O3. ‘Benefit‑linked opt in creates undue inducement and inequity.’

    Reply: This concern is credible, particularly where APCs are high and funding is unequal. Evidence suggests that APC systems can affect geographic diversity and participation (Ellers et al., 2017; Smith et al., 2021). The Third Option therefore requires safeguards: post‑acceptance consent, editorial separation, needs‑based waiver routes independent of consent and monitoring of disparate uptake (McGregor, 2005).

  • O4. ‘Withdrawal is meaningless if unlearning is infeasible.’

    Reply: Withdrawal must be defined with honest semantics. A feasible definition of governance is future‑only exclusion from subsequent dataset releases, with versioning, notifications and disclosure that the reversal of prior training may be impractical (Xu et al., 2023).

  • O5. ‘This legitimizes monetization without fair value distribution.’

    Reply: This is a governance design critique, not a refutation of workflow integration. The Third Option should specify benefit allocation and commit to periodic transparency reporting to sustain social licence (Muller et al., 2021).

Propositions

  • P1. A post‑acceptance, opt‑in addendum can reduce transaction costs relative to retroactive outreach and thereby support scalable author governance (Coase, 1937; Williamson, 2010).

  • P2. Bundling opt‑in governance with provenance, rights hygiene and dataset documentation can differentiate curated dataset products from scraped collections or permission‑only framings (Bender & Friedman, 2018; Buneman & Tan, 2019; Gebru et al., 2021).

  • P3. Under specified boundary conditions, early adoption can generate cumulative advantage through the accumulation of provenance‑rich, consented content (Garfield, 2006; Merton, 1968).

  • P4. Withdrawal can be defined coherently as future‑only exclusion from later dataset releases via versioning, while acknowledging the practical limits of machine unlearning (Buneman & Tan, 2019; Xu et al., 2023).

  • P5. Without safeguards and monitoring, benefit‑linked opt in can create inequitable uptake patterns and undermine legitimacy through undue inducement dynamics (Ellers et al., 2017; McGregor, 2005; Smith et al., 2021).

Evaluation protocol (a methods‑style proposal for future testing)

The Third Option is a conceptual institutional design and requires empirical evaluation. Table 5 provides a minimal evaluation matrix specifying observable indicators and feasible study designs for each proposition.

Table 5

Evaluation protocol (propositions, indicators, designs and boundary conditions)

PropositionObservable indicators (examples)Candidate study designsBoundary conditions to report explicitly
P1Uptake rate among accepted articles; time‑to‑consent completion; administrative cost per consented article.Journal roll‑out evaluation using administrative records; comparative case studies of workflow designs.Disclosure clarity; workflow integration; separation of editorial and licensing systems.
P2Provenance completeness; documentation quality; third‑party exclusion rates; audit clause prevalence in contracts.Document analysis of dataset artefacts and contracts; interviews with compliance and procurement stakeholders.Field heterogeneity in third‑party content prevalence and documentation norms.
P3Growth of consented corpus over time; partnership frequency; any association with submission trends.Time‑series analysis across venues; matched comparisons of early and late adopters.Tools must differentiate provenance‑rich datasets; market demand for auditability must be present.
P4Withdrawal requests; time‑to‑exclusion from subsequent releases; version traceability and notification logs.Implementation audits; stakeholder surveys on acceptability and understanding.Withdrawal cannot imply guaranteed model untraining; open‑licence terms remain unaffected.
P5Uptake disparities by region, funding context or career stage; waiver use; complaint rates and perceived pressure.Equity monitoring with privacy‑preserving aggregate reporting; quasi‑experimental comparison of benefit and waiver policies.Reporting must protect privacy and interpret disparities in context (discipline, mandates and funding).

For feasibility, administrative indicators (uptake, time‑to‑consent completion and staff effort) can be measured from workflow logs, while author understanding and perceived pressure can be assessed via short surveys. Where venues adopt the Third Option at different times, a staggered rollout can support comparative analysis (e.g. matched comparisons or difference‑in‑differences) under the boundary conditions stated in Table 5.

Limitations and boundary conditions

  • Legal heterogeneity: training‑data governance differs across jurisdictions, so incentives will not generalize uniformly.

  • Open‑licence contexts: for CC BY works, the addendum adds governance, provenance and assurance, not new permission.

  • Market uncertainty: willingness to pay for curated datasets depends on evolving compliance, procurement and reputational incentives.

  • Technical limits: withdrawal can govern future releases but cannot guarantee the reversal of past training (Xu et al., 2023).

  • Ethical risks: safeguards are necessary to reduce pressure and inequity, but they may also reduce uptake.

Conclusion

Scholarly publishing faces an institutional deadlock in governing AI training uses of scholarly texts: unilateral agreements can undermine legitimacy, while retroactive opt‑in campaigns often fail to scale. The Third Option offers a workflow‑integrated, post‑acceptance opt‑in mechanism that treats legitimacy and scalability as coequal requirements, emphasizes editorial independence, rights hygiene and provenance and defines withdrawal as future‑only exclusion from subsequent dataset releases. This article has formalized five propositions and addressed the key objections. The propositions identify the empirical claims that future pilots should test, while the objections specify the legitimacy, equity and implementation risks that the mechanism must address. An evaluation protocol has been specified to enable empirical testing without assuming particular market or legal outcomes.

AI Usage

ChatGPT (OpenAI) was used to assist with translation into English and with mechanical editing (e.g. grammar, clarity and consistency). The conceptual design, claims and final wording are the author’s responsibility. ChatGPT is not listed as an author, and the author remains fully responsible for the content.

Abbreviations and Acronyms

A list of the abbreviations and acronyms used in this and other Insights articles can be accessed here – click on the following URL and then select the ‘full list of industry A&As’ link: http://www.uksg.org/publications#aa.

Competing interests

The author has declared no competing interests.

DOI: https://doi.org/10.1629/uksg.776 | Journal eISSN: 2048-7754
Language: English
Page range: 16 - 16
Submitted on: Feb 25, 2026
Accepted on: Apr 20, 2026
Published on: Jul 17, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Masaya Ochiai, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.