
The Third Option: integrating AI training‑data licensing into scholarly publishing workflows
Abstract
Large language models are increasing demand for high‑quality training corpora, yet scholarly publishing lacks scalable and legitimacy‑preserving mechanisms for governing the inclusion of research articles in training datasets. Back‑catalogue agreements concluded without transparent governance can damage trust, while retroactive opt‑in campaigns impose high transaction costs and uncertain uptake. This article proposes the ‘Third Option’: a post‑acceptance, opt‑in addendum in which authors authorize defined artificial intelligence (AI) training uses of the accepted article in exchange for an author‑facing benefit. The addendum is designed to be choice‑expanding and editorially independent, and to operate across publishing models, including diamond open access journals that do not charge article processing charges. Its value is not limited to copyright permission: it can bundle provenance metadata, rights hygiene for third‑party content, dataset documentation, versioning and contractual assurances on permitted uses and auditability. The author defines withdrawal semantics as future‑only exclusion from subsequent dataset releases, while acknowledging the practical limits of machine unlearning. This article formalizes five propositions, addresses key objections and outlines a methods‑style evaluation protocol for empirical testing.
© 2026 Masaya Ochiai, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.