Skip to main content
Have a personal or library account? Click to login
The Labeling Dilemma: Study of Clickbait Datasets and their Methodologies Cover

The Labeling Dilemma: Study of Clickbait Datasets and their Methodologies

Open Access
|Jun 2026

Abstract

In today’s digital ecosystem, there is a dominance of attention driven content platforms that promote sensationalism over informational quality. These platforms use various means to manipulate users. Clickbait is among them. It often uses misleading or exaggerated headlines to lure people to click on a link. This leaves them in discontent as the promises are never met. The aim is to gain user engagement by either routing them to a page with lots of advertisements that, in turn, boost their revenue or simply spreading misinformation. This necessitate the development of automatic clickbait detection models. This article serves as a systematic review of the work done in this domain, focusing on two important areas: existing datasets with their labeling techniques, and the evolution of various clickbait detection models from ML to DL to new pre-trained language models such as BERT and RoBERTa. This paper aims to serve academic researchers and industry professionals seeking an overview of clickbait detection methods with particular emphasis on ground truth datasets generation and their labeling strategies.

DOI: https://doi.org/10.2478/fcds-2026-0008 | Journal eISSN: 2300-3405 | Journal ISSN: 0867-6356
Language: English
Page range: 239 - 257
Submitted on: Nov 18, 2025
Accepted on: Mar 30, 2026
Published on: Jun 26, 2026
In partnership with: Paradigm Publishing Services

© 2026 Avinash Shrivastava, Anamika Gupta, Aayush Arora, Anjali Tomar, published by Poznan University of Technology
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 3.0 License.