Reliability-Screened Human Evaluation of Four Conversational AI Platforms: Chatgpt, Perplexity, Claude and Gemini
Abstract
Human evaluation of large language models is widely used but its reliability is rarely verified. This study evaluated four conversational platforms, ChatGPT 5, Perplexity, Claude Sonnet 4 and Gemini 2.5 Flash, accessed on 20 August 2025. Twenty prompts covering factual, creative, technical and ethical-critical tasks were submitted identically to all four systems. Ten independent evaluators ranked each set of four responses on four criteria, yielding 800 forced-ranking blocks. Concordance was screened before any comparison. Only the factual set reached agreement above chance (Kendall’s W = 0.0379; mean pairwise agreement 0.561); the creative, technical and ethical-critical sets were statistically indistinguishable from random ranking. On the factual set Perplexity and ChatGPT outranked Claude and Gemini (Friedman chi-square = 22.30, p < 0.001), an effect carried almost entirely by response structure. Three of four dimensions in routine use were therefore unusable as measurements. Reliability screening should precede interpretation in human evaluation of language models.
© 2026 Emanuel Mihăluţe, Nicolae-Răzvan Mititelu, published by Gheorghe Asachi Technical University of Iasi
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.