Skip to main content
Have a personal or library account? Click to login
Reliability-Screened Human Evaluation of Four Conversational AI Platforms: Chatgpt, Perplexity, Claude and Gemini Cover

Reliability-Screened Human Evaluation of Four Conversational AI Platforms: Chatgpt, Perplexity, Claude and Gemini

Open Access
|Sep 2026

Abstract

Human evaluation of large language models is widely used but its reliability is rarely verified. This study evaluated four conversational platforms, ChatGPT 5, Perplexity, Claude Sonnet 4 and Gemini 2.5 Flash, accessed on 20 August 2025. Twenty prompts covering factual, creative, technical and ethical-critical tasks were submitted identically to all four systems. Ten independent evaluators ranked each set of four responses on four criteria, yielding 800 forced-ranking blocks. Concordance was screened before any comparison. Only the factual set reached agreement above chance (Kendall’s W = 0.0379; mean pairwise agreement 0.561); the creative, technical and ethical-critical sets were statistically indistinguishable from random ranking. On the factual set Perplexity and ChatGPT outranked Claude and Gemini (Friedman chi-square = 22.30, p < 0.001), an effect carried almost entirely by response structure. Three of four dimensions in routine use were therefore unusable as measurements. Reliability screening should precede interpretation in human evaluation of language models.

DOI: https://doi.org/10.2478/bipcm-2026-0028 | Journal eISSN: 2537-4869 | Journal ISSN: 1011-2855
Language: English
Page range: 107 - 119
Submitted on: Aug 3, 2026
Accepted on: Aug 20, 2026
Published on: Sep 21, 2026
Published by: Gheorghe Asachi Technical University of Iasi
In partnership with: Paradigm Publishing Services
Publication frequency: 4 issues per year

© 2026 Emanuel Mihăluţe, Nicolae-Răzvan Mititelu, published by Gheorghe Asachi Technical University of Iasi
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.