Skip to main content
Have a personal or library account? Click to login
Reliability-Screened Human Evaluation of Four Conversational AI Platforms: Chatgpt, Perplexity, Claude and Gemini Cover

Reliability-Screened Human Evaluation of Four Conversational AI Platforms: Chatgpt, Perplexity, Claude and Gemini

Open Access
|Sep 2026

References

  1. Amidei J., Piwek P., Willis A., Rethinking the Agreement in Human Evaluation Tasks, Proc. of the 27th Internat. Conf. on Computat. Linguistics, COLING 2018, August 2018, Santa Fe, USA, Association for Computational Linguistics, 3318-3329 (2018).
  2. Bender E.M., Gebru T., McMillan-Major A., Mitchell M., On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?, Proc. of the 2021 ACM Conf. on Fairness, Accountability, and Transparency, FAccT 2021, March 2021, Virtual Event, Canada, ACM, New York, 610-623 (2021).
  3. Brown T.B., Mann B., Ryder N., Subbiah M., Kaplan J., Dhariwal P., Neelakantan A., Shyam P., Sastry G., Askell A., Language Models are Few-Shot Learners, Adv. in Neural Inform. Process. Syst., 33, 1877-1901 (2020).
  4. Chen L., Zaharia M., Zou J., How Is ChatGPT’s Behavior Changing over Time?, Harvard Data Sci. Rev., 6, 2 (2024).
  5. Dwivedi Y.K., Kshetri N., Hughes L., Janssen M., Drăgoicea M., Wright R., So What if ChatGPT Wrote It? Multidisciplinary Perspectives on Opportunities, Challenges and Implications of Generative AI for Research, Practice and Policy, Internat. J. of Inform. Manag., 71, 102642 (2023).
  6. Friedman M., The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance, J. of the Amer. Statist. Assoc., 32, 200, 675-701 (1937).
  7. Hendrycks D., Burns C., Basart S., Zou A., Mazeika M., Song D., Steinhardt J., Measuring Massive Multitask Language Understanding, Proc. of the 9th Internat. Conf. on Learning Representations, ICLR 2021, May 2021, Virtual Event (2021).
  8. Howcroft D.M., Belz A., Clinciu M.-A., Gkatzia D., Hasan S.A., Mahamood S., Mille S., van Miltenburg E., Santhanam S., Rieser V., Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions, Proc. of the 13th Internat. Conf. on Natural Language Generation, INLG 2020, December 2020, Dublin, Ireland, Association for Computational Linguistics, 169-182 (2020).
  9. Jain S., A Comparative Assessment of Advanced Conversational Agents: A Multifaceted Evaluation of ChatGPT, Gemini, Perplexity, and Claude, Internat. J. of Emerg. Technol. and Adv. Engng., 14, 2, 59-69 (2024).
  10. Kasneci E., Sessler K., Küchemann S., Bannert M., Dementieva D., Fischer F., Kasneci G., ChatGPT for Good? On Opportunities and Challenges of Large Language Models for Education, Learning and Individ. Differences, 103, 102274 (2023).
  11. Kendall M.G., Babington Smith B., The Problem of m Rankings, The Ann. of Math. Statist., 10, 3, 275-287 (1939).
  12. Lewis P., Perez E., Piktus A., Petroni F., Karpukhin V., Goyal N., Küttler H., Lewis M., Yih W.-t., Rocktäschel T., Riedel S., Kiela D., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Adv. in Neural Inform. Process. Syst., 33, 9459-9474 (2020).
  13. Liang P., Bommasani R., Lee T., Tsipras D., Soylu D., Yasunaga M., Zhang Y., Narayanan D., Wu Y., Kumar A., Holistic Evaluation of Language Models, arXiv:2211.09110 (2022).
  14. Liu N.F., Zhang T., Liang P., Evaluating Verifiability in Generative Search Engines, Findings of the Assoc. for Computat. Linguistics, EMNLP 2023, December 2023, Singapore, Association for Computational Linguistics, 7001-7025 (2023).
  15. Mihăluţe E., Mititelu N.-R., The Limitations of Existing Software in Identifying Content Created with the Help of AI, Bul. Inst. Polit. Iaşi, 71, 3, s. Constr. de Maşini, 9-18 (2025).
  16. Mirowski P., Mathewson K.W., Pittman J., Evans R., Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry Professionals, arXiv:2209.14958 (2023).
  17. Mititelu N.-R., Hobjala R., Ermolai V., Rîpanu I.M., Bisog C., Use of AI Language Models in Generating 3D Printable Models, Acta Technica Napocensis -Series: Applied Mathematics, Mechanics, and Engineering, vol. 68, no. 1S. Available: https://atnamam.utcluj.ro/index.php/Acta/article/view/2779. (2025)
  18. Radford A., Wu J., Child R., Luan D., Amodei D., Sutskever I., Language Models are Unsupervised Multitask Learners, OpenAI Technical Report (2019).
  19. Shuster K., Poff S., Chen M., Kiela D., Weston J., Retrieval Augmentation Reduces Hallucination in Conversation, Findings of the Assoc. for Computat. Linguistics, EMNLP 2021, November 2021, Punta Cana, Dominican Republic, Association for Computational Linguistics, 3784-3803 (2021).
  20. Srivastava A., Rastogi A., Rao A., Shoeb A.A.M., Abid A., Fisch A., Brown A.R., Santoro A., Gupta A., Garriga-Alonso A., Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models, arXiv:2206.04615 (2022).
  21. Van der Lee C., Gatt A., van Miltenburg E., Krahmer E., Human Evaluation of Automatically Generated Text: Current Trends and Best Practice Guidelines, Comput. Speech and Lang., 67, 101151 (2021).
  22. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser L., Polosukhin I., Attention Is All You Need, Adv. in Neural Inform. Process. Syst., 30, 5998-6008 (2017).
  23. Weidinger L., Mellor J., Rauh M., Griffin C., Uesato J., Huang P.-S., Cheng M., Glaese M., Balle B., Kasirzadeh A., Ethical and Social Risks of Harm from Language Models, arXiv:2112.04359 (2021).
  24. Wu M., Aji A.F., Style Over Substance: Evaluation Biases for Large Language Models, arXiv:2307.03025 (2023).
  25. Zheng L., Chiang W.-L., Sheng Y., Zhuang S., Wu Z., Zhuang Y., Lin Z., Li Z., Li D., Xing E.P., Zhang H., Gonzalez J.E., Stoica I., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Adv. in Neural Inform. Process. Syst., 36, Datasets and Benchmarks Track (2023).
DOI: https://doi.org/10.2478/bipcm-2026-0028 | Journal eISSN: 2537-4869 | Journal ISSN: 1011-2855
Language: English
Page range: 107 - 119
Submitted on: Aug 3, 2026
Accepted on: Aug 20, 2026
Published on: Sep 21, 2026
Published by: Gheorghe Asachi Technical University of Iasi
In partnership with: Paradigm Publishing Services
Publication frequency: 4 issues per year

© 2026 Emanuel Mihăluţe, Nicolae-Răzvan Mititelu, published by Gheorghe Asachi Technical University of Iasi
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.