Skip to main content
Have a personal or library account? Click to login
Comparing AI-generated care plans with gold-standard care manager plans: a two-layer blinded pilot study Cover

Comparing AI-generated care plans with gold-standard care manager plans: a two-layer blinded pilot study

Open Access
|Sep 2026

Abstract

Background: and Aim: Large language models (LLMs), commonly referred to as artificial intelligence (AI), can assist care managers in preparing structured health plans by synthesising clinical and social data. However, their practical value and safety remain to be determined. This small-scale pilot study evaluated whether care plans generated with a general-purpose language model (ChatGPT-5, OpenAI, Plus subscription) were comparable in quality and safety to those produced by experienced care managers (specialist nurses or social workers).

 

Methods: A paired-sample design was applied using data from 11 patients participating in an integrated care programme at Viljandi Hospital. For each patient, two plans were created: by care manager (CM-CP) and by the language model (AI-CP). Both plans were based on identical inputs—hospital discharge summaries, national e-health records, prescription data, and the holistic InterRAI assessment results.

The key difference was that care managers had direct patient contact and performed the InterRAI assessments. All clinical data were pseudonymised before processing. AI plans were generated iteratively in ChatGPT-5 using standardised prompts, with each patient handled in a separate “chat”.

 

A two-tier blinded peer review was conducted by three independent care managers (tier 1) and two internist physicians (tier 2), all experienced in care management. An independent investigator normalised all plans for structure and appearance to conceal authorship. Experts rated safety, structural completeness, practicality, level of individualisation, person-centredness, and integration of medical and social aspects using a five-point scale (1 = poor/unsafe; 5 = excellent/no risk).

 

Results: Reviewers were generally able to distinguish AI- and human-authored plans (82-97%).

Tier 1 reviewers (n = 66 total ratings) found CM-CPs mostly safe: 55 % of plans were rated risk-free and 21 % carried only minor risks. In contrast, AI-CPs included a notable share of critical or major risks (12 % each). Regarding specificity and individualisation, 100 % of CM-CPs were rated good or excellent, compared with 57 % of AI-CPs; one-third of AI-CPs showed substantial deficiencies. Person-centredness was rated excellent in 46 % of CM-CPs versus 30 % of AI-CPs.

Tier 2 physician reviewers (n = 44 ratings) observed similar trends: 77 % of CM-CPs were free of risk compared with only 36 % of AI-CPs, which more frequently contained medication-related and self-monitoring errors. For specificity and individualisation, 82 % of CM-CPs scored the maximum of 5 versus 27 % of AI-CPs. Person-centredness achieved the highest rating (5) in 91 % of CM-CPs but only 23 % of AI-CPs.

 

Conclusions: AI-generated care plans remain inferior to the “gold standard” plans produced by specialist care managers. Nevertheless, current off-the-shelf LLMs can create coherent structural frameworks and synthesise clinical information to create a decent care-plan. Without professional editing and physician validation, however, AI-generated plans introduce clinically relevant safety risks and lack sufficient individualisation.

 

Future integrated-care workflows will likely combine AI-generated drafts with expert human review to enhance efficiency while maintaining patient safety and person-centredness. Further research with larger samples is required to explore the ethical, practical, and cost-effectiveness implications of such hybrid AI–human models in real-world care coordination.

Journal eISSN: 1568-4156
Language: English
Page range: 089 - 089
Published on: Sep 11, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Liis Puis, Kadri Oras, Aive Purason, Luule Vitsur, Diana Palumäe, Margit Aab, Elvi Link, Karoliina Hunt, Mann Randaru, Maret Moisa, Mart Kull, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.