
Figure 1
CAR-E system architecture showing the flow from user input through processing, coaching generation, and safety monitoring within a FERPA-compliant environment. The coaching agent (GPT-4) integrates retrieval-augmented generation (RAG) for evidence-based coaching methods and dual memory systems for longitudinal conversation continuity. A safety inference layer screens all outputs before delivery. Authentication via institutional Active Directory ensures appropriate governance and privacy protection.
Table 1
CAR-E Iterative Design and Refinement Log. The development of CAR-E followed an iterative, human-centered design process. Early cycles were rapid and informal, conducted with the core team and small groups of faculty and trainees; later cycles incorporated larger feedback bursts from volunteer users across the institution. The log below summarizes each cycle for transparency and transferability and reflects descriptive observation rather than formal qualitative analysis.
| CYCLE | WHO TESTED | DOMINANT ISSUE OBSERVED | CHANGE IMPLEMENTED | EVIDENCE THAT SUGGESTED IMPROVEMENT |
|---|---|---|---|---|
| 1 | One faculty member and the core design team | Foundational build and first internal testing: interface and layout, voice features, memory function, technical glitches, and the need to ground responses in coaching methodology rather than generic LLM behavior. | Built the minimalist chat interface with a conversation-topic menu and goal-setting tools; implemented the dual short and long-term memory architecture; assembled a retrieval-augmented generation (RAG) knowledge base of coaching resources to ground responses in evidence-based coaching practice and curb default advice-giving; resolved speech-to-text transcription and audio-playback glitches. | Core team confirmed reliable navigation, transcription, and session-to-session memory recall, and that responses reflected a reflective, coaching-oriented style rather than generic advice-giving. |
| 2 | Two faculty and one resident | Responses too verbose; multiple questions asked at once; CAR-E interrupted users at natural speech pauses; tone not conversational enough | Revised the system prompt to shorten responses and limit each turn to a single focused question; raised the voice model’s silence threshold so CAR-E waits through natural pauses before replying; tuned tone toward a more conversational style. | Subsequent transcripts showed shorter, single question turns; testers reported noticeably fewer interruptions and a more natural exchange. |
| 3 | Four faculty and one resident | Conversations lacked a defined focus; no prompting to shift towards setting goals | Added a session-opening step that prompts users to identify the purpose of the conversation; surfaced goal-setting prompts and tools. | Testers more consistently articulated a focus in the session; transcripts showed CAR-E better at prompting user to consider specific goals. |
| 4 | Feedback burst — 37 faculty, residents, and students (volunteer convenience sample; pilot cohort summarized in Table 1) | Inadequate recognition of and response to red-flag disclosures; formulaic reflect/question pattern; user frustration when CAR-E resisted offering advice | Began building a safety inference layer to detect concerning disclosures (e.g., unhealthy coping) and trigger escalation pathways; began varying the affirmation–question rhythm to reduce formulaic pattern; scoped a hybrid approach permitting occasional acknowledgments, reflections or suggestions when users are stuck or explicitly request input. | Two authors consolidated the 37-user free-text feedback into strengths and areas for improvement (Table 1); these items directly defined the changes tested in further cycles. |
| 5 | Two faculty and one resident | Weak conversational arc without moving user forward; assumptions about user meaning or significance; insufficient reflections and acknowledgments before advancing | Began implementing a three-phase conversational arc (establishing focus, evoking awareness, and moving toward insight and action) drawn from International Coaching Federation competencies; refined CAR-E to acknowledge and reflect back users statements and remain open-ended without assumptions. | Returning testers observed a clearer beginning–middle–end structure and that CAR-E validated what was said before advancing; informed the more recent comparison conversation shown in Figure 2. |
| 6 | Feedback burst — 28 faculty, residents, and students | Shallow emotional validation; limited contextual understanding; absence of concrete recommendations; rarely challenges the user | Ongoing improvements to prompt CAR-E to name and reflect underlying emotions before questioning; strengthening use of long-term memory for contextual continuity; piloting the hybrid approach in which CAR-E offers a concrete observation or soft suggestions when a user appears stuck or explicitly ask. | Early results suggest empathy felt less formulaic and that input or suggestions were available on request. |
[i] Cycles 1–3 and 5 were small, rapid design cycles with the core team and returning testers; Cycles 4 and 6 were larger feedback bursts with volunteer users recruited by email invitation. “Evidence” entries are descriptive observations drawn from transcript review and free-text user feedback, consistent with the early-stage, non-formal nature of this innovation report.

Figure 2
Example coaching conversations illustrating CAR-E behavior before and after revisions, with annotations highlighting common issues. Panel 4 shows the revised CAR-E responding to a similar conversation (feedback volume) after the changes.
Table 2
Organized summary of strengths and opportunities for improvement from user feedback.
| STRENGTHS | OPPORTUNITIES FOR IMPROVEMENT |
|---|---|
| Promoted self-reflection and introspection | Repetitive and Circular Questioning |
| Questioning was effective | Lack of direct or practical advice |
| Relentlessly positive in validating concerns | Lack of empathy or human-like interaction |
| Interaction felt personalized and adaptive | Conversations stagnated without new insight or clear closure |
| Useful for planning and goal setting | No recognition of red flag responses |
