Table 1
Purposes and related interview questions aligned with the research questions.
| RESEARCH QUESTION | PURPOSE | INTERVIEW QUESTIONS | |
|---|---|---|---|
| 1. | How do learners with visual impairments use learning resources? | To identify learners’ prior habits, preferred resources, and prior attitudes towards audio content. | Which learning resources do you generally prefer in the learning process? How do you use these resources? Why? Have you used a narrated learning resource before? Which environments or formats do you prefer? Can you give examples of positive or negative experiences? |
| 2. | How do learners with visual impairments use AI-assisted audio content? | To understand technical interactions (speed, pausing, etc.) and functional usage patterns. | How did you use the AI-assisted audio content provided in the study?
|
| 3. | What are the reasons underlying preferences regarding voice gender? | To identify preferences regarding voice characteristics (female/male) and the reasons for these preferences. | What are your views on using a female or a male voice in audio content?
|
| 4. | What are cognitive and affective perceptions of AI-assisted audio content? | To examine the emotions, motivation, and levels of concentration that develop in response to an AI voice | How did you feel while using the AI-assisted audio content? How would you describe this experience?
|
| 5. | How do learners evaluate the structural characteristics of the AI voice (tone, speed, quality)? | To elicit views about the characteristics of the AI voice. | What do you think about the features of the audio content (volume, speaking speed, tone, voice quality, etc.)?
|
| 6. | What expectations and recommendations do learners have? | To evaluate autonomous learning capability with audio resources and the adequacy of the system in this regard. | How does using audio resources affect your study speed and routine? Are audio contents sufficient for you to study the course on your own without help from others? What are your recommendations for using audio contents more effectively in learning environments? Is there anything else you would like to add about audio content usage? |
[i] *During the interviews, follow-up probes such as “Could you elaborate on that?”, “Could you give an example?”, and “How did that make you feel?” were used to help participants describe their experiences in more detail.
Table 2
Emerging themes and categories identified in the study.
| NO | THEME | CATEGORIES |
|---|---|---|
| 1 | Access experiences in the learning process | Use of digital resources; Screen reader use and human support; Accessibility barriers. |
| 2 | Patterns of using AI-assisted audio content | Listening session preferences; Control and interaction with the content; Note-taking and re-access. |
| 3 | Preferences regarding voice gender | Preferences based on technical performance; Affective and symbolic perceptions; Perceptions of naturalness and artificiality; Single vs. multiple voices. |
| 4 | Cognitive load and affective perceptions | Mental fatigue; Need for forced focus; Need for segmentation and time management. |
| 5 | Evaluations of the structural features of the AI voice | Perceived lack of emotion; Robotic timbre and mechanical quality; Clarity and intelligibility; Stable voice that does not tire. |
| 6 | Instructional design expectations | Demand for past exam questions and mock tests; Overcoming PDF-related accessibility issues; Interface simplicity and access speed; Linguistic improvements. |
Table 3
Categories related to the theme ‘Access experiences in the learning process’.
| THEME | CATEGORY | N | DETAILS AND SOURCES |
|---|---|---|---|
| Access experiences in the learning process | Use of digital resources | 16 | Fully digital learning; mobile apps; lecture recordings (L1, L2, L3, L4). |
| Screen reader use and human support | 11 | VoiceOver/TalkBack use; audio-based learning (L3, L9); need for human support (L1). | |
| Accessibility barriers | 10 | Mispronunciation (L3); tone issues (L4); need to convert PDF incompatibility; structural barriers (L7, L10, L15). |
Table 4
Categories related to the theme ‘Patterns of using AI-assisted audio content’.
| THEME | CATEGORY | N | DETAILS AND SOURCES |
|---|---|---|---|
| Patterns of using AI-assisted audio content | Listening session preferences | 8 | One-session listening; segmented listening; comprehension-driven preference (L2, L6, L7). |
| Control and interaction with the content | 14 | Rewinding/pausing; speed adjustment (L4, L5, L13, L15); self-testing; control-based variation (L1, L3, L7). | |
| Note-taking and re-access | 7 | Reduced note-taking (L5); selective notes (L10, L12); re-access via recordings (L14). |
Table 5
Categories related to the theme ‘Preferences regarding voice gender’.
| THEME | CATEGORY | N | DETAILS AND SOURCES |
|---|---|---|---|
| Preferences regarding voice gender | Preferences based on technical performance | 9 | Clarity; stress accuracy; fluency; intelligibility as main criterion (L5, L6, L10, L12). |
| Affective and symbolic perceptions | 6 | Female: sincere/gentle; male: serious/reassuring; symbolic associations (L1, L8). | |
| Perceptions of naturalness and artificiality | 4 | Female: robotic/flat; male: more natural; link to performance (L13, L15) | |
| Single voice vs. multiple voices | 5 | Single voice for consistency (L1, L6, L12); voice changes causing unease (L12); multiple voices increase attention but may distract (L11, L13). |
Table 6
Categories related to the theme ‘Cognitive load and affective perceptions’.
| THEME | CATEGORY | N | DETAILS AND SOURCES |
|---|---|---|---|
| Cognitive load and affective perceptions | Mental fatigue | 10 | Long listening → fatigue; attention loss; difficulty processing words (L8, L12). |
| Need for forced focus | 7 | Extra effort due to monotony; difficulty identifying key points (L5, L8). | |
| Need for segmentation and time management | 4 | Preference for shorter segments; micro-learning need (L5, L12). |
Table 7
Categories related to the theme ‘Evaluations of the structural features of the AI-assisted voice’.
| THEME | CATEGORY | F | DETAILS AND SOURCES |
|---|---|---|---|
| Evaluations of the structural features of the AI-assisted voice | Perceived lack of emotion and ‘soullessness’ | 11 | Lack of warmth/sincerity; limited intonation; impersonal delivery (L2) |
| Robotic timbre and mechanical quality | 9 | Metallic/flat tone; monotony (L4, L9); distraction vs. attention effect (L15). | |
| Clarity and intelligibility | 7 | Clear pronunciation; fluent reading; fast information access (L2). | |
| Stable voice that does not tire (continuity) | 7 | Consistent performance in long readings (L4, L12). |
Table 8
Categories related to the theme ‘Instructional design expectations’.
| THEME | CATEGORY | F | DETAILS AND SOURCES |
|---|---|---|---|
| Instructional design expectations | Demand for past exam questions and mock tests | 12 | Need for independent practice; accessible question banks (L13, L15). |
| Overcoming PDF conversion difficulties | 8 | Inaccessibility of PDF-based materials (L13, L15). | |
| Interface simplicity and access speed | 6 | Simple interface; fast access; seamless integration (L5, L16). | |
| Linguistic improvements (pronunciation/stress) | 5 | Stress/intonation issues; need for natural pronunciation (L4, L13). |
