Skip to main content
Have a personal or library account? Click to login
Judging Students' Texts in a Digital Research Tool: Do Text Quality, Students' Gender, and Migration Background Impact Teachers' Text Assessments? Cover

Judging Students' Texts in a Digital Research Tool: Do Text Quality, Students' Gender, and Migration Background Impact Teachers' Text Assessments?

Open Access
|Dec 2025

Figures & Tables

Tab. 1:

Means and standard deviations for all assessment scales by text quality and student gender

Text quality
ScaleLowMediumHigh
GenderM (SD)M (SD)M (SD)
HolisticMale2.88 (1.20)2.80 (1.06)3.38 (0.91)
Female2.80 (1.26)3.24 (1.10)3.44 (0.93)
ContentMale2.55 (0.85)2.68 (0.87)2.76 (0.89)
Female2.55 (0.87)2.96 (0.80)2.87 (0.80)
StyleMale2.44 (0.98)2.50 (0.84)2.86 (0.80)
Female2.43 (0.95)2.80 (0.82)2.86 (0.75)
Linguistic AccuracyMale2.64 (1.05)2.44 (1.09)3.09 (0.68)
Female2.55 (1.07)2.86 (1.04)3.11 (0.65)

[i] Note. The holistic scale ranged from 1 to 5, and the analytic scales (content, style, linguistic accuracy) ranged from 1 to 4.

Tab. 2:

Means and standard deviations for all assessment scales by text quality and students' migration back-ground

Text quality
ScaleMigration BackgroundLowMediumHigh
M (SD)M (SD)M (SD)
HolisticWithout2.72 (1.17)2.96 (1.14)3.41 (0.88)
With2.74 (1.14)2.95 (1.12)3.28 (0.86)
ContentWithout2.40 (0.81)2.76 (0.90)2.84 (0.75)
With2.47 (0.79)2.68 (0.85)2.74 (0.77)
StyleWithout2.37 (0.91)2.70 (0.89)2.86 (0.77)
With2.29 (0.89)2.64 (0.70)2.82 (0.75)
Linguistic AccuracyWithout2.61 (1.16)2.62 (1.08)3.09 (0.69)
With2.44 (1.07)2.69 (1.01)3.09 (0.71)

[i] Note. The holistic scale ranged from 1 to 5, and the analytic scales (content, style, linguistic accuracy) ranged from 1 to 4.

Figure S1

Text assessment using the holistic scale in the Student Inventory

Table S1

Results of the exploratory analysis regarding students' gender bias for teachers' assessment of medium-quality texts

Assessment scaleMale nameFemale name
MSDMSDt(116)pCohen's d
Holistic2.801.063.241.102.80.01.26
Content2.680.872.960.802.38.02.22
Style2.500.842.800.822.60.01.24
Linguistic accuracy2.441.092.861.042.36.02.22

[i] Note. N = 117; paired t-test; comparison of medium-quality texts with either a male or female student name.

Table S2

Means and standard deviations of the four assessments for all three text quality levels of Study 1

Text quality
ScaleLow M (SD)Medium M (SD)High M (SD)
Holistic2.84 (1.23)3.02 (1.10)3.41 (0.92)
Content2.55 (0.85)2.82 (0.85)2.82 (0.85)
Style2.44 (0.96)2.65 (0.84)2.86 (0.77)
Linguistic Accuracy2.59 (1.06)2.65 (1.08)3.10 (0.66)

[i] Note. The holistic scale ranged from 1 to 5, and the analytic scales (content, style, linguistic accuracy) ranged from 1 to 4.

Table S3

Post-hoc pairwise comparisons for text quality assessments of Study 1

Quality comparisonMdiff95 % CIt(233)pd
Holistic Assessment
Low vs. medium−0.18[−0.38, 0.02]−1.77.236−0.12
Low vs. high−0.57[−0.76, −0.38]−6.01< .001−0.39
Medium vs. high−0.39[−0.56, −0.22]−4.40< .001−0.29
Content Assessment
Low vs. medium−0.27[−0.41, −0.13]−3.87< .001−0.25
Low vs. high−0.27[−0.40, −3.93]−3.93< .001−0.26
Medium vs. high0.00[−0.13, 0.14]0.061.000.00
Style Assessment
Low vs. medium−0.21[−0.37, −0.06]−2.74.018−0.18
Low vs. high−0.43[−0.58, −0.27]−5.45< .001−0.36
Medium vs. high−0.21[−0.35, 0.08]−3.05.008−0.20
Linguistic accuracy
Low vs. medium−0.05[−0.24, 0.14]−0.531.00−0.03
Low vs. high−0.50[−0.66, −0.35]−6.31< .001−0.41
Medium vs. high−0.45[−0.61, −0.30]−5.88< .001−0.38

[i] Note. Bonferroni-adjusted p-values were used for multiple comparisons. Mdiff = Mean difference (M1 – M2). CI = Confidence interval. Cohen's d was used to calculate effect sizes for dependent samples.

Table S4

Means and standard deviations of the four assessments for all three text quality levels of Study 2

Text quality
ScaleLow M (SD)Medium M (SD)High M (SD)
Holistic2.73 (1.16)2.96 (1.13)3.34 (0.87)
Content2.44 (0.80)2.72 (0.88)2.79 (0.76)
Style2.33 (0.90)2.67 (0.80)2.84 (0.76)
Linguistic Accuracy2.52 (1.11)2.65 (1.04)3.09 (0.70)

[i] Note. The holistic scale ranged from 1 to 5, and the analytic scales (content, style, linguistic accuracy) ranged from 1 to 4.

Table S5

Post-hoc pairwise comparisons for text quality assessments of Study 2

Quality comparisonMdiff95 % CIt(233)pd
Holistic Assessment
Low vs. medium−0.23[−0.42, 0.04]−2.34.060−0.15
Low vs. high−0.61[−0.77, −0.46]−7.94< .001−0.50
Medium vs. high−0.39[−0.55, −0.23]−4.74< .001−0.30
Content Assessment
Low vs. medium−0.28[−0.41, −0.15]−4.23< .001−0.27
Low vs. high−0.35[−0.48, −0.23]−5.55< .001−0.35
Medium vs. high−0.07[−0.20, 0.05]−1.15.753−0.07
Style Assessment
Low vs. medium−0.34[−0.49, −0.19]−4.47< .001−0.28
Low vs. high−0.51[−0.64, −0.37]−7.46< .001−0.47
Medium vs. high−0.17[−0.30, −0.03]−2.48.041−0.16
Linguistic accuracy
Low vs. medium−0.13[−0.32, 0.06]−1.37.513−0.09
Low vs. high−0.56[−0.71, −0.42]−7.52< .001−0.47
Medium vs. high−0.43[−0.59, −0.28]−5.53< .001−0.35

[i] Note. Bonferroni-adjusted p-values were used for multiple comparisons. Mdiff = Mean difference (M1 – M2). CI = Confidence interval. Cohen's d was used to calculate effect sizes for dependent samples.

Language: English
Page range: 47 - 61
Published on: Dec 31, 2025
Published by: Gesellschaft für Fachdidaktik (GfD e.V.)
In partnership with: Paradigm Publishing Services
Publication frequency: 1 issue per year

© 2025 Frederike Strahl, Jörg Kilian, Jens Möller, published by Gesellschaft für Fachdidaktik (GfD e.V.)
This work is licensed under the Creative Commons Attribution-NonCommercial 4.0 License.