Abstract
Calculating the reliability of experimental tasks can be surprisingly difficult using existing tools. Although R packages such as psych are robust, they often require data in wide format and assume carefully selected items that avoid floor and ceiling effects. To encourage the reporting of task reliability in experimental research, I have written an R function, ICC_participants_long, which uses the intraclass correlation coefficient (ICC) to measure the reliability of participant scores directly from data in long format. Applying this function revealed that the current split-half approach may underestimate the reliability of experimental tasks. Furthermore, the model-based approach makes it possible to generate Best Linear Unbiased Predictions (BLUPs) as estimates of participants’ scores, which provides an informative supplement to the raw means.
© 2026 Marc Brysbaert, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.
