Table 1
Examples of person-centered and data-driven analyses.
| Analysis | Description | Reference |
|---|---|---|
| Topological Data Analysis | Used to identify geometric patterns in multivariate data. Continuous structures are built on top of the data and geometric information is extracted from the created structures and used to identify groups. For more information, see the example from engineering education provided below. | Chazal & Michel, 2017 |
| Cluster Analysis | Used to create groups according to similarity between observations in a dataset, often through the algorithm K-means clustering. Groups are created according to their distance from the center of a cluster and group assignment is not probabilistic. | Garcia-Dias et al., 2020 |
| Gaussian Mixture Modeling | Used to create groups according to similarity between observations in a dataset. Unlike cluster analysis, this technique accounts for variance in the data, and thus allows for more variability in group shape and size while providing probabilistic assignment to groups. | McNicholas, 2010 |
| Latent Profile/Class Analysis | Used to recover hidden groups from multivariate data. Falls within the larger umbrella of mixture modeling. Can be used with continuous or categorical data, and results in probability-based assignment to groups. | Oberski, 2016 |
| Growth Mixture Modeling | Similar to latent profile/class analysis but used with longitudinal data. Can be used to identify groups and then track individual movement across group lines or can be used to identify groups that emerge over time. | Ram & Grimm, 2009 |
| Artificial Neural Networks | A machine-learning classical algorithm that performs tasks using methods derived from studies of the human brain. Can be used to recognize patterns or classify data. Self-Organizing Maps (Saxxo, Motta, You, Bertolazzo, Carini, & Ma, 2017) are a form of person-centered neural networking that can be used to convert complex multivariate data into two-dimensional maps that emphasize the relationships between observations. | Abiodun et al., 2018 |
| Principal Component Analysis | Used to collapse correlated multivariate data into smaller composite components that maximize the total variance (aka dimension reduction). Often used to reduce a large number of variables to a more manageable number. For non-continuous data, categorical principal component analysis can be used. Data-driven but not person-centered. | Kherif & Latypova, 2020 |
| Multidimensional Scaling | Another form of dimension reduction, but with a focus on graphics and the visual analysis of data. Multivariate data is collapsed into two dimensions by computing the distance between variables and plotting the resulting output. Data-driven but not person-centered. | Hout et al., 2013 |
| Exploratory Factor Analysis | Used to identify latent factors or variables in correlated multivariate data. Often used in scale development or when analyzing constructs that cannot be measured directly. Data-driven but not person-centered. | Sellbom & Tellegen, 2019 |

Figure 1
The map represents students’ self-reported home Zip Codes from a national survey. Each dot may represent more than one student. This image was generated in R (R Core Team, 2018) using the ggplot2 package (Wickham, 2009).

Figure 2
TDA map generated from the analyses, including groupings based on the distribution of the network of nodes. The colors shown in the map above represent the density of the map. The blue nodes denote a population of approximately 200 students, while the red nodes denote a smaller population of approximately three to five students. Our final parameters included a k-nearest neighbors filtering method, a single-linkage hierarchical agglomerative clustering method, 35 filter slices (n), a 50% overlap in data, and a 4.0 cut height (ɛ).

Figure 3
Spider plot of average student responses on factors within TDA. Measures include disciplinary role identity constructs: Math_Int = mathematics interest; Math_PC = mathematics performance/competence beliefs; Math_Rec = mathematics recognition; Phys_Int = physics interest; Phys_PC = physics performance/competence beliefs; Phys_Rec = physics recognition; Eng_Int = engineering interest; Eng_PC = engineering performance/competence beliefs; Eng_Rec = engineering recognition. Two factors from the Big Five Personality measure were used: Ocean_NC = conscientiousness and Ocean_Neu = neuroticism. Belonging was measured in two contexts: Bel_Fac1 = in the engineering classroom and Bel_Fac2 = in engineering as a field. Students’ motivation was captured by Motiv_CR1 = controlled regulation for engaging in courses; Motiv_CR2 = controlled regulation for completing course requirements; and Motiv_AR2 = autonomous regulation for completing course requirements. Students’ epistemic beliefs (Epis_Fac4) captured the certainty of engineering knowledge (i.e., absolute to emergent).

Figure 4
Differences in controlled regulation for classroom engagement by intersectional gender and race/ethnicity groups. Groups with large enough samples for comparisons include: WW = White women, AW = Asian women, BW = Black women, LW = Latinas, WM = White men, AM = Asian men, BM = Black men, and LM = Latinos.
