Skip to main content
Have a personal or library account? Click to login
Cost Sensitive and Explainable Machine Learning for Heart Condition Prediction on Imbalanced Data Cover

Cost Sensitive and Explainable Machine Learning for Heart Condition Prediction on Imbalanced Data

Open Access
|Mar 2026

Abstract

Heart disease remains a leading cause of mortality worldwide, highlighting the need for predictive models that can reliably detect high risk individuals under severe class imbalance. This study presents a cost sensitive and explainable learning framework for heart condition prediction using machine learning, deep learning, and hybrid models, with particular emphasis on how dataset size influences both model effectiveness and model selection. The 2022 CDC Behavioral Risk Factor Surveillance System (BRFSS) dataset, consisting of 246,013 records with 40 features and 8.8% positive cases, was used to generate subsets ranging from 5K to the full dataset while preserving the original class distribution. Models were evaluated under Synthetic Minority Oversampling Technique (SMOTE) and Cost Sensitive Learning (CSL), and performance was assessed using recall, F1 score, and AUC–ROC. Across all model categories, CSL consistently outperformed SMOTE by improving minority class detection while maintaining data integrity. On the full dataset, logistic regression improved from a minority class F1 score of 0.320 with recall of 0.624 under SMOTE to an F1 score of 0.364 and recall of 0.782 under CSL, corresponding to an improvement of approximately 14% in F1 score. Random forest showed a stronger absolute gain, with F1 increasing from 0.173 to 0.359 and recall from 0.111 to 0.778 under CSL. Dataset size effects were evident when analyzed using CSL results: for the 5K subset, logistic regression achieved a minority class F1 score of 0.389 and the multi layer perceptron achieved 0.407, whereas on the full dataset the best performing models were random forest and TabNet, with minority-class F1 scores in the range of 0.34–0.39, while the multi layer perceptron showed reduced generalization. These findings demonstrate that model performance is not constant across dataset sizes and that optimal prediction requires selecting the most suitable model within each paradigm. Finally, a quantum support vector machine is explored as a proof of concept to compare quantum and classical kernel behavior, highlighting future potential for quantum enhanced medical analytics.

Language: English
Page range: 132 - 157
Published on: Mar 31, 2026
Published by: The Institute of Applied Statistics, Sri Lanka
In partnership with: Paradigm Publishing Services

© 2026 K. A. M. N. Aryawansha, C. H. Magalla, N. Hettiarachchi, published by The Institute of Applied Statistics, Sri Lanka
This work is licensed under the Creative Commons Attribution 4.0 License.