Early Stroke Risk Prediction Using Optimized Machine Learning Models and Balanced Clinical Data

Authors

  • Sharmin Sultana Akhi , Md. Samiul Alom

Keywords:

Stroke prediction, machine learning, clinical data, SMOTE balancing, supervised learning, model tuning, health analytics

Abstract

Stroke is one of the leading causes of death and long-term disability, and many early symptoms often remain unnoticed until critical complications occur. Early identification of risk patterns is therefore essential for preventing severe outcomes. This study presents a complete machine learning framework for early stroke risk prediction using structured clinical and demographic data. The workflow includes preprocessing, class balancing with SMOTE, supervised learning with Support Vector Machine, Random Forest, and XGBoost, and tuning based optimization through GridSearchCV with five fold cross validation. The experimental results show a clear improvement after tuning. Random Forest achieved the highest performance with an accuracy of 99.0%, an F1 score of 99.3%, a ROC AUC of 99.0%, and an MCC of 98.1%. SVM also showed strong results with an accuracy of 98.8% and a ROC AUC of 99.3%. XGBoost delivered reliable performance with an accuracy of 98.4% and an F1 score of 98.7%. These findings confirm that balanced learning and proper tuning can significantly enhance predictive ability and support early stroke screening in clinical environments.

Downloads

Published

2023-03-26

How to Cite

Sharmin Sultana Akhi , Md. Samiul Alom. (2023). Early Stroke Risk Prediction Using Optimized Machine Learning Models and Balanced Clinical Data. Journal of Computational Analysis and Applications (JoCAAA), 31(3), 805–818. Retrieved from https://eudoxuspress.com/index.php/pub/article/view/4424

Issue

Section

Articles