Big data classification is a challenging task, especially when dealing with imbalanced datasets where minority class instances are significantly underrepresented. Traditional machine learning models often exhibit bias toward the majority class, leading

Authors

  • M. Vamshi Krishna, Dr. Dara Eshwar

Keywords:

Big Data, Imbalance Classification, Optimization, Ensemble Learning Model, Particle Swarm Optimization, Genetic Algorithm.

Abstract

Big data classification is a challenging task, especially when dealing with imbalanced datasets where minority class instances are significantly underrepresented. Traditional machine learning models often exhibit bias toward the majority class, leading to suboptimal classification performance. This research proposes a novel optimization-tuned machine learning framework that enhances classification accuracy in imbalanced big data. The proposed approach integrates ensemble learning with an innovative hybrid
optimization algorithm that combines Particle Swarm Optimization (PSO) and Genetic Algorithm (GA) to optimize hyperparameters dynamically. The study employs various benchmark datasets and evaluates the framework using performance metrics such as Precision, Recall, F1-score, and the Area Under the Receiver Operating Characteristic Curve (AUC-ROC). Experimental results demonstrate that the proposed approach outperforms conventional machine learning methods in terms of classification
accuracy and minority class detection. The findings contribute to the advancement of imbalanced big data classification by providing an optimized and scalable solution.

References

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32.

Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321-357.

Domingos, P. (1999). Metacost: A general method for making classifiers cost-sensitive. Proceedings of the Fifth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 155-164.

Downloads

Published

2019-05-10

How to Cite

M. Vamshi Krishna, Dr. Dara Eshwar. (2019). Big data classification is a challenging task, especially when dealing with imbalanced datasets where minority class instances are significantly underrepresented. Traditional machine learning models often exhibit bias toward the majority class, leading . Journal of Computational Analysis and Applications (JoCAAA), 27(5), 1–5. Retrieved from https://eudoxuspress.com/index.php/pub/article/view/1901

Issue

Section

Articles