Scalable Data Processing Frameworks: From MapReduce to Distributed Gradient Boosting
Keywords:
distributed computing; MapReduce; Apache Spark; gradient boosting; XGBoost; LightGBM; scalability; big dataAbstract
The last two decades of large-scale data processing trace an arc from coarse-grained batch computationtoward fine-grained, communication-aware machine learning. This article reviews that arc, connecting two literatures that are usually treated separately
References
J. Dean and S. Ghemawat, “MapReduce: Simplified data processing on large clusters,” in Proc. 6th USENIX Symp. Operating Systems Design and Implementation (OSDI), 2004, pp. 137–150
Downloads
Published
2022-01-20
How to Cite
Ravi Teja Gurram. (2022). Scalable Data Processing Frameworks: From MapReduce to Distributed Gradient Boosting. Journal of Computational Analysis and Applications (JoCAAA), 30(2), 1201–1210. Retrieved from https://eudoxuspress.com/index.php/pub/article/view/5617
Issue
Section
Articles


