Scalable Data Processing Frameworks: From MapReduce to Distributed Gradient Boosting

Authors

  • Ravi Teja Gurram

Keywords:

distributed computing; MapReduce; Apache Spark; gradient boosting; XGBoost; LightGBM; scalability; big data

Abstract

The last two decades of large-scale data processing trace an arc from coarse-grained batch computationtoward fine-grained, communication-aware machine learning. This article reviews that arc, connecting two literatures that are usually treated separately

References

J. Dean and S. Ghemawat, “MapReduce: Simplified data processing on large clusters,” in Proc. 6th USENIX Symp. Operating Systems Design and Implementation (OSDI), 2004, pp. 137–150

Downloads

Published

2022-01-20

How to Cite

Ravi Teja Gurram. (2022). Scalable Data Processing Frameworks: From MapReduce to Distributed Gradient Boosting. Journal of Computational Analysis and Applications (JoCAAA), 30(2), 1201–1210. Retrieved from https://eudoxuspress.com/index.php/pub/article/view/5617

Issue

Section

Articles