TEXT CLASSIFICATION USING LANGUAGE INDEPENDENT DATA AUGUMENTATION
Keywords:
Text classification, Low resource language, Labeled data, Data augmentation, Synthetic data, Language-independent, Multilingual language model, Sentence embedding, LSTM model, BERT model.Abstract
Developing a high-performance textclassification model in a low resource language ischallenging due to the lack of labeled data.Meanwhile, collecting large amounts of labeled datais cost inefficient. One approach to increase
References
X. Zhang, J. Zhao, and Y. LeCun, ‘‘Characterlevel convolutional networks for text classification,’’ in Proc. 28th Int. Conf. Neural Inf. Process. Syst.
(NIPS), vol. 1. Cambridge, MA, USA: MIT Press, 2015, pp. 649–657.
J. Mueller and A. Thyagarajan, ‘‘Siamese recurrent architectures for learning sentence similarity,’’ in Proc. 13th AAAI Conf. Artif. Intell.,
, pp. 2786–2792.


