A Comparative Analysis of the Application of the Random Forest and Logistic Regression Machine Learning Algorithms in Text Classification
DOI:
https://doi.org/10.26439/interfases2026.n023.8710Keywords:
machine learning, natural language processing, clinical text classification, logistic regression, random forestAbstract
Managing the vast amount of unstructured data we generate today would be unthinkable without automatic text classification. This tool has become indispensable for tasks as varied as sentiment analysis and combating fake news. However, a significant technical dilemma arises: despite the vast array of algorithms designed for these purposes, there is still no clear consensus on which is the most effective for each specific dataset. To address this issue, the present study compares the robustness of two models: Random Forest, which operates using a set of trees, and Logistic Regression, a linear model. The corpus used is Symptom2Disease, a dataset of 1,200 text records containing descriptions of symptoms in natural language, evenly distributed across 24 diagnostic categories, covering conditions such as psoriasis, diabetes, dengue, and malaria. For this purpose, a factorial experimental design was adopted, incorporating standard preprocessing, data augmentation to mitigate sample size limitations, and representation analysis via N-grams. In other words, it appears that Logistic Regression performs better than Random Forest, achieving a maximum F1 score of 0.9797 when used in conjunction with data augmentation and unigrams. For short medical texts, the linearity of the model and synthetic corpus enrichment prove to be a more efficient and accurate solution compared to non-linear models.
Downloads
References
Aggarwal, C. C., & Zhai, C. (2012). A Survey of Text Classification Algorithms. En C. C. Aggarwal & C. Zhai (Eds.), Mining Text Data (pp. 163-222). Springer US. https://doi.org/10.1007/978-1-4614-3223-4_6
Alsaeedi, A. (2020). A survey of term weighting schemes for text classification. International Journal of Data Mining, Modelling and Management, 12(2), 237. https://doi.org/10.1504/IJDMMM.2020.10028060
Báez, P., Arancibia, A. P., Chaparro, M. I., Bucarey, T., Núñez, F., & Dunstan, J. (2022). Procesamiento de lenguaje natural para texto clínico en español: el caso de las listas de espera en Chile. Revista Médica Clínica Las Condes, 33(6), 576-582. https://doi.org/10.1016/J.RMCLC.2022.10.002
Barman, N., Karim, F., & Sharma, K. (2023). Symptom2Disease: Diseases and Natural Language Symptom Descriptions. https://www.kaggle.com/datasets/niyarrbarman/symptom2disease?resource=download
Calero Sánchez, M., González González, J. C., Sánchez Berriel, I., Burillo-Putze, G., & Roda García, J. L. (2024). Natural language processing for reviewing search results from PubMed. Revista Española de Urgencias y Emergencias. https://doi.org/10.55633/S3ME/REUE030.2024
Cardoso, F. E., Ferneda, E., & Botega, L. (2023). Clasificación de textos: un enfoque con uso de machine learning. Revista EDICIC, 3(3), 1–17. https://doi.org/10.62758/re.v3i3.212
Chapman, P., Clinton, J., Kerber, R., Khabaza, T., Reinartz, T., Shearer, C., & Wirth, R. (2000). CRISP-DM 1.0: Step-by-step data mining guide. SPSS Inc. DaimlerChrysler. https://public.dhe.ibm.com/software/analytics/spss/documentation/modeler/14.2/es/CRISP-DM.pdf
Genkin, A., Lewis, D. D., & Madigan, D. (2007). Large-Scale Bayesian Logistic Regression for Text Categorization. Technometrics, 49(3), 291-304. https://doi.org/10.1198/004017007000000245
Kowsari, K., Jafari Meimandi, K., Heidarysafa, M., Mendu, S., Barnes, L., & Brown, D. (2019). Text Classification Algorithms: A Survey. Information, 10(4). https://doi.org/10.3390/info10040150
Li, Q., Peng, H., Li, J., Xia, C., Yang, R., Sun, L., Yu, P. S., & He, L. (2022). A Survey on Text Classification: From Traditional to Deep Learning. ACM Trans. Intell. Syst. Technol., 13(2). https://doi.org/10.1145/3495162
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press. https://doi.org/10.1017/CBO9780511809071
Pal, M. (2005). Random forest classifier for remote sensing classification. International Journal of Remote Sensing - INT J REMOTE SENS, 26, 217-222. https://doi.org/10.1080/01431160412331269698
Powers, D. (2008). Evaluation: From Precision, Recall and F-Factor to ROC, Informedness, Markedness & Correlation. Mach. Learn. Technol., 2. https://doi.org/10.48550/arXiv.2010.16061
Pranckevičius, T., & Marcinkevičius, V. (2017). Comparison of Naive Bayes, Random Forest, Decision Tree, Support Vector Machines, and Logistic Regression Classifiers for Text Reviews Classification. Baltic Journal of Modern Computing, 5(2). https://doi.org/10.22364/bjmc.2017.5.2.05
Sun, Y., Li, Y., Qingtao, Z., & Bian, Y. (2020). Application Research of Text Classification Based on Random Forest Algorithm. 370-374. https://doi.org/10.1109/AEMCSE50948.2020.00086
Vabalas, A., Gowen, E., Poliakoff, E., & Casson, A. (2019). Machine learning algorithm validation with a limited sample size. PLOS ONE, 14(11), 1-20. https://doi.org/10.1371/journal.pone.0224365
Wei, J., & Zou, K. (2019). EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks. En K. Inui, J. Jiang, V. Ng, & X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 6382-6388). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1670
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under an Attribution 4.0 International (CC BY 4.0) License. that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
Last updated 03/05/21


