The Impact of Signal Deletion in Text Preprocessing on Timestamp Comment Detection under Severe Class Imbalance

decha danillo novicahyanto, Albertus Dwivoga Widiantoro

Abstract


Timestamping in YouTube comment sections is a distinctive form of participatory content curation. This study examines why four text classification algorithms, namely Logistic Regression, Support Vector Machine, Random Forest, and XGBoost, fail completely at detecting timestamp-bearing comments in an informal Indonesian-language corpus. Across 14,149 comments from Ruangguru Clash of Champions videos, the imbalance ratio reaches 73.86:1. Every model returns an accuracy near 98.7% with a minority-class recall of zero, a Balanced Accuracy of 0.50, and Cohen’s Kappa at or below zero, the textbook signature of the Accuracy Paradox. The contribution is diagnostic: class imbalance is not the sole cause. A token-level trace shows that punctuation removal followed by standalone-numeral removal deterministically destroys the pattern defining the positive label, while sparse-term pruning at a 0.99 threshold discards vocabulary occurring in fewer than 1% of documents, above the minority prevalence of 1.34%. Recovering the confusion matrices permits a complete imbalance-robust metric set to be computed, confirming that all four models are indistinguishable from a constant classifier that ignores its input. The study specifies the ablation required to separate these causes, and establishes that before minority-class failure is attributed to imbalance, researchers must verify that preprocessing has not deleted the signal being learned.

Full Text:

PDF

References


M. Kaur, H. S. Pannu, and A. K. Malhi, "A systematic review on imbalanced data challenges in machine learning: Applications and solutions," ACM Computing Surveys, vol. 52, no. 4, pp. 1–36, 2019.

H. Jenkins, Convergence Culture: Where Old and New Media Collide. New York: New York University Press, 2006.

A. M. Barik, R. Mahendra, and M. Adriani, "Normalization of Indonesian-English code-mixed Twitter data," in Proc. 5th Workshop on Noisy User-generated Text (W-NUT 2019), Hong Kong, 2019, pp. 417–424. doi: 10.18653/v1/D19-5554.

G. Haixiang, L. Yijing, J. Shang, G. Mingyun, H. Yuanyue, and G. Bing, "Learning from class-imbalanced data: Review of methods and applications," Expert Systems with Applications, vol. 73, pp. 220–239, 2017.

A. Fernandez, S. Garcia, F. Herrera, and N. V. Chawla, "SMOTE for learning from imbalanced data: Progress and challenges, marking the 15-year anniversary," Journal of Artificial Intelligence Research, vol. 61, pp. 863–905, 2018.

D. J. Hand, "Measuring classifier performance: A coherent alternative to the area under the ROC curve," Machine Learning, vol. 77, no. 1, pp. 103–123, 2009.

N. Japkowicz and S. Stephen, "The class imbalance problem: A systematic study," Intelligent Data Analysis, vol. 6, no. 5, pp. 429–449, 2002.

H. He and E. A. Garcia, "Learning from imbalanced data," IEEE Transactions on Knowledge and Data Engineering, vol. 21, no. 9, pp. 1263–1284, 2009.

N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002.

H. He, Y. Bai, E. A. Garcia, and S. Li, "ADASYN: Adaptive synthetic sampling approach for imbalanced learning," in Proc. IEEE Int. Joint Conf. Neural Networks, Hong Kong, 2008, pp. 1322–1328.

K. M. Ting, "A comparative study of cost-sensitive algorithms for instance ranking," in Proc. 6th IEEE Int. Conf. Data Mining, Hong Kong, 2006, pp. 1001–1006.

A. Luque, A. Carrasco, A. Martin, and A. de las Heras, "The impact of class imbalance in classification performance metrics based on the binary confusion matrix," Pattern Recognition, vol. 91, pp. 216–231, 2019.

D. Chicco and G. Jurman, "The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation," BMC Genomics, vol. 21, no. 1, pp. 1–13, 2020.

J. Davis and M. Goadrich, "The relationship between precision-recall and ROC curves," in Proc. 23rd Int. Conf. Machine Learning, Pittsburgh, 2006, pp. 233–240.

T. Fawcett, "An introduction to ROC analysis," Pattern Recognition Letters, vol. 27, no. 8, pp. 861–874, 2006.

F. Koto and G. Y. Rahmaningtyas, "InSet lexicon: Evaluation of a word list for Indonesian sentiment analysis in microblogs," in Proc. Int. Conf. Asian Language Processing (IALP), Singapore, 2017, pp. 391–394.

F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, "IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP," in Proc. 28th Int. Conf. Computational Linguistics, Barcelona (Online), 2020, pp. 757–770.

B. Wilie, K. Vincentio, G. I. Winata, S. Cahyawijaya, X. Li, Z. Y. Lim, S. Soleman, R. Mahendra, P. Fung, S. Bahar, and A. Purwarianti, "IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding," in Proc. 1st Conf. Asia-Pacific Chapter of the ACL and 10th Int. Joint Conf. Natural Language Processing, Suzhou, 2020, pp. 843–857.

T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, "Distributed representations of words and phrases and their compositionality," in Proc. 26th Int. Conf. Neural Information Processing Systems, Lake Tahoe, 2013, pp. 3111–3119.

D. D. Novicahyanto, "A comparative analysis of text classification algorithms for the information-sharing (timestamping) culture in YouTube comments on the Ruangguru Clash of Champions," B.S. thesis, Faculty of Computer Science, Soegijapranata Catholic University, Semarang, 2026.

J. Cohen, "A coefficient of agreement for nominal scales," Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960.




DOI: https://doi.org/10.15548/isrj.v6i02.14780

Refbacks

  • There are currently no refbacks.


Gedung Fakultas Sains dan Teknologi 
Kampus III Universitas Islam Negeri Imam Bonjol Padang
Sungai Bangek, Kec. Koto Tangah, Kota Padang, Sumatera Barat

Creative Commons License This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.