Dimensionality reduction in machine learning for arabic text classification

Loading...
Thumbnail Image

Date

Journal Title

Journal ISSN

Volume Title

Publisher

Setif 1 University - Ferhat ABBAS , Faculty of Sciences

Abstract

Text classification is the automated process of assigning predefined labels or categories to text based on its content. This process helps organize vast amounts of textual data, simplifies management, enables efficient searches, and extracts valuable knowledge. The computa-tional analysis of the Arabic language plays a crucial role in addressing its growing global significance. As the fourth most widely used language online, Arabic has driven the emer-gence of Arabic Text Classification (ATC) as a key research area. However, the field of ATC faces considerable challenges, primarily due to the linguistic complexity of the language and the high demands of its processing, which can impact the performance of real-time systems. This dissertation aims to bridge the gap between effectiveness and efficiency in ATC, particularly in resource-constrained environments. The first objective of this research is to review existing ATC techniques, including pre-processing methods, vectorization strategies, dimensionality reduction techniques, and both classical machine learning and deep learning models, in order to provide a comprehensive understanding of current approaches. The second objective is to propose three innovative methods to enhance computational efficiency through dimensionality reduction while im- proving or at least maintaining high classification effectiveness. These methods are specif-ically designed for Modern Standard Arabic (MSA) text classification and are evaluated against state-of-the-art methods.

Description

Citation

Endorsement

Review

Supplemented By

Referenced By