High-performance hate speech detection models achieve high accuracy; however, their complex architectures make it difficult to interpret the basis of their predictions. Although methods focusing on explainability have been proposed, their effectiveness remains limited. This study aims to develop a hate speech detection model with improved explainability.
This study proposes an enhanced two-stage learning model for hate speech detection, consisting of rationale learning in the first stage and fine-tuning in the second stage. The model incorporates two components: multi-class rationale learning to capture fine-grained rationale semantics and multi-task learning to capture target-related contextual cues through joint hate speech detection and target classification.
The experimental results suggest that the proposed method can improve explainability while maintaining competitive detection performance. Additional analysis further supports the effectiveness of both components in improving explainability.
This study is limited to English datasets, and its generalizability to other languages remains to be explored. Practical deployment may also involve challenges related to scalability, latency, and ethical considerations such as bias and fairness.
By introducing three-class rationale labels, the proposed method captures a broader range of rationales than conventional approaches. In addition, integrating multi-task learning into an explainable framework further enhances the quality of explanations.
