Business & Management Studies

Detection of Homophobia & Transphobia in Malayalam and Tamil: Exploring Deep Learning Methods

Detection of Homophobia & Transphobia in Malayalam and Tamil: Exploring Deep Learning Methods

This paper attempts to explore applicability of different deep learning models for classification of the social media comments in Malayalam and Tamil languages as homophobic, transphobic and non-anti-LGBT + content.

Vedika Gupta, Assistant Professor, Jindal Global Business School, O.P. Jindal Global University, Sonipat, Haryana, India.

Deepawali Sharma, Department of Computer Science, Banaras Hindu University, Varanasi, India.

Vivek Kumar Singh, Department of Computer Science, Banaras Hindu University, Varanasi, India.

Summary

The increase in abusive content on online social media platforms is impacting the social life of online users. Use of offensive and hate speech has been making social media toxic. Homophobia and transphobia constitute offensive comments against LGBT + community. It becomes imperative to detect and handle these comments, to timely flag or issue a warning to users indulging in such behaviour.

However, automated detection of such content is a challenging task, more so in Dravidian languages which are identified as low resource languages. Motivated by this, the paper attempts to explore applicability of different deep learning models for classification of the social media comments in Malayalam and Tamil languages as homophobic, transphobic and non-anti-LGBT + content.

The popularly used deep learning models-Convolutional Neural Network (CNN), Long Short Term Memory (LSTM) using GloVe embedding and transformer-based learning models (Multilingual BERT and IndicBERT) are applied to the classification problem.

Results obtained show that IndicBERT outperforms the other implemented models, with obtained weighted average F1-score of 0.86 and 0.77 for Malayalam and Tamil, respectively. Therefore, the present work confirms higher performance of IndicBERT on the given task on selected Dravidian languages.

Published in: Woungang, I., Dhurandher, S.K., Pattanaik, K.K., Verma, A., Verma, P. (eds) Advanced Network Technologies and Intelligent Computing. ANTIC 2022. Communications in Computer and Information Science, vol 1798. Springer, Cham.

To read the full paper, please click here.