Threatening language detection from Urdu data with deep sequential model

The Urdu language is spoken and written on different social media platforms like Twitter, WhatsApp, Facebook, and YouTube. However, due to the lack of Urdu Language Processing (ULP) libraries, it is quite challenging to identify threats from textual and sequential data on the social media provided i...

Full description

Saved in:

Bibliographic Details
Published in:	PloS one Vol. 19; no. 6; p. e0290915
Main Authors:	Ullah, Ashraf, Khan, Khair Ullah, Khan, Aurangzeb, Bakhsh, Sheikh Tahir, Rahman, Atta Ur, Akbar, Sajida, Saqia, Bibi
Format:	Journal Article
Language:	English
Published:	United States Public Library of Science 06-06-2024
Subjects:	Biology and Life Sciences Bullying Computational linguistics Computer and Information Sciences Cybercrime Deep Learning Dictionaries Digital media Engineering and Technology Humans Internet Language Language processing Libraries Linguistics Long short-term memory Machine Learning Natural language interfaces Natural Language Processing Social Media Social networks Social Sciences Threats Urdu language United Kingdom
Online Access:	Get full text
Tags:	Add Tag No Tags, Be the first to tag this record!

Description
Summary:	The Urdu language is spoken and written on different social media platforms like Twitter, WhatsApp, Facebook, and YouTube. However, due to the lack of Urdu Language Processing (ULP) libraries, it is quite challenging to identify threats from textual and sequential data on the social media provided in Urdu. Therefore, it is required to preprocess the Urdu data as efficiently as English by creating different stemming and data cleaning libraries for Urdu data. Different lexical and machine learning-based techniques are introduced in the literature, but all of these are limited to the unavailability of online Urdu vocabulary. This research has introduced Urdu language vocabulary, including a stop words list and a stemming dictionary to preprocess Urdu data as efficiently as English. This reduced the input size of the Urdu language sentences and removed redundant and noisy information. Finally, a deep sequential model based on Long Short-Term Memory (LSTM) units is trained on the efficiently preprocessed, evaluated, and tested. Our proposed methodology resulted in good prediction performance, i.e., an accuracy of 82%, which is greater than the existing methods.
Bibliography:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 23 Competing Interests: The authors have declared that no competing interests exist.
ISSN:	1932-6203 1932-6203
DOI:	10.1371/journal.pone.0290915