TY - GEN
T1 - Ax-to-Grind Urdu
T2 - 22nd IEEE International Conference on Trust, Security and Privacy in Computing and Communications, TrustCom 2023
AU - Harris, Sheetal
AU - Liu, Jinshuo
AU - Hadi, Hassan Jalil
AU - Cao, Yue
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2024/5/29
Y1 - 2024/5/29
N2 - Misinformation can seriously impact society, affecting anything from public opinion to institutional confidence and the political horizon of a state. Fake News (FN) proliferation on online websites and Online Social Networks (OSNs) has increased profusely. Various fact-checking websites include news in English and barely provide information about FN in regional languages. Thus the Urdu FN purveyors cannot be discerned using fact-checking portals. State-of-the-art (SOTA) approaches for Fake News Detection (FND) count upon appropriately labelled and large datasets. FND in regional and resource-constrained languages lags due to the lack of limited-sized datasets and legitimate lexical resources. The previous datasets for Urdu FND are limited-sized, domain-restricted, publicly unavailable and not manually verified where the news is translated from English into Urdu. In this paper, we curate and contribute the first largest publicly available dataset for Urdu FND, "Ax-to-Grind Urdu", to bridge the identified gaps and limitations of existing Urdu datasets in the literature. It constitutes 10,083 fake and real news on fifteen domains collected from leading and authentic Urdu newspapers and news channel websites in Pakistan and India. FN for the Ax-to-Grind dataset is collected from websites and crowdsourcing. The dataset contains news items in Urdu from the year 2017 to the year 2023. Expert journalists annotated the dataset. We benchmark the dataset with an ensemble model of mBERT, XLNet, and XLM-RoBERTa. The selected models are originally trained on multilingual large corpora. The results of the proposed model are based on performance metrics, F1-score, accuracy, precision, recall and MCC value. F1-score of 0.924, accuracy of 0.956, precision of 0.942, recall of 0.940 and an MCC value of 0.902 demonstrate the effectiveness of the proposed approach for Urdu FND. Comparison analysis with SOTA ML and DL models and existing Urdu benchmark datasets exhibit that the ensemble model outperforms them for Urdu FND. The dataset used for our experiments is publicly available at https://github.com/HjH-Whu-CRC/Ax-to-Grind-Urdu for further analysis and validation.
AB - Misinformation can seriously impact society, affecting anything from public opinion to institutional confidence and the political horizon of a state. Fake News (FN) proliferation on online websites and Online Social Networks (OSNs) has increased profusely. Various fact-checking websites include news in English and barely provide information about FN in regional languages. Thus the Urdu FN purveyors cannot be discerned using fact-checking portals. State-of-the-art (SOTA) approaches for Fake News Detection (FND) count upon appropriately labelled and large datasets. FND in regional and resource-constrained languages lags due to the lack of limited-sized datasets and legitimate lexical resources. The previous datasets for Urdu FND are limited-sized, domain-restricted, publicly unavailable and not manually verified where the news is translated from English into Urdu. In this paper, we curate and contribute the first largest publicly available dataset for Urdu FND, "Ax-to-Grind Urdu", to bridge the identified gaps and limitations of existing Urdu datasets in the literature. It constitutes 10,083 fake and real news on fifteen domains collected from leading and authentic Urdu newspapers and news channel websites in Pakistan and India. FN for the Ax-to-Grind dataset is collected from websites and crowdsourcing. The dataset contains news items in Urdu from the year 2017 to the year 2023. Expert journalists annotated the dataset. We benchmark the dataset with an ensemble model of mBERT, XLNet, and XLM-RoBERTa. The selected models are originally trained on multilingual large corpora. The results of the proposed model are based on performance metrics, F1-score, accuracy, precision, recall and MCC value. F1-score of 0.924, accuracy of 0.956, precision of 0.942, recall of 0.940 and an MCC value of 0.902 demonstrate the effectiveness of the proposed approach for Urdu FND. Comparison analysis with SOTA ML and DL models and existing Urdu benchmark datasets exhibit that the ensemble model outperforms them for Urdu FND. The dataset used for our experiments is publicly available at https://github.com/HjH-Whu-CRC/Ax-to-Grind-Urdu for further analysis and validation.
KW - Ensemble model
KW - FND
KW - NLP
KW - Urdu Corpus
UR - https://www.scopus.com/pages/publications/85195495268
UR - https://www.scopus.com/pages/publications/85195495268#tab=citedBy
U2 - 10.1109/TrustCom60117.2023.00343
DO - 10.1109/TrustCom60117.2023.00343
M3 - Conference proceeding (ISBN)
AN - SCOPUS:85195495268
T3 - Proceedings - 2023 IEEE 22nd International Conference on Trust, Security and Privacy in Computing and Communications, TrustCom/BigDataSE/CSE/EUC/iSCI 2023
SP - 2440
EP - 2447
BT - Proceedings - 2023 IEEE 22nd International Conference on Trust, Security and Privacy in Computing and Communications, TrustCom/BigDataSE/CSE/EUC/iSCI 2023
A2 - Hu, Jia
A2 - Min, Geyong
A2 - Wang, Guojun
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 1 November 2023 through 3 November 2023
ER -