A context-aware attention and graph neural network-based multimodal framework for misogyny detection

Rehman, Mohammad Zia Ur; Kumar, Nagendra

Please use this identifier to cite or link to this item: https://dspace.iiti.ac.in/handle/123456789/14760

Full metadata record

DC Field	Value	Language
dc.contributor.author	Rehman, Mohammad Zia Ur	en_US
dc.contributor.author	Kumar, Nagendra	en_US
dc.date.accessioned	2024-10-25T05:51:01Z	-
dc.date.available	2024-10-25T05:51:01Z	-
dc.date.issued	2025	-
dc.identifier.citation	Rehman, M. Z. U., Zahoor, S., Manzoor, A., Maqbool, M., & Kumar, N. (2025). A context-aware attention and graph neural network-based multimodal framework for misogyny detection. Information Processing and Management. Scopus. https://doi.org/10.1016/j.ipm.2024.103895	en_US
dc.identifier.issn	0306-4573	-
dc.identifier.other	EID(2-s2.0-85204472925)	-
dc.identifier.uri	https://doi.org/10.1016/j.ipm.2024.103895	-
dc.identifier.uri	https://dspace.iiti.ac.in/handle/123456789/14760	-
dc.description.abstract	A substantial portion of offensive content on social media is directed towards women. Since the approaches for general offensive content detection face a challenge in detecting misogynistic content, it requires solutions tailored to address offensive content against women. To this end, we propose a novel multimodal framework for the detection of misogynistic and sexist content. The framework comprises three modules: the Multimodal Attention module (MANM), the Graph-based Feature Reconstruction Module (GFRM), and the Content-specific Features Learning Module (CFLM). The MANM employs adaptive gating-based multimodal context-aware attention, enabling the model to focus on relevant visual and textual information and generating contextually relevant features. The GFRM module utilizes graphs to refine features within individual modalities, while the CFLM focuses on learning text and image-specific features such as toxicity features and caption features. Additionally, we curate a set of misogynous lexicons to compute the misogyny-specific lexicon score from the text. We apply test-time augmentation in feature space to better generalize the predictions on diverse inputs. The performance of the proposed approach has been evaluated on two multimodal datasets, MAMI, and MMHS150K, with 11,000 and 13,494 samples, respectively. The proposed method demonstrates an average improvement of 11.87% and 10.82% in macro-F1 over existing multimodal methods on the MAMI and MMHS150K datasets, respectively. © 2024 Elsevier Ltd	en_US
dc.language.iso	en	en_US
dc.publisher	Elsevier Ltd	en_US
dc.source	Information Processing and Management	en_US
dc.subject	Data fusion	en_US
dc.subject	Deep learning	en_US
dc.subject	Hate speech against women	en_US
dc.subject	Misogyny detection	en_US
dc.subject	Multimodal learning	en_US
dc.subject	Sexism detection	en_US
dc.title	A context-aware attention and graph neural network-based multimodal framework for misogyny detection	en_US
dc.type	Journal Article	en_US
Appears in Collections:	Department of Computer Science and Engineering

Files in This Item:

There are no files associated with this item.

Show simple item record

Altmetric Badge: