Physical Sciences › Computer Science › Artificial Intelligence
Hate Speech and Cyberbullying Detection
390 papers indexed
Analyzing publications on arXiv in the field of artificial intelligence reveals a body of work dedicated to detecting hate speech and cyberbullying. These studies explore methods for identifying toxic content across various contexts, whether in textual messages, dialogues generated by language models, or even visual elements such as memes. The approaches examined include classification techniques, contextual analysis frameworks, and strategies for adapting moderation tools to the cultural and linguistic specificities of the targeted communities.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States34% · 84 papers
- China21% · 52 papers
- India10% · 25 papers
- Germany6.8% · 17 papers
- United Kingdom6% · 15 papers
- Bangladesh5.2% · 13 papers
- Italy4.8% · 12 papers
- Canada4.4% · 11 papers
Across 249 papers on this subject with at least one lab located. 59 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- The Surge of Anti-Semitism in German Social Media following the October 7 Attacks
Gregor Wiedemann, Daniel Wehrend · 1 October 2026
We investigate the extent to which the Hamas attacks on Israel of October 7, 2023, have affected German social media debates about Judaism and Israel. For this, we develop an approach to detect 26 anti-Semitic categories in user postings via large language models (LLMs). The approach is applied to F…
- Look What You Made Us Cluster: Hate Narrative Extraction from Reddit Discourse
Annabelle K. L. Chua, Forster J. Khoo, Joel C. R. Tan, Huey Ting Ang, Kheng Hwee Tan, Joel Y. A. Sim, Shirley W. H. Ow, Ria Mundhra, Elsie C. K. Toh, Youfeng Xu, Lynnette H. X. Ng · 1 October 2026
Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational approaches, however, are limited in precision as they rely on semantic representations, which tend to capture only surface-level meaning. To detect mor…
- N\"urnberg NLP at ChildSafeAds 2026: Structurally Dissimilar Voter Ensembles under Four Levels of Data Access
Philipp Steigerwald, Eric Rudolph, Jens Albrecht · 30 September 2026
We describe the N\"urnberg NLP system for ChildSafeAds 2026. The shared task asks what a monitoring system for commercial content in child-facing YouTube videos can achieve at a given level of data access. We answer with per-subtask ensembles of nine voters, organised into three branches that differ…
- In-game Toxic Detection: Bi-directional Representations with Attention Residuals
Yuanzhe Jia · 30 September 2026
In-game toxic language has emerged as a critical concern in the gaming industry and community. While several frameworks and models for online game toxicity analysis have been proposed, detecting toxicity in player chat utterances remains a formidable challenge: stemming not only from the extremely s…
- Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Girish A. Koushik, Diptesh Kanojia, Helen Treharne · 30 September 2026
When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, a…
- Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm
Laura Wagner, Eva Cetinic · 29 September 2026
Open-source text-to-image (TTI) pipelines have become dominant in the landscape of AI-generated visual content, driven by technological advances that enable users to personalize models through adapters tailored to specific tasks. While personalization methods such as LoRA offer unprecedented creativ…
- MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante · 28 September 2026
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the t…
- ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
Firoj Alam, Md. Rafiul Biswas, Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Hamdy Mubarak, George Mikros, Abul Hasnat, Wajdi Zaghouani · 25 September 2026
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In total, 58 teams registered, 35 participat…
- An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection
Rameesha Zia, Muhammad Shahid Iqbal Malik · 25 September 2026
Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary classification, evaluated their frameworks on a single dataset, and provid…
- Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier
Adam G\'orski, Mateusz J\k{a}kalak, Rafa{\l} Jakubowski · 25 September 2026
We develop a multi-label Polish content-safety classifier by fine-tuning allegro/herbert-base-cased (124M) across five categories (hate, vulgarity, sexual content, crime, self-harm) using a Focal + R-Drop objective, and evaluate the resulting model against Bielik Guard (S\'ojka) on the shared out-of…
- SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments
Zixiang Xu, Yanbo Wang, Yue Huang, Haomin Zhuang, Yujun Zhou, Jiayi Ye, Sixian Li, Zirui Song, Lang Gao, Chenxi Wang, Zhaorun Chen, Wang Pan, Yue Zhao, Jieyu Zhao, Xiangliang Zhang, Xiuying Chen · 22 September 2026
Large language models (LLMs) are increasingly deployed in socially grounded applications, where success requires interpreting context, inferring others' mental states, and reasoning about unreliable information. Yet existing benchmarks rarely evaluate these demands jointly in complex, evolving setti…
- Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation
Zeeshan Ahmed, Yang Qin, Hanqing Huang · 22 September 2026
Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining for every policy change and suffering from label scarcity since multimedia cannot be meaningfully augmented. We propose Summarize-Judge-Refine (SJR), a t…
- Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection
Ashanvi Yadav, Shubham Bhardwaj · 22 September 2026
Lexicon-driven misogyny detectors cannot, by construction, distinguish a slur used against a woman from the same slur mentioned in counter-speech ("don't call her that") -- yet exactly this distinction governs whether moderation protects or silences the people discussing abuse. We study this problem…
- MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction
Akshit Sharma, Prashant W. Patil · 21 September 2026
The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers…
- Unifying Models of Intergroup Hostility in Online Discourse
Patrick Gerard, Julia Mendelsohn, Kristina Lerman · 18 September 2026
Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational theories in social and moral psychology, and political science. Howe…
- Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection
Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia · 18 September 2026
Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explaina…
- Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Sarah Wyer, Sue Black, Noura Al Moubayed · 18 September 2026
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{har…
- Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators
Yibo Hu · 17 September 2026
Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks (BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat, and ToxiGen) into a …
- Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues
Arefeh Kazemi, Hamza Qadeer, Sinan Asci, Joachim Wagner, Brian Davis · 17 September 2026
Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether …
- Control-Theoretic Content Moderation
Benedetta Tessa, Serena Tardelli, Marco Avvenuti, Anna Monreale, Stefano Cresci · 17 September 2026
A sizable literature studies content moderation locally, at the level of individual moderation decisions, for example by measuring or predicting the effects of specific interventions. However, the problem of how such decisions should be combined into effective platform-level moderation strategies is…
- Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech
Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso, Serena Villata · 16 September 2026
Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, frequently produce generic, ineffective replies that fail to target the implicit…
- ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian
Zahra Bokaei, Walid Magdy, Bonnie Webber · 16 September 2026
We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identifica…
- Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truic\u{a}, Elena-Simona Apostol · 16 September 2026
Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposing a human-centered multimodal framework that analyses and incorporates the psychological and demographic characteristics of human annotators into the …
- Automated Comment Moderation Enhances Social Media Advertising Performance
Jiwoon Park, Julian De Freitas · 16 September 2026
Social media advertising exposes brands not only to potential customers but also to unfiltered consumer discourse in the form of user comments. While comments can enhance authenticity and engagement, they also introduce reputational risks through spam, hate speech, and negative user-generated conten…
- How User-AI Mistreatment Occurs and Matters in Conversational Systems?
Fanqi Zeng, Sadid A. Hasan, Chaocheng He · 15 September 2026
Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper,…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
