Physical Sciences › Computer Science › Artificial Intelligence
Hate Speech and Cyberbullying Detection
390 indexierte Paper
Die Beobachtung von Veröffentlichungen auf arXiv im Bereich der künstlichen Intelligenz zeigt eine Reihe von Arbeiten, die sich der Erkennung von Hassrede und Cybermobbing widmen. Diese Forschungen untersuchen Methoden zur Identifizierung toxischer Inhalte in verschiedenen Kontexten, sei es in textuellen Nachrichten, Dialogen, die von Sprachmodellen generiert werden, oder sogar in visuellen Elementen wie Memes. Die untersuchten Ansätze umfassen Klassifizierungstechniken, kontextuelle Analyseframeworks sowie Strategien zur Anpassung von Moderationswerkzeugen an die kulturellen und sprachlichen Besonderheiten der betroffenen Communities.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten34 % · 84 Artikel
- China21 % · 52 Artikel
- Indien10 % · 25 Artikel
- Deutschland6,8 % · 17 Artikel
- Vereinigtes Königreich6 % · 15 Artikel
- Bangladesch5,2 % · 13 Artikel
- Italien4,8 % · 12 Artikel
- Kanada4,4 % · 11 Artikel
Über 249 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 59 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm
Laura Wagner, Eva Cetinic · 29. September 2026
Open-source text-to-image (TTI) pipelines have become dominant in the landscape of AI-generated visual content, driven by technological advances that enable users to personalize models through adapters tailored to specific tasks. While personalization methods such as LoRA offer unprecedented creativ…
- MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante · 28. September 2026
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the t…
- ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
Firoj Alam, Md. Rafiul Biswas, Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Hamdy Mubarak, George Mikros, Abul Hasnat, Wajdi Zaghouani · 25. September 2026
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In total, 58 teams registered, 35 participat…
- An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection
Rameesha Zia, Muhammad Shahid Iqbal Malik · 25. September 2026
Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary classification, evaluated their frameworks on a single dataset, and provid…
- Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier
Adam G\'orski, Mateusz J\k{a}kalak, Rafa{\l} Jakubowski · 25. September 2026
We develop a multi-label Polish content-safety classifier by fine-tuning allegro/herbert-base-cased (124M) across five categories (hate, vulgarity, sexual content, crime, self-harm) using a Focal + R-Drop objective, and evaluate the resulting model against Bielik Guard (S\'ojka) on the shared out-of…
- SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments
Zixiang Xu, Yanbo Wang, Yue Huang, Haomin Zhuang, Yujun Zhou, Jiayi Ye, Sixian Li, Zirui Song, Lang Gao, Chenxi Wang, Zhaorun Chen, Wang Pan, Yue Zhao, Jieyu Zhao, Xiangliang Zhang, Xiuying Chen · 22. September 2026
Large language models (LLMs) are increasingly deployed in socially grounded applications, where success requires interpreting context, inferring others' mental states, and reasoning about unreliable information. Yet existing benchmarks rarely evaluate these demands jointly in complex, evolving setti…
- Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation
Zeeshan Ahmed, Yang Qin, Hanqing Huang · 22. September 2026
Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining for every policy change and suffering from label scarcity since multimedia cannot be meaningfully augmented. We propose Summarize-Judge-Refine (SJR), a t…
- Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection
Ashanvi Yadav, Shubham Bhardwaj · 22. September 2026
Lexicon-driven misogyny detectors cannot, by construction, distinguish a slur used against a woman from the same slur mentioned in counter-speech ("don't call her that") -- yet exactly this distinction governs whether moderation protects or silences the people discussing abuse. We study this problem…
- MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction
Akshit Sharma, Prashant W. Patil · 21. September 2026
The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers…
- Unifying Models of Intergroup Hostility in Online Discourse
Patrick Gerard, Julia Mendelsohn, Kristina Lerman · 18. September 2026
Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational theories in social and moral psychology, and political science. Howe…
- Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection
Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia · 18. September 2026
Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explaina…
- Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Sarah Wyer, Sue Black, Noura Al Moubayed · 18. September 2026
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{har…
- Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators
Yibo Hu · 17. September 2026
Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks (BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat, and ToxiGen) into a …
- Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues
Arefeh Kazemi, Hamza Qadeer, Sinan Asci, Joachim Wagner, Brian Davis · 17. September 2026
Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether …
- Control-Theoretic Content Moderation
Benedetta Tessa, Serena Tardelli, Marco Avvenuti, Anna Monreale, Stefano Cresci · 17. September 2026
A sizable literature studies content moderation locally, at the level of individual moderation decisions, for example by measuring or predicting the effects of specific interventions. However, the problem of how such decisions should be combined into effective platform-level moderation strategies is…
- Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech
Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso, Serena Villata · 16. September 2026
Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, frequently produce generic, ineffective replies that fail to target the implicit…
- ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian
Zahra Bokaei, Walid Magdy, Bonnie Webber · 16. September 2026
We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identifica…
- Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truic\u{a}, Elena-Simona Apostol · 16. September 2026
Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposing a human-centered multimodal framework that analyses and incorporates the psychological and demographic characteristics of human annotators into the …
- Automated Comment Moderation Enhances Social Media Advertising Performance
Jiwoon Park, Julian De Freitas · 16. September 2026
Social media advertising exposes brands not only to potential customers but also to unfiltered consumer discourse in the form of user comments. While comments can enhance authenticity and engagement, they also introduce reputational risks through spam, hate speech, and negative user-generated conten…
- How User-AI Mistreatment Occurs and Matters in Conversational Systems?
Fanqi Zeng, Sadid A. Hasan, Chaocheng He · 15. September 2026
Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper,…
- Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
Joshua Wolfe Brook, Ilia Markov · 11. September 2026
This paper investigates the use of an LLM to generate auxiliary background context for social media posts, and explores four methods to incorporate this context into the input of an SBERT-based Hate Speech Detection (HSD) classifier. These are: text concatenation, embedding concatenation, a hierarch…
- Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms
Pushpdeep Singh, Sayeh Jarollahi, Ayan Majumdar, Vabuk Pahari, Abhijnan Chakraborty, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash · 11. September 2026
Empirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for indepe…
- MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short
Kristin Gnadt, Maximilian Meidinger, Matthias A{\ss}enmacher · 10. September 2026
With hate speech being ubiquitous online, automatic detection is crucial, in particular when it comes to criminally relevant social media posts. We study a variety of retrieval-based in-context learning (RetICL) strategies for detecting defamatory offences under {\S}{\S} 185-187 StGB (the subject of…
- BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language Models
Yongxin Deng (University of Technology Sydney), Xiaoyu Tan (National University of Singapore), Jing Pan (Monash University), Ling Chen (University of Technology Sydney), Zhen Fang (University of Technology Sydney), Xihe Qiu (National University of Singapore) · 10. September 2026
Large language models (LLMs) may encode biased associations from heterogeneous training corpora that are not immediately visible under ordinary prompting, but can surface when the model is steered toward particular demographic personas. Such behavior often manifests not as explicit toxic output, but…
- Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety
Hamed Jelodar, Amir Firouzi, Yen-Wu Lo, Maryam Tanha, Sajjad Dadkhah · 10. September 2026
Healthcare systems, mental health, and public well-being are increasingly affected by cyberbullying and harmful online interactions. This paper presents CareGuard, an early-warning framework designed to support healthcare-driven mental health protection and proactive online safety through the detect…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.136 Papiere / 12 Monate+653 %
- Adversarial Robustness in Machine Learning3.480 Papiere / 12 Monate+491 %
- Reinforcement Learning in Robotics2.430 Papiere / 12 Monate+329 %
- Explainable Artificial Intelligence (XAI)2.262 Papiere / 12 Monate+521 %
- Domain Adaptation and Few-Shot Learning1.999 Papiere / 12 Monate+281 %
- Advanced Graph Neural Networks1.875 Papiere / 12 Monate+162 %
