Physical Sciences › Computer Science › Artificial Intelligence
Hate Speech and Cyberbullying Detection
390 artículos indexados
El análisis de las publicaciones en arXiv en el ámbito de la inteligencia artificial muestra un conjunto de trabajos dedicados a la detección de discursos de odio y ciberacoso. Estas investigaciones exploran métodos para identificar contenidos tóxicos en contextos variados, ya sean mensajes textuales, diálogos generados por modelos de lenguaje o incluso elementos visuales como los memes. Los enfoques estudiados incluyen técnicas de clasificación, marcos de análisis contextual o estrategias para adaptar las herramientas de moderación a las especificidades culturales y lingüísticas de las comunidades involucradas.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos34 % · 84 artículos
- China21 % · 52 artículos
- India10 % · 25 artículos
- Alemania6,8 % · 17 artículos
- Reino Unido6 % · 15 artículos
- Bangladés5,2 % · 13 artículos
- Italia4,8 % · 12 artículos
- Canadá4,4 % · 11 artículos
Sobre 249 artículos de este tema con al menos un laboratorio localizado. 59 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information
Tian Lan, Xiaoqing Cheng, Han Zhang, Jiang Li · 2 de octubre de 2026
Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at ca…
- The Surge of Anti-Semitism in German Social Media following the October 7 Attacks
Gregor Wiedemann, Daniel Wehrend · 1 de octubre de 2026
We investigate the extent to which the Hamas attacks on Israel of October 7, 2023, have affected German social media debates about Judaism and Israel. For this, we develop an approach to detect 26 anti-Semitic categories in user postings via large language models (LLMs). The approach is applied to F…
- Look What You Made Us Cluster: Hate Narrative Extraction from Reddit Discourse
Annabelle K. L. Chua, Forster J. Khoo, Joel C. R. Tan, Huey Ting Ang, Kheng Hwee Tan, Joel Y. A. Sim, Shirley W. H. Ow, Ria Mundhra, Elsie C. K. Toh, Youfeng Xu, Lynnette H. X. Ng · 1 de octubre de 2026
Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational approaches, however, are limited in precision as they rely on semantic representations, which tend to capture only surface-level meaning. To detect mor…
- N\"urnberg NLP at ChildSafeAds 2026: Structurally Dissimilar Voter Ensembles under Four Levels of Data Access
Philipp Steigerwald, Eric Rudolph, Jens Albrecht · 30 de septiembre de 2026
We describe the N\"urnberg NLP system for ChildSafeAds 2026. The shared task asks what a monitoring system for commercial content in child-facing YouTube videos can achieve at a given level of data access. We answer with per-subtask ensembles of nine voters, organised into three branches that differ…
- In-game Toxic Detection: Bi-directional Representations with Attention Residuals
Yuanzhe Jia · 30 de septiembre de 2026
In-game toxic language has emerged as a critical concern in the gaming industry and community. While several frameworks and models for online game toxicity analysis have been proposed, detecting toxicity in player chat utterances remains a formidable challenge: stemming not only from the extremely s…
- Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Girish A. Koushik, Diptesh Kanojia, Helen Treharne · 30 de septiembre de 2026
When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, a…
- Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm
Laura Wagner, Eva Cetinic · 29 de septiembre de 2026
Open-source text-to-image (TTI) pipelines have become dominant in the landscape of AI-generated visual content, driven by technological advances that enable users to personalize models through adapters tailored to specific tasks. While personalization methods such as LoRA offer unprecedented creativ…
- MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante · 28 de septiembre de 2026
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the t…
- ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
Firoj Alam, Md. Rafiul Biswas, Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Hamdy Mubarak, George Mikros, Abul Hasnat, Wajdi Zaghouani · 25 de septiembre de 2026
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In total, 58 teams registered, 35 participat…
- An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection
Rameesha Zia, Muhammad Shahid Iqbal Malik · 25 de septiembre de 2026
Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary classification, evaluated their frameworks on a single dataset, and provid…
- Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier
Adam G\'orski, Mateusz J\k{a}kalak, Rafa{\l} Jakubowski · 25 de septiembre de 2026
We develop a multi-label Polish content-safety classifier by fine-tuning allegro/herbert-base-cased (124M) across five categories (hate, vulgarity, sexual content, crime, self-harm) using a Focal + R-Drop objective, and evaluate the resulting model against Bielik Guard (S\'ojka) on the shared out-of…
- SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments
Zixiang Xu, Yanbo Wang, Yue Huang, Haomin Zhuang, Yujun Zhou, Jiayi Ye, Sixian Li, Zirui Song, Lang Gao, Chenxi Wang, Zhaorun Chen, Wang Pan, Yue Zhao, Jieyu Zhao, Xiangliang Zhang, Xiuying Chen · 22 de septiembre de 2026
Large language models (LLMs) are increasingly deployed in socially grounded applications, where success requires interpreting context, inferring others' mental states, and reasoning about unreliable information. Yet existing benchmarks rarely evaluate these demands jointly in complex, evolving setti…
- Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation
Zeeshan Ahmed, Yang Qin, Hanqing Huang · 22 de septiembre de 2026
Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining for every policy change and suffering from label scarcity since multimedia cannot be meaningfully augmented. We propose Summarize-Judge-Refine (SJR), a t…
- Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection
Ashanvi Yadav, Shubham Bhardwaj · 22 de septiembre de 2026
Lexicon-driven misogyny detectors cannot, by construction, distinguish a slur used against a woman from the same slur mentioned in counter-speech ("don't call her that") -- yet exactly this distinction governs whether moderation protects or silences the people discussing abuse. We study this problem…
- MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction
Akshit Sharma, Prashant W. Patil · 21 de septiembre de 2026
The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers…
- Unifying Models of Intergroup Hostility in Online Discourse
Patrick Gerard, Julia Mendelsohn, Kristina Lerman · 18 de septiembre de 2026
Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational theories in social and moral psychology, and political science. Howe…
- Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection
Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia · 18 de septiembre de 2026
Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explaina…
- Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Sarah Wyer, Sue Black, Noura Al Moubayed · 18 de septiembre de 2026
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{har…
- Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators
Yibo Hu · 17 de septiembre de 2026
Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks (BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat, and ToxiGen) into a …
- Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues
Arefeh Kazemi, Hamza Qadeer, Sinan Asci, Joachim Wagner, Brian Davis · 17 de septiembre de 2026
Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether …
- Control-Theoretic Content Moderation
Benedetta Tessa, Serena Tardelli, Marco Avvenuti, Anna Monreale, Stefano Cresci · 17 de septiembre de 2026
A sizable literature studies content moderation locally, at the level of individual moderation decisions, for example by measuring or predicting the effects of specific interventions. However, the problem of how such decisions should be combined into effective platform-level moderation strategies is…
- Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech
Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso, Serena Villata · 16 de septiembre de 2026
Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, frequently produce generic, ineffective replies that fail to target the implicit…
- ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian
Zahra Bokaei, Walid Magdy, Bonnie Webber · 16 de septiembre de 2026
We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identifica…
- Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truic\u{a}, Elena-Simona Apostol · 16 de septiembre de 2026
Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposing a human-centered multimodal framework that analyses and incorporates the psychological and demographic characteristics of human annotators into the …
- Automated Comment Moderation Enhances Social Media Advertising Performance
Jiwoon Park, Julian De Freitas · 16 de septiembre de 2026
Social media advertising exposes brands not only to potential customers but also to unfiltered consumer discourse in the form of user comments. While comments can enhance authenticity and engagement, they also introduce reputational risks through spam, hate speech, and negative user-generated conten…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
