Social Sciences › Social Sciences › Safety Research
Ethics and Social Impacts of AI
1,406 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States48% · 436 papers
- United Kingdom14% · 125 papers
- China12% · 106 papers
- Germany11% · 96 papers
- Canada7.5% · 68 papers
- India4.4% · 40 papers
- Australia4.1% · 37 papers
- Switzerland4% · 36 papers
Across 906 papers on this subject with at least one lab located. 80 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- SCM-based Fairness and Faithful Explainability for Legal Document Classification
Yasmina El Kacemi, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag · 2 October 2026
Transformer models such as LegalBERT are increasingly used in legal decision support, raising concerns about both fairness and the transparency of model explanations. These properties are usually evaluated separately, leaving open whether a debiasing intervention that changes fairness also changes h…
- Multicalibration for Unbiased Model-Based Prevalence Estimation
Fridolin Linder, Thomas Leeper, Daniel Haimovich, Niek Tax, Lorenzo Perini, Milan Vojnovic · 2 October 2026
Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is fundamental to science, public health, and online trust and safety. Standard approaches correct for known device error rates but assume these rates…
- The AI Assessment Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes
Alessio Buscemi, German Castignani, Daniele Pagani, Maxime Cordy, Jordi Cabot · 2 October 2026
The EU's Artificial Intelligence Act requires all Member States to establish AI Regulatory Sandboxes (AIRS) by August 2027: supervised environments bringing together national Competent Authorities, technical experts, and the organisations under assessment. When AIRS engagements include structured te…
- Architecture Without an Architect? Global Governance of Artificial Intelligence in a Divided World
Simon Chesterman · 2 October 2026
Artificial intelligence presents an unusually difficult problem for global governance. The technology develops rapidly, crosses borders easily, and is shaped by actors whose resources and capabilities may rival those of states. Yet international responses remain fragmented, unevenly representative, …
- White Men Without Degrees Receive the Lowest Ratings from Large Language Models
Maxim Chupilkin · 2 October 2026
White men without an undergraduate degree receive the lowest average ratings among eight gender-race-education groups in controlled large-language-model evaluations of credit, hiring, and rental applications. We conduct full-factorial vignette experiments with 18 models from 12 developer groups, var…
- Efficient Active Auditing of Multi-Group Fairness with Bias Probes
Ayoub Ajarra, Debabrota Basu · 1 October 2026
Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable …
- Fairness Theatre: Evaluating Post-Hoc Fairness Interventions in Vendor-Controlled Early Warning Systems
Kelly McConvey, Angelina Zhai, Rebecca Li, Shion Guha · 1 October 2026
Public institutions increasingly procure AI systems whose design they cannot inspect or change. In higher education, proprietary Early Warning Systems (EWS) leave colleges with few options beyond adjusting model outputs to address inequity. This raises the question of how fairness work is coordinate…
- A Reusable Semantic Web Framework for Evidence-Grounded Fundamental Rights Impact Assessments under the EU AI Act
Faith Olopade, Delaram Golpayegani, David Lewis · 1 October 2026
The EU AI Act (Art. 27) requires deployers of high-risk AI systems to conduct Fundamental Rights Impact Assessments (FRIAs) before deployment, yet the evidence needed for credible assessments is fragmented across incompatible incident repositories, risk vocabularies, and legal texts. We present a re…
- The Concentration of Artificial Intelligence in Big Tech and Its Implications for Human Rights in the European Union
Marcin Marciniak · 1 October 2026
The development of advanced artificial intelligence is increasingly concentrated in a small group of vertically integrated technology companies. These firms control combinations of computing infrastructure, cloud services, data, foundation models, software ecosystems, and channels of distribution. T…
- The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents
Achint Mehta · 30 September 2026
Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screenshot tool. We call this set the agent's verification surface. This study asks whether increasing only that surface, with everything else held fixed, produ…
- A Comprehensive View of Fairness through Distributional Stability
Gayane Taturyan, Charlotte Laclau, Stephan Cl\'emencon · 30 September 2026
We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under th…
- An Empirical Study and Assessment of EU AI Act Compliance Checkers
Zhen Tao, Alize Kahraman, Shidong Pan, Zhenchang Xing, Chiara Ullstein, Jens Grossklags, Chunyang Chen · 30 September 2026
The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements are directly relevant to security and privacy, while also addressing closely related issues such as data governance, transparency, accuracy, and robu…
- Algorithmic Harms Associated with Generative Model-Augmented Recommendation Systems
Christine Herlihy, Xumei Xi, Shloka Desai, Kevin Bannerman Hutchful, Pedro Silva · 29 September 2026
In this work, we consider algorithmic harms that may arise as generative models are incorporated into machine learning platforms. We argue that existing harm taxonomies and threat models require extension to (1) address novel causal drivers of well-studied representational and quality-of-service har…
- Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks
Manya Singh, Arjun Pakrashi · 29 September 2026
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on …
- Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-based Robots
Xuancun Lu, Zhengxian Huang, Xinfeng Li, Chi Zhang, Xiaoyu Ji, Wenyuan Xu · 29 September 2026
LLM-based robots use Large Language Models (LLMs) as planners to translate natural language instructions into policies such as grasp(), move_to(), and open_gripper(). Jailbreak attacks on these robots extend the threat from generating malicious content to executing harmful behaviors. However, we fin…
- Generative AI & Two Forms of Decoupling
Shira Gur-Arieh, Sina Fazelpour · 29 September 2026
Textual artifacts are sometimes valued not only for the words on the page, but for the human activity involved in producing them. The effort invested in a carefully tailored email can signal genuine interest; composing an apology might involve attending to another person's hurt and deciding how to r…
- Adaptive Multi-Value Control in LLMs via Causal Activation Steering
Payel Bhattacharjee, Ravi Tandon · 28 September 2026
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. Howev…
- Statistical attribute alignment for black-box generative AI via output post-processing
Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski · 28 September 2026
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, whe…
- The Gold in Bias: Maturing the AI Design Process through Verification
Samira Maghool, Paolo Ceravolo · 25 September 2026
Bias in AI systems is typically framed as a flaw to be minimized, yet it also serves as a critical indicator of underlying weaknesses in data, modeling assumptions, and system design. Existing approaches often treat bias as an isolated problem rather than as evidence that can strengthen verification…
- When No One Owns the Judgment: Accountability Under Contribution Dissolution in Human-AI Collaboration
Hengzhi Ye · 25 September 2026
Communities often respond to potentially AI-assisted work by asking three questions: Was AI used? Was that use disclosed? Can hidden use be detected? These questions place AI use itself at the center of accountability while overlooking a deeper problem: unowned judgment. Evaluations, claims, decisio…
- Fair Like Us? Auditing LLM Alignment in Resource Allocation
Qishen Han, Hadi Hosseini, Joshua Kavner, Samarth Khanna, Sujoy Sikdar, Lirong Xia · 25 September 2026
Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they …
- Signed Exposure: Fair Routing of Algorithmic Attention When Attention Can Harm
Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov · 25 September 2026
Fairness-of-exposure treats algorithmic attention as a good to be distributed equitably. But when an autonomous agent initiates contact, attention is signed: it delivers value to a willing receiver and imposes a burden on an unwilling one. We formalize routing under signed exposure and show that a f…
- RADAR: Readiness for AI Discovery and Agentic Reach
Luke Jordan, Tiago C. Peixoto, Manuel Ramos-Maqueda · 25 September 2026
Governments increasingly meet citizens through an AI system rather than a website. RADAR (Readiness for AI Discovery and Agentic Reach) measures whether that system works, across 166 countries and on two tasks: whether a chatbot can give a correct, officially sourced, country-specific answer about a…
- When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations
Nithin Raghava Ramachandra Narla · 24 September 2026
Fairness audits in production ML typically occur once, at deployment, on a single domain. Both fail in practice: fairness can shift after retraining or a changing user base, and interventions validated on one dataset are rarely tested across the heterogeneous domains an organization deploys. We pres…
- Who Acts, Who Knows, Who Answers? A Corpus-Assisted Discourse Analysis of Agency, Epistemic Responsibility, and Accountability in Generative AI Higher Education Research
Biranchi Poudyal · 24 September 2026
Generative artificial intelligence (GenAI) is increasingly described in higher education as a tool, collaborator, evaluator, proxy, and infrastructure. These labels are not neutral: they shape who is seen as acting, knowing, and, crucially, being answerable when someone uses AI. This study examines …
