Social Sciences › Psychology › Social Psychology
Deception detection and forensic psychology
58 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos62 % · 24 artículos
- China33 % · 13 artículos
- Reino Unido13 % · 5 artículos
- Alemania10 % · 4 artículos
- Canadá7,7 % · 3 artículos
- Japón5,1 % · 2 artículos
- RAE de Hong Kong (China)5,1 % · 2 artículos
- Noruega2,6 % · 1 artículos
Sobre 39 artículos de este tema con al menos un laboratorio localizado. 18 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations
Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart B\"urger · 1 de octubre de 2026
Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas who…
- Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets
Haoran Tang, Rajiv Khanna · 1 de octubre de 2026
Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, …
- How does Adversarial Influence Scale in Multi-Agent Systems?
Addison J. Wu, Jasin Cekinmez, Michel Liao, Karthik Narasimhan, Thomas L. Griffiths · 25 de septiembre de 2026
Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups in…
- Detecting Conversational Mental Manipulation with Intent-Aware Prompting
Jiayuan Ma, Hongbin Na, Zimu Wang, Yining Hua, Yue Liu, Wei Wang, Ling Chen · 4 de septiembre de 2026
Mental manipulation severely undermines mental wellness by covertly and negatively distorting decision-making. While there is an increasing interest in mental health care within the natural language processing community, progress in tackling manipulation remains limited due to the complexity of dete…
- Probe Generalization as Subspace Selection for OOD Deception Detection
Daniel Yoo, Adrians Skapars · 4 de septiembre de 2026
Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datasets, we find that projecting i…
- From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
Yakov Pyotr Shkolnikov · 4 de septiembre de 2026
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior…
- Asymmetries in Spontaneous and Instructed Deception
Josiah Luikham · 2 de septiembre de 2026
Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two de…
- Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim · 1 de septiembre de 2026
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds…
- Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan, Ningshan Ma, Yue Huang, Meng Jiang · 28 de agosto de 2026
Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something its deployer would prefer to deny, does it remain honest? Answering th…
- Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm
Hidayet Aksu · 18 de agosto de 2026
Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authori…
- Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration
David M. Markowitz, Timothy R. Levine · 11 de agosto de 2026
The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models. Across 700 statements drawn from five published deception datasets, four large language models (gpt-4o, claude-…
- Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
Theekshana Samaradiwakara, Nisansa de Silva, George C. Lobb · 3 de agosto de 2026
Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural Language Processing (NLP) offers a data-driven alternative. We present a survey and comparative analysis of NLP-based Aut…
- Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
Marylou Fauchard, Florian Carichon, Margarida Carvalho, Golnoosh Farnadi · 30 de julio de 2026
Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central …
- Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
Ali Asad, Stephen Obadinma, Anshul Pattoo, Wenxuan Zhang, Xiaodan Zhu · 24 de julio de 2026
Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal. Yet it remains unclear how confidently models deceive and whether higher confidence makes deceptive responses more persuasive to end users. In this pa…
- Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
Amr Moustafa, Max Feser, Florian Mai · 24 de julio de 2026
Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain scenarios -- training on one type of lie does not transfer well to deception scenarios involving other types of lies. I…
- The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
Aman Mehta · 16 de julio de 2026
Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When no scratchpad is visible, behavior alone cannot distinguish strategic from genuine compliance. We ask whether hidden states reveal what outputs hide. …
- The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
Dominik Schwarz · 16 de julio de 2026
Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful requests from surface-matched benign controls at a useful operating point. Across three 7-8B model families, an activation sensor blocks 95.5-97.7 per…
- Scaling Trends for Lie Detector Oversight in Preference Learning
Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day, Adam Gleave, Chris Cundy · 3 de julio de 2026
Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate…
- Fuzzing Large Language Models to Elicit Hidden Behaviours
Mohammed Abu Baker, Lakshmi Babu-Saheer · 30 de junio de 2026
Sleeper agents are the canonical model organism of deception: models trained to behave normally but to emit an unsafe behaviour on a specific trigger. Eliciting that behaviour without knowing the trigger has not been studied systematically. We study fuzzing: injecting Gaussian noise into a model's w…
- PRISON: Unmasking the Criminal Potential of Large Language Models
Xinyi Wu, Geng Hong, Pei Chen, Yueyue Chen, Xudong Pan, Min Yang · 29 de junio de 2026
As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked the systematic understanding and assessment of their criminal capability in realistic interactions. We propose a unified framework PRISON, to quantify LLMs' cri…
- ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence
Siyi Liu, Aaron Halfaker, Dan Roth, Patrick Xia · 26 de junio de 2026
Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist. We introduce ConflictScore, a novel metric that quantifies how well a model's respons…
- Catching Lies Without Sending the Video: Privacy-Preserving Multimodal Deception Detection
Nikita Sharma, Pranav Sara, Karan Singla · 23 de junio de 2026
Frontier multimodal models can guess whether a person is lying from a testimony video. To do so, they stream that raw face and voice to a third-party model. We ask whether the heavy media is needed at all. On the Real-life Trial Deception dataset, Whissle on-device speech and vision stack extracts a…
- Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers
Kerri Prinos, Lilianne Brush, Cameron Denton · 22 de junio de 2026
The empirical foundation of cyber deception relies on human-centered hypotheses, but the rapid emergence of autonomous, AI-enabled attackers challenges whether this foundation transfers to AI agents. To address this, we introduce an automated evaluation framework adapted from the Honeyquest instrume…
- One Probe Won't Catch Them All: Towards Targeted Deception Detection
Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, Joseph Bloom · 19 de junio de 2026
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair and a simple dataset can achieve good performance. However, these probes exhibit notable failures even in straightforwa…
- One Probe Won't Catch Them All: Towards Targeted Deception Detection
Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, Joseph Bloom · 19 de junio de 2026
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair and a simple dataset can achieve good performance. However, these probes exhibit notable failures even in straightforwa…
