Physical Sciences › Computer Science › Information Systems
Digital Rights Management and Security
55 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility
Kajetan O\.z\'og, Alicja Wojciechowska, Dawid Malarz, Pawe{\l} Batorski, Artur Kasymov, Przemys{\l}aw Spurek · 1 de octubre de 2026
Establishing unbranding as a critical practice to prevent visual logos from acquiring negative connotations is standard in image generation. Large Language Models (LLMs) now face a parallel and emerging challenge. These models frequently generate brand descriptions within diverse contexts. This freq…
- Eternal Sunshine of the Spotless Mind: Systematically Erasing LLM's Memories
Olga Ohrimenko · 1 de octubre de 2026
We consider persistent LLMs that accumulate memories of their interactions with a user over time. Such LLMs maintain memories using external storage, which they can query to overcome the limitations of a fixed context window. Such systems have numerous practical applications, as they can draw on all…
- MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models
Yan Zhang, Chuming Wei, Ruien Li, Yaoyao Peng, Wusheng Zhang, Guangwen Yang · 1 de octubre de 2026
Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and bias, still remains a challenge. Current efforts mainly focus on detecting and filtering inputs and outputs of the trained models, rather than studyi…
- Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu · 30 de septiembre de 2026
As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introd…
- Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference
Taeung Yoon, Yupeng Zhang, Xiaojing Liao · 30 de septiembre de 2026
Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model paramete…
- Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time
Jacob Dineen, Silei Ren, Muhao Chen, Dan Roth, Ben Zhou · 29 de septiembre de 2026
In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model obs…
- From Latents to Wires: Surgical Post-Editing on Large Language Models
Jiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu, Dongdong Yang, Deyue Zhang, Quanchen Zou · 29 de septiembre de 2026
Given a large language model (LLM), can whoever holds the weights name a semantic target (e.g., the model's identity), locate the model components that produce it, and edit them so that the target no longer appears while other capability is preserved? We call such an edit on a trained model a post-e…
- Loyal Agents: Training LLM Agents to Protect Principal Interests Under Strategic Information Asymmetry
Zimeng Huang, Shilei Chen, Jiatong Zhao, Wenxin Xu, Tonghan Wang · 29 de septiembre de 2026
As LLMs increasingly act as delegated agents, they are expected to protect principals' interests when interacting with external parties. Standard alignment objectives, such as helpfulness, harmlessness, and honesty, do not specify how agents should protect principals' strategic interests under deleg…
- LLMs as Linguistic Chameleons: Decoupling Semantics and Structure for Privacy-Preserving Communication
Yuzhu Mao, Liang Zhao · 22 de septiembre de 2026
As Large Language Model (LLM) APIs become increasingly integrated into privacy-sensitive workflows, ensuring inference-time privacy without compromising task utility remains a major challenge. Existing approaches preserve most of the original semantic content to maintain downstream performance, but …
- PolicyMem: Geometric Policy Memory for LLM Governance
Yuanchen Bei, Zhengzhang Chen, Yanjun Zhao, Haoyu Wang, Hanghang Tong, Haifeng Chen · 15 de septiembre de 2026
As large language models (LLMs) are increasingly deployed in real-world high-stakes applications, effective governance has become essential. Existing safeguards largely follow two paradigms: learning-based guards provide strong semantic discrimination but couple policy behavior to trained models and…
- Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
Oliver Daniels, Perusha Moodley, Benjamin M. Marlin, David Lindner · 9 de septiembre de 2026
Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on larg…
- Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses
Cen (Mia), Zhao, Haibo Ruan, Wenjie Chen, Pei-fen Tu, Usman Abbasi, Joel Hesch · 9 de septiembre de 2026
LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness selection for fixed-model multi-turn tool agents, with the search surf…
- Uncensored Open-weight Models: Redistribution as the Persistence Layer
10a Labs, :, Juliette Garcia, Hailey May, Bobby McKenzie, David Pham, Matthew Swain, Joshua Valdez, Corie Wieland, Zachary Yahn · 7 de septiembre de 2026
A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471 original uncensored models …
- CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs
Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi, Martin Takac, Salem Lahlou, Nils Lukas · 2 de septiembre de 2026
Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Prefe…
- JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols
Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng, Qian Wang, Kwok-Yan Lam · 28 de agosto de 2026
Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not specifically target LLM…
- Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints
Hiroko Takano · 28 de agosto de 2026
Multi-stage LLM hiring pipelines (resume improvement, interview question generation, answer feedback) can fabricate credentials, inflate qualifiers, and invent experience. We evaluate two mitigations, prompt guardrails and human-in-the-loop (HITL) checkpoints, against a fully automated baseline. In …
- Training Alignment Auditors via Reinforcement Learning
Paul Rosu, Rowan Wang · 27 de agosto de 2026
Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training envi…
- Models in the Same Family are NOT Trust-Equivalent
Rohit Raj Rai, Chirag Kothari, Siddhesh Shelke, Yatika Jena, Amit Awekar · 25 de agosto de 2026
Within a model family, a smaller variant is often deployed as a drop-in replacement for a larger one when their performance is similar. However, performance alone does not tell the full story. We propose a framework to evaluate trust-equivalence between a larger model and a smaller one in the same f…
- Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet · 25 de agosto de 2026
Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying…
- Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning
Yajie Yin · 20 de agosto de 2026
Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to mean at least five different things: veri…
- Debate Training Reduces Reward Hacking in RLAIF
Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah · 19 de agosto de 2026
We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as tr…
- Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee · 19 de agosto de 2026
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to f…
- Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints
Man Liang, Xinzhao Cheng, Faizan Wajid · 19 de agosto de 2026
Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separatin…
- An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez · 19 de agosto de 2026
Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model sh…
- Synchronized Logit Steering: Real-world Steganography
Andrew Rufail, Aadi Dash, Onir Narahari, Ethan Mui, Mahi Gajare, Prakhar Tiwari, Shrija Makapothula, Nick Cui · 18 de agosto de 2026
Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the sender and receiver to share an identical prompt context, which is rarely guaranteed in production pipelines that use retrieval-augm…
Otros asuntos del tema Sistemas de información
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Software Engineering Research853 artículos / 12 meses+218 %
- Information Retrieval and Search Behavior727 artículos / 12 meses+506 %
- Recommender Systems and Techniques600 artículos / 12 meses+154 %
- Expert finding and Q&A systems205 artículos / 12 meses+1175 %
- Information and Cyber Security169 artículos / 12 meses+1500 %
- Blockchain Technology Applications and Security146 artículos / 12 meses+650 %
