Physical Sciences › Computer Science › Information Systems
Digital Rights Management and Security
51 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Derniers papiers
- Uncensored Open-weight Models: Redistribution as the Persistence Layer
10a Labs, :, Juliette Garcia, Hailey May, Bobby McKenzie, David Pham, Matthew Swain, Joshua Valdez, Corie Wieland, Zachary Yahn · 7 septembre 2026
A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471 original uncensored models …
- CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs
Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi, Martin Takac, Salem Lahlou, Nils Lukas · 2 septembre 2026
Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Prefe…
- JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols
Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng, Qian Wang, Kwok-Yan Lam · 28 août 2026
Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not specifically target LLM…
- Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints
Hiroko Takano · 28 août 2026
Multi-stage LLM hiring pipelines (resume improvement, interview question generation, answer feedback) can fabricate credentials, inflate qualifiers, and invent experience. We evaluate two mitigations, prompt guardrails and human-in-the-loop (HITL) checkpoints, against a fully automated baseline. In …
- Training Alignment Auditors via Reinforcement Learning
Paul Rosu, Rowan Wang · 27 août 2026
Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training envi…
- Models in the Same Family are NOT Trust-Equivalent
Rohit Raj Rai, Chirag Kothari, Siddhesh Shelke, Yatika Jena, Amit Awekar · 25 août 2026
Within a model family, a smaller variant is often deployed as a drop-in replacement for a larger one when their performance is similar. However, performance alone does not tell the full story. We propose a framework to evaluate trust-equivalence between a larger model and a smaller one in the same f…
- Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet · 25 août 2026
Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying…
- Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning
Yajie Yin · 20 août 2026
Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to mean at least five different things: veri…
- Debate Training Reduces Reward Hacking in RLAIF
Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah · 19 août 2026
We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as tr…
- Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee · 19 août 2026
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to f…
- Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints
Man Liang, Xinzhao Cheng, Faizan Wajid · 19 août 2026
Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separatin…
- An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez · 19 août 2026
Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model sh…
- Synchronized Logit Steering: Real-world Steganography
Andrew Rufail, Aadi Dash, Onir Narahari, Ethan Mui, Mahi Gajare, Prakhar Tiwari, Shrija Makapothula, Nick Cui · 18 août 2026
Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the sender and receiver to share an identical prompt context, which is rarely guaranteed in production pipelines that use retrieval-augm…
- Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
Aryan Luthra, Kshitij Jain, Siddharth Arya, Bobby Filar, Anna Bertiger · 17 août 2026
Agentic "Continual Learning Harnesses", systems that pair an LLM with retrieval or memory to improve from feedback without retraining, have shown growing value in cybersecurity. But their value is conventionally measured by gains against labeled benchmarks, an approach that often fails in operationa…
- Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation
Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer · 13 août 2026
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolva…
- PPDL: LLM-Based Flows as Probabilistic Programs
Louis Mandel, Guillaume Baudart, Mandana Vaziri, Martin Hirzel · 7 août 2026
Building reliable applications that leverage large language models (LLMs) remains a significant challenge. While LLMs offer impressive capabilities across diverse tasks, their outputs often lack accuracy and provide no clear measure of confidence. This uncertainty compounds in flows of multiple call…
- A Security-Oriented Lifecycle Model for Large Language Model Systems
Eleftherios Batzolis, George Drosatos, Vassilis Katsouros, Konstantinos Rantos · 5 août 2026
Large language models are being integrated into critical infrastructure and enterprise workflows at unprecedented scale,yet the lifecycle frameworks governing their development and operations were designed for operational efficiency rather than security analysis. As a result, security-relevant activ…
- Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
Jingxing Wang, Chenyu Zhou, Zhihui Fu, Jun Wang, Weiwen Liu, Weinan Zhang, Jianghao Lin · 30 juillet 2026
Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not necessarily yield greater agent capability. We call this challenge test-time compute-to-capability conversion and propose SkillTTA, which retrieves task-relevant tra…
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
Shida Wang, Chaohu Liu, Yubo Wang, Linli Xu · 30 juillet 2026
Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable intellectual assets. Nevertheless, these AI assets remain vulnerable to unauthorized redistribution and commercial exploitation through fine-tuning or black-b…
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He · 30 juillet 2026
Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation shows that LLM-authored skills deliver $+0.0$pp over no-skill baselines while human-curated ones deliver $+16.2$pp: the bottleneck is not skill autho…
- Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang, Kai Xie, Xue Liu, Chen Ma, Yiyan Qi · 30 juillet 2026
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to …
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
Md Rezwanul Haque, Md. Milon Islam, Fakhri Karray · 29 juillet 2026
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations…
- LLM Scheming Inversely Scales with Pretraining Language Coverage
Nathan Truong, Aryan Panda, Rayming Ye, Zoe Sun, Maheep Chaudhary · 29 juillet 2026
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning alignment -- in frontier language models, mo…
- From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
James Jewitt, Hao Li, Bram Adams, Gopi Krishnan Rajbahadur, Ahmed E. Hassan · 27 juillet 2026
Hidden license conflicts in the open-source AI ecosystem pose serious legal and ethical risks, exposing organizations to potential litigation and users to undisclosed risk. However, the field lacks a data-driven understanding of how frequently these conflicts occur, where they originate, and which c…
- Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
Kunfeng Lai, Zhenheng Tang, Xinglin Pan, Peijie Dong, Xiang Liu, Haolan Chen, Huacan Wang, Li Shen, Bo Li, Xiaowen Chu · 21 juillet 2026
Model merging aggregates Large Language Models (LLMs) finetuned on different tasks into a stronger one. However, parameter conflicts between models leads to performance degradation in averaging. While model routing addresses this issue by selecting individual models during inference, it imposes exce…
Autres sujets du thème Systèmes d'information
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Software Engineering Research799 papiers / 12 mois+577 %
- Information Retrieval and Search Behavior618 papiers / 12 mois+939 %
- Recommender Systems and Techniques567 papiers / 12 mois+408 %
- Expert finding and Q&A systems154 papiers / 12 mois+975 %
- Information and Cyber Security153 papiers / 12 mois+2600 %
- Blockchain Technology Applications and Security131 papiers / 12 mois+550 %
