Physical Sciences › Engineering › Control and Systems Engineering
Advanced Research in Systems and Signal Processing
0 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Noch nicht genug Historie, um die Kurve zu zeichnen.
Neueste Paper
- Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
Sara Rajaram, R. James Cotton, Fabian H. Sinz · 17. Juli 2026
Preference-based Reinforcement Learning (PbRL) entails a variety of approaches for aligning models with human intent to alleviate the burden of reward engineering. However, most previous PbRL work has not investigated the robustness to labeler errors, inevitable with labelers who are non-experts or …
- Adaptive Reinforcement Learning for Unobservable Random Delays
John Wikman, Alexandre Proutiere, David Broman · 14. Juli 2026
In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instantaneously, selects an action without delay, and executes it immediately. In re…
- Inverse Reinforcement Learning with Just Classification and a Few Regressions
Lars van der Laan, Nathan Kallus, Aurelien Bibaut · 11. Mai 2026
Inverse reinforcement learning (IRL) aims to infer rewards from observed behavior, but rewards are not identified from the policy alone: many reward--value pairs can rationalize the same actions. Meaningful reward recovery therefore requires a normalization, yet existing normalized IRL methods often…
- When Are Two RLHF Objectives the Same?
Madhava Gaikwad · 6. Februar 2026
The preference optimization literature contains many proposed objectives, often presented as distinct improvements. We introduce Opal, a canonicalization algorithm that determines whether two preference objectives are algebraically equivalent by producing either a canonical form or a concrete witnes…
- AI Prior Art Search: Semantic Clusters and Evaluation Infrastructure
Boris Genin (Division for the Design of Information Search Systems, Federal Institute of Industrial Property, Berezhkovskaya nab. 30-1, Moscow, 125993, Russian Federation), Alexander Gorbunov (Development Centre for Artificial Intelligence, Federal Institute of Industrial Property, Berezhkovskaya nab. 30-1, Moscow, 125993, Russian Federation), Dmitry Zolkin (Division for the Design of Information Search Systems, Federal Institute of Industrial Property, Berezhkovskaya nab. 30-1, Moscow, 125993, Russian Federation), Igor Nekrasov (Division for the Design of Information Search Systems, Federal Institute of Industrial Property, Berezhkovskaya nab. 30-1, Moscow, 125993, Russian Federation) · 6. Januar 2026
The key to success in automating prior art search in patent research using artificial intelligence (AI) lies in developing large datasets for machine learning (ML) and ensuring their availability. This work is dedicated to providing a comprehensive solution to the problem of creating infrastructure …
- Concentration of Cumulative Reward in Markov Decision Processes
Borna Sayedana, Peter E. Caines, Aditya Mahajan · 4. Dezember 2025
In this paper, we investigate the concentration properties of cumulative reward in Markov Decision Processes (MDPs), focusing on both asymptotic and non-asymptotic settings. We introduce a unified approach to characterize reward concentration in MDPs, covering both infinite-horizon settings (i.e., a…
- On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
Till Freihaut, Giorgia Ramponi · 26. November 2025
Multi-agent Inverse Reinforcement Learning (MAIRL) aims to recover agent reward functions from expert demonstrations. We characterize the feasible reward set in Markov games, identifying all reward functions that rationalize a given equilibrium. However, equilibrium-based observations are often ambi…
- On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian, Nhat Ho, Alessandro Rinaldo · 25. November 2025
The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt, functioning as a new, trainable expert. Despite its populari…
- On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers
Miroslav \v{S}trupl, Oleg Szehr, Francesco Faccio, Dylan R. Ashley, Rupesh Kumar Srivastava, J\"urgen Schmidhuber · 12. November 2025
This article provides a rigorous analysis of convergence and stability of Episodic Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning and Online Decision Transformers. These algorithms performed competitively across various benchmarks, from games to robotic tasks, but their the…
- Structured Reinforcement Learning for Combinatorial Decision-Making
Heiko Hoppe, L\'eo Baty, Louis Bouvier, Axel Parmentier, Maximilian Schiffer · 29. Oktober 2025
Reinforcement learning (RL) is increasingly applied to real-world problems involving complex and structured decisions, such as routing, scheduling, and assortment planning. These settings challenge standard RL algorithms, which struggle to scale, generalize, and exploit structure in the presence of …
Weitere Unterthemen aus Regelungs- und Systemtechnik
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Robot Manipulation and Learning842 Papiere / 12 Monate+833 %
- Human Motion and Animation476 Papiere / 12 Monate+200 %
- Machine Fault Diagnosis Techniques113 Papiere / 12 Monate+550 %
- Traffic control and management91 Papiere / 12 Monate−69 %
- Fault Detection and Control Systems80 Papiere / 12 Monate+450 %
- Smart Grid Security and Resilience61 Papiere / 12 Monate−43 %
