Social Sciences › Decision Sciences › Management Science and Operations Research
Multi-Criteria Decision Making
27 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Neueste Paper
- Masking Frequent Tokens Sharpens Direct Preference Optimization
Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu · 30. September 2026
Direct Preference Optimization (DPO) aligns language models by optimizing over sequence-level sums of token-wise implicit reward differences. However, we identify a pervasive pathology in this formulation: a disproportionately small subset of high-frequency token types dominates cumulative sequence …
- From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact
Jiangtao Lin, Bangyang Wei, Siyi Liu, Yihang Ding, Yuhan Dong · 29. September 2026
Soft preference targets specify supervision strength, and reward objectives convert that strength into learned reward signals. A central design question remains: how does assigning a fixed set of preference strengths to different response pairs change the rewards produced by different objectives? We…
- Matrix Aggregation Operators
Inmaculada Guti\'errez (Faculty of Statistical Studies, Complutense University of Madrid, Instituto Universitario de Estad\'istica y Ciencia de Datos, Complutense University of Madrid), Asier Urio-Larrea (Department of Statistics, Computer Science and Mathematics, Universidad P\'ublica de Navarra, Institute of Smart Cities), J. Tinguaro Rodr\'iguez (Faculty of Mathematics, Complutense University of Madrid, Instituto de Matem\'atica Interdisciplinar, Complutense University of Madrid), Daniel G\'omez (Faculty of Statistical Studies, Complutense University of Madrid, Instituto Universitario de Estad\'istica y Ciencia de Datos, Complutense University of Madrid), Javier Montero (Faculty of Mathematics, Complutense University of Madrid, Instituto de Matem\'atica Interdisciplinar, Complutense University of Madrid), Humberto Bustince (Department of Statistics, Computer Science and Mathematics, Universidad P\'ublica de Navarra, Institute of Smart Cities) · 25. September 2026
Aggregation theory has traditionally focused on operators defined over vectors. However, many applications-including Multi-Criteria Decision Making, Group Decision Making, Fuzzy Rule-Based Classification Systems, and overlap/grouping indices-require aggregating information naturally structured as a …
- Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment
David Tsoi, Esra D\"onmez · 24. September 2026
People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective wei…
- Multiple latent orderings better predict language model preferences
Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young · 22. September 2026
Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats suc…
- Direct Preference Density Alignment for Conversational Audio Equalization
Ioannis Stylianou, Sven Ewan Shepstone, Jon Francombe, Pablo Martinez Nuevo, Zheng-Hua Tan · 14. September 2026
Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they…
- Emotional Preferences as Goal-Priority Regulation
Shiqi Liu, Yihua Tan, Hu Fu, Guanyu Qi · 28. August 2026
A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. Under changing external environments and evolving i…
- Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment
Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang · 27. August 2026
We consider the problem of learning a mixture of $k$ Plackett-Luce models given multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Brad…
- Triangular Fuzzy Rescaling Distance
Eddy Soria, Aida Valls, Ana Beatriz Hern\'andez-Lara · 21. August 2026
Decision-making in complex systems often involves dealing with imprecise or uncertain information, frequently represented using fuzzy sets, particularly Triangular Fuzzy Numbers (TFNs). A crucial aspect of many fuzzy methods is the quantification of distance between TFNs. Many distance measures assu…
- MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
Tony Tu, Sayan Chakraborty, Ruomeng Xu, Tony Qin, Austin Tian · 18. August 2026
Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no re…
- A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions
Di Yang Shi, W. Bradley Knox · 13. August 2026
We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural language, our process produces a linear reward function in three steps…
- Population-Level Generative Modeling for Ranking Data
Zhaoyang Shi · 11. August 2026
Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback. Existing statistical work has primarily focused on inference tasks such as preference estimation, rank aggreg…
- Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process
Vitaliy Tsyganok, Sergii Kadenko, Oleh Andriichuk · 7. August 2026
The paper addresses several ranking-dependent decision support methods. Ordinal information on compared objects can be used to improve the quality of expert data during estimation and help reduce the number of comparisons that the experts need to perform. In the paper we compare three incomplete ran…
- Decision-oriented joint optimization of evidence fusion based on event-conditioned credibility
Chaoxiong Ma, Yan Liang, Huixia Zhang, Hao Sun · 30. Juli 2026
In decision-level fusion tasks involving heterogeneous sources with unequal precision and potential anomalies, evidence deviating from the majority may be either critical evidence supporting the correct decision or anomalous evidence supporting an incorrect event. Existing credible evidence fusion m…
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Siheng Wang, Haoyan Xu, Yuqi Li, Chenhao Wei, Zhengdao Li, Rongchao Zhang, Guang Yang, Yidong Wang, Junhao Dong · 29. Juli 2026
Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimizati…
- Geometric mean-based pairwise comparison method with the reference values -- statistical approach
Konrad Ku{\l}akowski, Jacek Szybowski · 14. Juli 2026
For many years, the pairwise comparison method has been widely used for decision-making involving experts. The best-known example of this method is the Analytic Hierarchy Process (AHP). In this now classic approach, the weights of alternatives are calculated using the principal eigenvector of the co…
- POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy Process
Kevin Kam Fung Yuen · 9. Juli 2026
Pairwise comparison (PC) via pairwise reciprocal matrices (PRMs) is central to the Analytic Hierarchy Process (AHP). Although the traditional eigenvector method is widely applied to derive priorities, its theoretical robustness in reflecting true priority vectors remains debated. Building upon a pre…
- Double Fuzzy Probabilistic Interval Linguistic Term Set and a Dynamic Fuzzy Decision Making Model based on Markov Process with tts Application in Multiple Criteria Group Decision Making
Zongmin Liu · 7. Juli 2026
The probabilistic linguistic term has been proposed to deal with probability distributions in provided linguistic evaluations. However, because it has some fundamental defects, it is often difficult for decision-makers to get reasonable information of linguistic evaluations for group decision making…
- A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu · 10. Juni 2026
With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) has emerged as a promising approach for alignment, acting as an RL-free alternative to Reinforcement Learning from Human Fe…
- TOPSIS-RAD: Ranking According to Desires
Leonardo Fernandes Costa, Helder Gomes Costa, Diogo Lima, Brunno Rodrigues · 8. Juni 2026
Traditional TOPSIS derives its reference points -- the Positive Ideal Solution ($PIS$) and Negative Ideal Solution ($NIS$) -- from the observed alternative set, making rankings susceptible to misalignment with decision-maker (DM) requirements, sensitivity to outlier performances, and rank reversal. …
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
Hyung Gyu Rho · 2. Juni 2026
Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models. However, its reliance on a fixed temperature parameter leads to suboptimal training on diverse preference data, causing overfitting on easy examples and under-learning from informati…
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
Xuan Qi, Rongwu Xu, Zhijing Jin · 19. Mai 2026
Aligning large language models (LLMs) with human preferences is a critical challenge in AI research. While methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) are widely used, they often rely on large, costly preference datasets. The current work l…
- Unweighted ranking for value-based decision making with uncertainty
Aar\'on L\'opez Garc\'ia, Natalia Criado, Jose Such · 14. Mai 2026
As intelligent systems are increasingly implemented in our society to make autonomous decisions, their commitment to human values raises serious concerns. Their alignment with human values remains a critical challenge because it can jeopardise the integrity and security of citizens. For this reason,…
- Sufficient conditions for a Heuristic Rating Estimation Method application
Jacek Szybowski, Konrad Ku{\l}akowski, Jiri Mazurek · 12. Mai 2026
A series of papers has introduced the Heuristic Rating Estimation method, which evaluates a set of alternatives based on pairwise comparisons and the weights of reference alternatives. We formulate the conditions under which the HRE method can be applied correctly. The research considers both arithm…
- Contextual Online Uncertainty-Aware Preference Learning for Human Feedback
Nan Lu, Ethan Lee, Ethan X. Fang, Junwei Lu · 1. Mai 2026
Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm in artificial intelligence to align large models with human preferences. In this paper, we propose a novel statistical framework to simultaneously conduct the online decision-making and statistical inference on the optim…
Weitere Unterthemen aus Operations Research und Managementwissenschaft
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Advanced Bandit Algorithms Research697 Papiere / 12 Monate+31 %
- Stock Market Forecasting Methods391 Papiere / 12 Monate+420 %
- Forecasting Techniques and Applications300 Papiere / 12 Monate+700 %
- Data Quality and Management254 Papiere / 12 Monate+1650 %
- Auction Theory and Applications98 Papiere / 12 Monate+100 %
- Risk and Portfolio Optimization98 Papiere / 12 Monate+233 %
