Social Sciences › Decision Sciences › Management Science and Operations Research
Multi-Criteria Decision Making
27 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- GAW-PO: Preference Optimization with Gradient-Aligned Token Weights
Andreea Dutulescu, Stefan Ruseti, Mihai Masala, Traian Rebedea, Mihai Dascalu · 2 de octubre de 2026
Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, inc…
- Robust Nash Alignment under Preference Uncertainty
Shihab Ahmed, Debamita Ghosh, David Tang, Yudan Wang, Alvaro Velasquez, Yue Wang · 2 de octubre de 2026
Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for al…
- Gradient-Aligned Pair Selection for Personalized Preference Optimization
Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou · 2 de octubre de 2026
Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends…
- Masking Frequent Tokens Sharpens Direct Preference Optimization
Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu · 30 de septiembre de 2026
Direct Preference Optimization (DPO) aligns language models by optimizing over sequence-level sums of token-wise implicit reward differences. However, we identify a pervasive pathology in this formulation: a disproportionately small subset of high-frequency token types dominates cumulative sequence …
- From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact
Jiangtao Lin, Bangyang Wei, Siyi Liu, Yihang Ding, Yuhan Dong · 29 de septiembre de 2026
Soft preference targets specify supervision strength, and reward objectives convert that strength into learned reward signals. A central design question remains: how does assigning a fixed set of preference strengths to different response pairs change the rewards produced by different objectives? We…
- Matrix Aggregation Operators
Inmaculada Guti\'errez (Faculty of Statistical Studies, Complutense University of Madrid, Instituto Universitario de Estad\'istica y Ciencia de Datos, Complutense University of Madrid), Asier Urio-Larrea (Department of Statistics, Computer Science and Mathematics, Universidad P\'ublica de Navarra, Institute of Smart Cities), J. Tinguaro Rodr\'iguez (Faculty of Mathematics, Complutense University of Madrid, Instituto de Matem\'atica Interdisciplinar, Complutense University of Madrid), Daniel G\'omez (Faculty of Statistical Studies, Complutense University of Madrid, Instituto Universitario de Estad\'istica y Ciencia de Datos, Complutense University of Madrid), Javier Montero (Faculty of Mathematics, Complutense University of Madrid, Instituto de Matem\'atica Interdisciplinar, Complutense University of Madrid), Humberto Bustince (Department of Statistics, Computer Science and Mathematics, Universidad P\'ublica de Navarra, Institute of Smart Cities) · 25 de septiembre de 2026
Aggregation theory has traditionally focused on operators defined over vectors. However, many applications-including Multi-Criteria Decision Making, Group Decision Making, Fuzzy Rule-Based Classification Systems, and overlap/grouping indices-require aggregating information naturally structured as a …
- Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment
David Tsoi, Esra D\"onmez · 24 de septiembre de 2026
People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective wei…
- Multiple latent orderings better predict language model preferences
Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young · 22 de septiembre de 2026
Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats suc…
- Direct Preference Density Alignment for Conversational Audio Equalization
Ioannis Stylianou, Sven Ewan Shepstone, Jon Francombe, Pablo Martinez Nuevo, Zheng-Hua Tan · 14 de septiembre de 2026
Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they…
- Emotional Preferences as Goal-Priority Regulation
Shiqi Liu, Yihua Tan, Hu Fu, Guanyu Qi · 28 de agosto de 2026
A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. Under changing external environments and evolving i…
- Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment
Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang · 27 de agosto de 2026
We consider the problem of learning a mixture of $k$ Plackett-Luce models given multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Brad…
- Triangular Fuzzy Rescaling Distance
Eddy Soria, Aida Valls, Ana Beatriz Hern\'andez-Lara · 21 de agosto de 2026
Decision-making in complex systems often involves dealing with imprecise or uncertain information, frequently represented using fuzzy sets, particularly Triangular Fuzzy Numbers (TFNs). A crucial aspect of many fuzzy methods is the quantification of distance between TFNs. Many distance measures assu…
- MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
Tony Tu, Sayan Chakraborty, Ruomeng Xu, Tony Qin, Austin Tian · 18 de agosto de 2026
Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no re…
- A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions
Di Yang Shi, W. Bradley Knox · 13 de agosto de 2026
We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural language, our process produces a linear reward function in three steps…
- Population-Level Generative Modeling for Ranking Data
Zhaoyang Shi · 11 de agosto de 2026
Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback. Existing statistical work has primarily focused on inference tasks such as preference estimation, rank aggreg…
- Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process
Vitaliy Tsyganok, Sergii Kadenko, Oleh Andriichuk · 7 de agosto de 2026
The paper addresses several ranking-dependent decision support methods. Ordinal information on compared objects can be used to improve the quality of expert data during estimation and help reduce the number of comparisons that the experts need to perform. In the paper we compare three incomplete ran…
- Decision-oriented joint optimization of evidence fusion based on event-conditioned credibility
Chaoxiong Ma, Yan Liang, Huixia Zhang, Hao Sun · 30 de julio de 2026
In decision-level fusion tasks involving heterogeneous sources with unequal precision and potential anomalies, evidence deviating from the majority may be either critical evidence supporting the correct decision or anomalous evidence supporting an incorrect event. Existing credible evidence fusion m…
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Siheng Wang, Haoyan Xu, Yuqi Li, Chenhao Wei, Zhengdao Li, Rongchao Zhang, Guang Yang, Yidong Wang, Junhao Dong · 29 de julio de 2026
Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimizati…
- Geometric mean-based pairwise comparison method with the reference values -- statistical approach
Konrad Ku{\l}akowski, Jacek Szybowski · 14 de julio de 2026
For many years, the pairwise comparison method has been widely used for decision-making involving experts. The best-known example of this method is the Analytic Hierarchy Process (AHP). In this now classic approach, the weights of alternatives are calculated using the principal eigenvector of the co…
- POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy Process
Kevin Kam Fung Yuen · 9 de julio de 2026
Pairwise comparison (PC) via pairwise reciprocal matrices (PRMs) is central to the Analytic Hierarchy Process (AHP). Although the traditional eigenvector method is widely applied to derive priorities, its theoretical robustness in reflecting true priority vectors remains debated. Building upon a pre…
- Double Fuzzy Probabilistic Interval Linguistic Term Set and a Dynamic Fuzzy Decision Making Model based on Markov Process with tts Application in Multiple Criteria Group Decision Making
Zongmin Liu · 7 de julio de 2026
The probabilistic linguistic term has been proposed to deal with probability distributions in provided linguistic evaluations. However, because it has some fundamental defects, it is often difficult for decision-makers to get reasonable information of linguistic evaluations for group decision making…
- A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu · 10 de junio de 2026
With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) has emerged as a promising approach for alignment, acting as an RL-free alternative to Reinforcement Learning from Human Fe…
- TOPSIS-RAD: Ranking According to Desires
Leonardo Fernandes Costa, Helder Gomes Costa, Diogo Lima, Brunno Rodrigues · 8 de junio de 2026
Traditional TOPSIS derives its reference points -- the Positive Ideal Solution ($PIS$) and Negative Ideal Solution ($NIS$) -- from the observed alternative set, making rankings susceptible to misalignment with decision-maker (DM) requirements, sensitivity to outlier performances, and rank reversal. …
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
Hyung Gyu Rho · 2 de junio de 2026
Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models. However, its reliance on a fixed temperature parameter leads to suboptimal training on diverse preference data, causing overfitting on easy examples and under-learning from informati…
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
Xuan Qi, Rongwu Xu, Zhijing Jin · 19 de mayo de 2026
Aligning large language models (LLMs) with human preferences is a critical challenge in AI research. While methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) are widely used, they often rely on large, costly preference datasets. The current work l…
Otros asuntos del tema Investigación operativa y ciencias de la gestión
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Advanced Bandit Algorithms Research697 artículos / 12 meses+31 %
- Stock Market Forecasting Methods391 artículos / 12 meses+420 %
- Forecasting Techniques and Applications300 artículos / 12 meses+700 %
- Data Quality and Management254 artículos / 12 meses+1650 %
- Auction Theory and Applications98 artículos / 12 meses+100 %
- Risk and Portfolio Optimization98 artículos / 12 meses+233 %
