Social Sciences › Decision Sciences › Management Science and Operations Research
Advanced Bandit Algorithms Research
697 indexierte Paper
Bandit-Algorithmen untersuchen, wie optimale Entscheidungen unter Unsicherheit getroffen werden können, indem sie zwischen der Nutzung bekannter Optionen und der Entdeckung neuer abwägen. Dieses Gebiet erstreckt sich auf Varianten wie multi-armed bandits, contextual bandits oder multiplayer bandits, bei denen die Entscheidungen an Einschränkungen wie begrenzte Ressourcen, verzögerte Rückmeldungen oder adversarische Umgebungen angepasst werden müssen. Aktuelle Arbeiten befassen sich auch mit Erweiterungen hin zu quantenbasierten Modellen, multi-agent-Dynamiken oder robusten Evaluierungsmethoden, insbesondere für Kontexte, in denen Daten unvollständig oder verzerrt sind.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten46 % · 224 Artikel
- China18 % · 86 Artikel
- Frankreich11 % · 53 Artikel
- Indien7,5 % · 36 Artikel
- Vereinigtes Königreich7,3 % · 35 Artikel
- Italien5 % · 24 Artikel
- Japan4,4 % · 21 Artikel
- Südkorea3,7 % · 18 Artikel
Über 482 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 49 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit
Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu · 1. Oktober 2026
We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 202…
- On the Complexity of Preference-Based Bandits
Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY) · 1. Oktober 2026
We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and …
- Reserve-Aware Contrast Certificates for Conservative Bandits with Uncertain Baselines
Qinchuan Cheng · 1. Oktober 2026
Conservative bandits must improve an incumbent policy without exhausting a prescribed performance budget. When the incumbent is uncertain, separately bounding candidate and baseline rewards can charge twice for shared estimation error. We develop Reserve-C4B around the baseline-relative contrast its…
- Adaptive Random Matrices in Gaussian Bandits: Spectral Universality and Selection-Induced Outliers
Sudarshan Manikantan (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute) · 28. September 2026
Adaptive arm selection changes the distribution of the observations collected by a bandit algorithm, but it need not change their limiting empirical spectrum. We study Gaussian bandit designs in which the dimension and the number of observations grow proportionally. A quantitative coupling theorem c…
- Cost-Aware Best-LLM Identification using Dueling Feedback
Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair · 28. September 2026
Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust p…
- Cost-Sensitive Online Window Size Selection for Portfolio Management
Yi-Chen Liu, Chung-Han Hsieh · 25. September 2026
This paper investigates cost-sensitive online window size selection for portfolio management under changing market conditions. Specifically, we propose a two-level framework that constructs portfolios using candidate window sizes and dynamically aggregates them through online learning. By treating c…
- Exact Bayes Regret and Asymptotic Optimality in High-Dimensional Gaussian Bandits
Prakhar Singhvi (Abstract Math Institute), Yi Zou (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute) · 25. September 2026
We study Bayesian linear bandits with an isotropic Gaussian parameter, independent Gaussian candidate arms, and Gaussian reward noise when the horizon is proportional to the dimension. The normalized posterior uncertainty has an explicit limit that is uniform over all causal policies. Gaussian poste…
- Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits
Michael Jerge, Suman Jana · 25. September 2026
Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix defines a node, and its continuations form a subtree below it.…
- Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant
Yang Cai, Vineet Gupta, Yanchen Jiang, Christopher Liaw, Aranyak Mehta, Grigoris Velegkas, Di Wang · 24. September 2026
Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate …
- Efficient Linear Bandits via Cluster-Aware Sketching
Hantao Yang, Hong Xie, Defu Lian · 24. September 2026
We study the problem of computational efficiency for linear bandits in high-dimensional settings with a finite arm set. In linear bandits, the increase in the dimension $d$ of the feature vectors leads to growing computational costs of $O(d^2)$ at each round of update. Traditional sketching-based me…
- Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices
Qian Xie, Yueli He, Nairen Cao · 23. September 2026
Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval, which draws on the Bayesian-optimal Gittins policy to determine which c…
- Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
Jiazheng Sun, Weixin Wang, Pan Xu · 23. September 2026
We provide a unified algorithmic framework for ensemble sampling in nonlinear contextual bandits and develop corresponding regret bounds for two most common nonlinear contextual bandit settings: Generalized Linear Model Ensemble Sampling (GLM-ES) for generalized linear contextual bandits and Neural …
- Improved Multiplayer Bandit Algorithm for Bernoulli Rewards
Khang Nguyen, Ricardo Parada, William Chang · 23. September 2026
We study the multiplayer multi-armed bandit problem with information asymmetry under Bernoulli rewards, for three information structures: asymmetry in actions, in rewards, and in both. Replacing the Hoeffding-style confidence intervals of prior work with Kullback--Leibler (KL) divergence-based bound…
- Optimal No-Regret Learning for Repeated Prophet Inequality
Kun Wang · 22. September 2026
We study repeated prophet inequalities under prefix feedback. In each of $T$ rounds, a learner encounters fresh values drawn independently from $n$ boxes with unknown $[0,1]$-supported distributions in a fixed order and must irrevocably accept one, observing only the prefix up to its stopping box. R…
- The Role of Coordinates in Pareto Regret for Adversarial Multi-Objective Bandits
Changkun Guan, Mengfan Xu · 22. September 2026
Adversarial multi-objective bandits hold the potential to help us optimize choices (arms) whose reward is a multidimensional vector chosen by an adversary and whose performance is measured by Pareto regret. We define loss as one minus reward and measure the easiness of a coordinate by the smallest c…
- Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping
Huikang Liu, Zhengchao Wang, Daniel Kuhn, Wolfram Wiesemann · 22. September 2026
We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems. Each SAB problem involves choosing between an unknown Bernoulli arm and a known reward. We show that minimizing worst-case regret of SAB problems over all non-antic…
- Rollout Total Correlation for Deep Reinforcement Learning
Bang You, Huaping Liu, Jan Peters, Oleg Arenz · 21. September 2026
Learning task-relevant representations is crucial for reinforcement learning. Recent approaches aim to learn such representations by improving the temporal consistency in the observed transitions. However, they only consider individual transitions and can fail to achieve long-term consistency. Inste…
- From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences
Yibo Wang, Wenhao Yang, Sifan Yang, Yuanyu Wan, Lijun Zhang · 21. September 2026
In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intri…
- Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits
Sulgi Kim · 18. September 2026
Batched multi-armed bandits update on a service's own schedule, and the usual implementation carries each arm's absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thom…
- Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits
Zhaojun Peng · 16. September 2026
Decentralized bandit systems often contain heterogeneous agents: rewards can change at individual agents even when the best action for the network stays the same. These local changes may cancel when rewards are averaged across agents, so the number of local changes $\Stloc$ can be much larger than t…
- Gap Entropy and Almost Instance-Wise Optimal Best-Arm Identification
Jiarui Yao, Jiaxi Zhao, Xiangxin Zhou · 15. September 2026
In the best-arm identification problem, we are given $n$ stochastic arms with unknown means and wish to identify the arm with the largest mean with probability at least $1-\delta$, using as few samples as possible. We consider independent Gaussian rewards with unit variance and means in $[0,1]$. Che…
- Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients
Johannes M\"uller, Semih \c{C}ayc{\i}, Guido Mont\'ufar · 11. September 2026
Kakade's natural policy gradient method has been studied extensively in recent years, showing linear convergence with and without regularization. We study another natural gradient method based on the Fisher information matrix of the state-action distributions which has received little attention from…
- Thompson Sampling for Non-Monotone Convex Ridge Bandits: Monotonicity Is Not Needed for Polynomial Regret
Xuan Li · 11. September 2026
Bakhtiari, Lattimore and Szepesv\'ari (COLT 2025) proved that Thompson sampling (TS) has Bayesian regret $\tilde O(d^{5/2}\sqrt n)$ for bandit convex optimisation with convex \emph{monotone} ridge losses $f(x)=\ell(\ip{x}{\theta})$, and asked whether monotonicity of the link is necessary. We give a …
- A positive resolution of the gap-entropy conjecture
P. M. Aronow, Nathan Kallus, Patrick Lopatto · 10. September 2026
We prove the gap-entropy conjecture for fixed-confidence best-arm identification with independent unit-variance Gaussian arms, means in $[0,1]$, and a unique optimal arm. For each suboptimal arm $i$, let $\Delta_i=\mu_*-\mu_i$ be its gap from the optimal mean, and write $H=\sum_{i\ne *}\Delta_i^{-2}…
- Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits
Hao Li, Jie Xu, Zheng Xie · 10. September 2026
Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action set…
Weitere Unterthemen aus Operations Research und Managementwissenschaft
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Stock Market Forecasting Methods391 Papiere / 12 Monate+420 %
- Forecasting Techniques and Applications300 Papiere / 12 Monate+700 %
- Data Quality and Management254 Papiere / 12 Monate+1650 %
- Auction Theory and Applications98 Papiere / 12 Monate+100 %
- Risk and Portfolio Optimization98 Papiere / 12 Monate+233 %
- Game Theory and Applications75 Papiere / 12 Monate+200 %
