Social Sciences › Decision Sciences › Management Science and Operations Research
Advanced Bandit Algorithms Research
697 artículos indexados
Los algoritmos de bandits exploran cómo tomar decisiones óptimas bajo incertidumbre, equilibrando entre explotar las opciones conocidas y descubrir nuevas. Este campo se extiende a variantes como los multi-armed bandits, los contextual bandits o los multiplayer bandits, donde las elecciones deben adaptarse a restricciones como recursos limitados, recompensas diferidas o entornos adversarios. Los trabajos recientes abordan también extensiones hacia modelos cuánticos, dinámicas multi-agent o métodos de evaluación robusta, en particular para contextos donde los datos son incompletos o están sesgados.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos46 % · 224 artículos
- China18 % · 86 artículos
- Francia11 % · 53 artículos
- India7,5 % · 36 artículos
- Reino Unido7,3 % · 35 artículos
- Italia5 % · 24 artículos
- Japón4,4 % · 21 artículos
- Corea del Sur3,7 % · 18 artículos
Sobre 482 artículos de este tema con al menos un laboratorio localizado. 49 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit
Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu · 1 de octubre de 2026
We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 202…
- On the Complexity of Preference-Based Bandits
Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY) · 1 de octubre de 2026
We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and …
- Reserve-Aware Contrast Certificates for Conservative Bandits with Uncertain Baselines
Qinchuan Cheng · 1 de octubre de 2026
Conservative bandits must improve an incumbent policy without exhausting a prescribed performance budget. When the incumbent is uncertain, separately bounding candidate and baseline rewards can charge twice for shared estimation error. We develop Reserve-C4B around the baseline-relative contrast its…
- Adaptive Random Matrices in Gaussian Bandits: Spectral Universality and Selection-Induced Outliers
Sudarshan Manikantan (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute) · 28 de septiembre de 2026
Adaptive arm selection changes the distribution of the observations collected by a bandit algorithm, but it need not change their limiting empirical spectrum. We study Gaussian bandit designs in which the dimension and the number of observations grow proportionally. A quantitative coupling theorem c…
- Cost-Aware Best-LLM Identification using Dueling Feedback
Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair · 28 de septiembre de 2026
Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust p…
- Cost-Sensitive Online Window Size Selection for Portfolio Management
Yi-Chen Liu, Chung-Han Hsieh · 25 de septiembre de 2026
This paper investigates cost-sensitive online window size selection for portfolio management under changing market conditions. Specifically, we propose a two-level framework that constructs portfolios using candidate window sizes and dynamically aggregates them through online learning. By treating c…
- Exact Bayes Regret and Asymptotic Optimality in High-Dimensional Gaussian Bandits
Prakhar Singhvi (Abstract Math Institute), Yi Zou (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute) · 25 de septiembre de 2026
We study Bayesian linear bandits with an isotropic Gaussian parameter, independent Gaussian candidate arms, and Gaussian reward noise when the horizon is proportional to the dimension. The normalized posterior uncertainty has an explicit limit that is uniform over all causal policies. Gaussian poste…
- Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits
Michael Jerge, Suman Jana · 25 de septiembre de 2026
Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix defines a node, and its continuations form a subtree below it.…
- Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant
Yang Cai, Vineet Gupta, Yanchen Jiang, Christopher Liaw, Aranyak Mehta, Grigoris Velegkas, Di Wang · 24 de septiembre de 2026
Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate …
- Efficient Linear Bandits via Cluster-Aware Sketching
Hantao Yang, Hong Xie, Defu Lian · 24 de septiembre de 2026
We study the problem of computational efficiency for linear bandits in high-dimensional settings with a finite arm set. In linear bandits, the increase in the dimension $d$ of the feature vectors leads to growing computational costs of $O(d^2)$ at each round of update. Traditional sketching-based me…
- Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices
Qian Xie, Yueli He, Nairen Cao · 23 de septiembre de 2026
Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval, which draws on the Bayesian-optimal Gittins policy to determine which c…
- Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
Jiazheng Sun, Weixin Wang, Pan Xu · 23 de septiembre de 2026
We provide a unified algorithmic framework for ensemble sampling in nonlinear contextual bandits and develop corresponding regret bounds for two most common nonlinear contextual bandit settings: Generalized Linear Model Ensemble Sampling (GLM-ES) for generalized linear contextual bandits and Neural …
- Improved Multiplayer Bandit Algorithm for Bernoulli Rewards
Khang Nguyen, Ricardo Parada, William Chang · 23 de septiembre de 2026
We study the multiplayer multi-armed bandit problem with information asymmetry under Bernoulli rewards, for three information structures: asymmetry in actions, in rewards, and in both. Replacing the Hoeffding-style confidence intervals of prior work with Kullback--Leibler (KL) divergence-based bound…
- Optimal No-Regret Learning for Repeated Prophet Inequality
Kun Wang · 22 de septiembre de 2026
We study repeated prophet inequalities under prefix feedback. In each of $T$ rounds, a learner encounters fresh values drawn independently from $n$ boxes with unknown $[0,1]$-supported distributions in a fixed order and must irrevocably accept one, observing only the prefix up to its stopping box. R…
- The Role of Coordinates in Pareto Regret for Adversarial Multi-Objective Bandits
Changkun Guan, Mengfan Xu · 22 de septiembre de 2026
Adversarial multi-objective bandits hold the potential to help us optimize choices (arms) whose reward is a multidimensional vector chosen by an adversary and whose performance is measured by Pareto regret. We define loss as one minus reward and measure the easiness of a coordinate by the smallest c…
- Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping
Huikang Liu, Zhengchao Wang, Daniel Kuhn, Wolfram Wiesemann · 22 de septiembre de 2026
We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems. Each SAB problem involves choosing between an unknown Bernoulli arm and a known reward. We show that minimizing worst-case regret of SAB problems over all non-antic…
- Rollout Total Correlation for Deep Reinforcement Learning
Bang You, Huaping Liu, Jan Peters, Oleg Arenz · 21 de septiembre de 2026
Learning task-relevant representations is crucial for reinforcement learning. Recent approaches aim to learn such representations by improving the temporal consistency in the observed transitions. However, they only consider individual transitions and can fail to achieve long-term consistency. Inste…
- From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences
Yibo Wang, Wenhao Yang, Sifan Yang, Yuanyu Wan, Lijun Zhang · 21 de septiembre de 2026
In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intri…
- Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits
Sulgi Kim · 18 de septiembre de 2026
Batched multi-armed bandits update on a service's own schedule, and the usual implementation carries each arm's absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thom…
- Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits
Zhaojun Peng · 16 de septiembre de 2026
Decentralized bandit systems often contain heterogeneous agents: rewards can change at individual agents even when the best action for the network stays the same. These local changes may cancel when rewards are averaged across agents, so the number of local changes $\Stloc$ can be much larger than t…
- Gap Entropy and Almost Instance-Wise Optimal Best-Arm Identification
Jiarui Yao, Jiaxi Zhao, Xiangxin Zhou · 15 de septiembre de 2026
In the best-arm identification problem, we are given $n$ stochastic arms with unknown means and wish to identify the arm with the largest mean with probability at least $1-\delta$, using as few samples as possible. We consider independent Gaussian rewards with unit variance and means in $[0,1]$. Che…
- Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients
Johannes M\"uller, Semih \c{C}ayc{\i}, Guido Mont\'ufar · 11 de septiembre de 2026
Kakade's natural policy gradient method has been studied extensively in recent years, showing linear convergence with and without regularization. We study another natural gradient method based on the Fisher information matrix of the state-action distributions which has received little attention from…
- Thompson Sampling for Non-Monotone Convex Ridge Bandits: Monotonicity Is Not Needed for Polynomial Regret
Xuan Li · 11 de septiembre de 2026
Bakhtiari, Lattimore and Szepesv\'ari (COLT 2025) proved that Thompson sampling (TS) has Bayesian regret $\tilde O(d^{5/2}\sqrt n)$ for bandit convex optimisation with convex \emph{monotone} ridge losses $f(x)=\ell(\ip{x}{\theta})$, and asked whether monotonicity of the link is necessary. We give a …
- A positive resolution of the gap-entropy conjecture
P. M. Aronow, Nathan Kallus, Patrick Lopatto · 10 de septiembre de 2026
We prove the gap-entropy conjecture for fixed-confidence best-arm identification with independent unit-variance Gaussian arms, means in $[0,1]$, and a unique optimal arm. For each suboptimal arm $i$, let $\Delta_i=\mu_*-\mu_i$ be its gap from the optimal mean, and write $H=\sum_{i\ne *}\Delta_i^{-2}…
- Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits
Hao Li, Jie Xu, Zheng Xie · 10 de septiembre de 2026
Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action set…
Otros asuntos del tema Investigación operativa y ciencias de la gestión
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Stock Market Forecasting Methods391 artículos / 12 meses+420 %
- Forecasting Techniques and Applications300 artículos / 12 meses+700 %
- Data Quality and Management254 artículos / 12 meses+1650 %
- Auction Theory and Applications98 artículos / 12 meses+100 %
- Risk and Portfolio Optimization98 artículos / 12 meses+233 %
- Game Theory and Applications75 artículos / 12 meses+200 %
