Paper Sets Minimax Regret Bounds for Bandits with Probing Feedback
A new arXiv paper studies a bandit setting where a learner may probe up to k of n arms per round and only observes the highest reward among those probed, rather than each individual reward. The authors derive two minimax laws characterizing when probing yields a statistical advantage over standard bandit learning, covering independent stochastic reward models. They also identify limits on what can be learned from this winner-only feedback.