A Numerical Analysis of Allocation Strategies for the Multi Armed Bandit Problem under Delayed Rewards Conditions in Digital Campaign Management
收藏资源简介:
In this paper, we propose a set of allocation strategies to deal with the multi-armed bandit problem, the possibilistic reward (PR) methods. First, we use possibilistic reward distributions to model the uncer- tainty about the expected rewards from the arm, derived from a set of infinite confidence intervals nested around the expected value. Depending on the inequality used to compute the confidence intervals, there are three possible PR methods with different f eatures. Next, we use a pignistic probability transformation to convert these possibilistic functions into probability distributions following the insufficient reason princi- ple . Finally, Thompson sampling techniques are used to identify the arm with the higher expected reward and play that arm. A numerical study analyses the performance of the proposed methods with respect to other policies in the literature. Two PR methods perform well in all representative scenarios under consideration, and are the best allocation strategies if truncated poisson or exponential distributions in [0,10] are considered for the arms.



