2000 character limit reached
On the connection between Bregman divergence and value in regularized Markov decision processes
Published 21 Oct 2022 in cs.LG, cs.AI, and math.OC | (2210.12160v4)
Abstract: In this short note we derive a relationship between the Bregman divergence from the current policy to the optimal policy and the suboptimality of the current value function in a regularized Markov decision process. This result has implications for multi-task reinforcement learning, offline reinforcement learning, and regret analysis under function approximation, among others.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.