Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access
Apr 30, 2026·
,,,·
0 min read
Daniel Ebi
Damien Ernst
Klemens Böhm
Gaspard Lambrechts
Abstract
Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability. Existing asymmetric actor-critic methods typically assume access to the full environment state to condition the critic during training, which is often unrealistic in practice. We introduce the informed asymmetric actor-critic framework that allows the critic to be conditioned on arbitrary state-dependent privileged signals, and show that any such signal yields unbiased policy gradient estimates. This substantially expands the set of admissible privileged information and raises the problem of selecting the most informative signals for learning. To this end, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a test based on improvements in value prediction that can be applied post hoc. Experiments on partially observable benchmarks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.
Type
Publication
In Forty-third International Conference on Machine Learning (ICML 2026)