We extend asymmetric actor-critic methods to leverage arbitrary state-dependent privileged signals and introduce criteria to identify which signals most effectively improve learning in partially observable environments.
Apr 30, 2026
We present an asymmetric actor-critic method for partially observable environments that leverages arbitrary privileged information, without requiring full-state access, while preserving unbiased policy gradient estimates.
Jul 17, 2025