<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Preprint | Daniel Ebi</title><link>https://ebida.github.io/tags/preprint/</link><atom:link href="https://ebida.github.io/tags/preprint/index.xml" rel="self" type="application/rss+xml"/><description>Preprint</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 05 Feb 2026 00:00:00 +0000</lastBuildDate><image><url>https://ebida.github.io/media/icon_hu8655581192770655996.png</url><title>Preprint</title><link>https://ebida.github.io/tags/preprint/</link></image><item><title>[Updated Preprint] Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access</title><link>https://ebida.github.io/posts/preprint-iaac-informativeness_v2/</link><pubDate>Thu, 05 Feb 2026 00:00:00 +0000</pubDate><guid>https://ebida.github.io/posts/preprint-iaac-informativeness_v2/</guid><description>&lt;p>&lt;strong>Preprint.&lt;/strong> &lt;em>Not peer-reviewed.&lt;/em>&lt;/p>
&lt;h2 id="abstract">Abstract&lt;/h2>
&lt;p>Asymmetric actor-critic methods are widely used in partially observable reinforcement learning, but typically assume full state observability to condition the critic during training, which is often unrealistic in practice.&lt;/p>
&lt;p>We introduce the informed asymmetric actor-critic framework, allowing the critic to be conditioned on arbitrary state-dependent privileged signals without requiring access to the full state. We show that any such privileged signal yields unbiased policy gradient estimates, substantially expanding the set of admissible privileged information. This raises the problem of selecting the most adequate privileged information in order to improve learning. For this purpose, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a criterion based on improvements in value prediction accuracy that can be applied post-hoc.&lt;/p>
&lt;p>Empirical results on partially observable benchmark tasks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.&lt;/p>
&lt;h2 id="read-more">Read more&lt;/h2>
&lt;ul>
&lt;li>Access the preprint on &lt;a href="https://arxiv.org/abs/2509.26000v2">arXiv&lt;/a>.&lt;/li>
&lt;/ul></description></item><item><title>[New Preprint] Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access</title><link>https://ebida.github.io/posts/preprint-iaac-informativeness_v1/</link><pubDate>Tue, 30 Sep 2025 00:00:00 +0000</pubDate><guid>https://ebida.github.io/posts/preprint-iaac-informativeness_v1/</guid><description>&lt;p>&lt;strong>Preprint.&lt;/strong> &lt;em>Not peer-reviewed.&lt;/em>&lt;/p>
&lt;h2 id="abstract">Abstract&lt;/h2>
&lt;p>Reinforcement learning in partially observable environments requires agents to act under uncertainty from noisy, incomplete observations. Asymmetric actor-critic methods leverage privileged information during training to improve learning under these conditions. However, existing approaches typically assume full-state access during training.&lt;/p>
&lt;p>In this work, we challenge this assumption by proposing a novel actor-critic framework, called informed asymmetric actor-critic, that enables conditioning the critic on arbitrary privileged signals without requiring access to the full state. We show that policy gradients remain unbiased under this formulation, extending the theoretical foundation of asymmetric methods to the more general case of privileged partial information. To quantify the impact of such signals, we propose informativeness measures based on kernel methods and return prediction error, providing practical tools for evaluating training-time signals. We validate our approach empirically on benchmark navigation tasks and synthetic partially observable environments, showing that our informed asymmetric method improves learning efficiency and value estimation when informative privileged inputs are available.&lt;/p>
&lt;p>Our findings challenge the necessity of full-state access and open new directions for designing asymmetric reinforcement learning methods that are both practical and theoretically sound.&lt;/p>
&lt;h2 id="read-more">Read more&lt;/h2>
&lt;ul>
&lt;li>Access the preprint on &lt;a href="https://arxiv.org/abs/2509.26000v1">arXiv&lt;/a>.&lt;/li>
&lt;/ul></description></item></channel></rss>