<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Poster Presentation | Daniel Ebi</title><link>https://ebida.github.io/tags/poster-presentation/</link><atom:link href="https://ebida.github.io/tags/poster-presentation/index.xml" rel="self" type="application/rss+xml"/><description>Poster Presentation</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sat, 11 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>https://ebida.github.io/media/icon_hu8655581192770655996.png</url><title>Poster Presentation</title><link>https://ebida.github.io/tags/poster-presentation/</link></image><item><title>"Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access" goes EWRL 2026</title><link>https://ebida.github.io/posts/accepted-iaac-informativeness-ewrl2026/</link><pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate><guid>https://ebida.github.io/posts/accepted-iaac-informativeness-ewrl2026/</guid><description>&lt;p>&lt;strong>Date:&lt;/strong> 5th October 2026, 16:00-18:00 (Poster Session B)&lt;br>
&lt;strong>Location&lt;/strong>: LILLIAD learning center, University of Lille, Lille, France&lt;br>&lt;/p>
&lt;h2 id="abstract">Abstract&lt;/h2>
&lt;p>Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability. Existing asymmetric actor-critic methods typically assume access to the full environment state to condition the critic during training, which is often unrealistic in practice.&lt;/p>
&lt;p>We introduce the informed asymmetric actor-critic framework that allows the critic to be conditioned on arbitrary state-dependent privileged signals, and show that any such signal yields unbiased policy gradient estimates. This substantially expands the set of admissible privileged information and raises the problem of selecting the most informative signals for learning. To this end, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a test based on improvements in value prediction that can be applied post hoc.&lt;/p>
&lt;p>Experiments on partially observable benchmarks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.&lt;/p>
&lt;h2 id="read-more">Read more&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://ewrl-org.github.io/ewrl-2026">Workshop website&lt;/a>&lt;/li>
&lt;li>Read the &lt;a href="../../publications/iaac-informativeness-ewrl2026/IAAC_Ebi_Ernst_Boehm_Lambrechts_EWRL2026_CR.pdf">camera-ready version&lt;/a>.&lt;/li>
&lt;li>Access the &lt;a href="https://openreview.net/forum?id=7G43VVT4SI">conference version&lt;/a> published at ICML 2026.&lt;/li>
&lt;/ul></description></item></channel></rss>