Off-policy evaluation is known to suffer from high variance in large action spaces. Recent estimators leverage existing structure to reduce the problem dimensionality by using action embeddings. Yet, which properties lead embeddings to be useful for downstream evaluation remains an open question. To answer it, we benchmark several embeddings in a variety of synthetic environments. We observe that even if they exist, the causal action embeddings may not lead to the lowest error in downstream estimation. We then analyze the sensitivity of our findings through several ablations, and highlight that the presence of flat, redundant regions in the reward function, as well the dependency of embeddings on the reward, are key to reducing the variance of embeddings-based estimators.

Related Research