Off-policy evaluation is known to suffer from high variance in large action spaces. Recent estimators leverage existing structure to reduce the problem dimensionality by using action embeddings. Yet, which properties lead embeddings to be useful for downstream evaluation remains an open question. To answer it, we benchmark several embeddings in a variety of synthetic environments. We observe that even if they exist, the causal action embeddings may not lead to the lowest error in downstream estimation. We then analyze the sensitivity of our findings through several ablations, and highlight that the presence of flat, redundant regions in the reward function, as well the dependency of embeddings on the reward, are key to reducing the variance of embeddings-based estimators.
Related Research
-
Can LLMs Access their World Knowledge for Event Prediction?
Can LLMs Access their World Knowledge for Event Prediction?
Z. Wu, Z. Zhang, H. Hajimirsadeghi, N. Dvornik, and A. Pashevich. EMNLP
Publications
-
Jump Start or False Start? A Theoretical and Empirical Evaluation of LLM-initialized Bandits
Jump Start or False Start? A Theoretical and Empirical Evaluation of LLM-initialized Bandits
A. Bailey, X. Zhu, R. Aoki, Y. Cao, and K. Wilson. TMLR
Publications
-
Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization
Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization
Y. Wen, Y. Cao, and L. Mou. CIKM
Publications