Learning expressive and multimodal policies is essential for solving complex continuous control tasks. However, most reinforcement learning (RL) algorithms rely on unimodal or factorized Gaussian policies, limiting their representational flexibility. While sof... ...