Questions for Research Directions on DreamerV3
I’m researching in Model-bases RL. I implement DreamerV3 and train on DeepMind Control Suite. I benchmark on 4 environments. I try some research directions like representation collapse, compounding error/stability, adaptive imagination horizon, reconstruction-free imagination quality, prior-rollout reward-overestimation. But it failed with 3 reasons: Variance swamps small effects. Two near-identical configs, same seed, differed 2–4× at a checkpoint on a small (size-1m) model. 10–30% sample-eff gains are basically unmeasurable here without many-seed sweeps I can’t afford everywhere. The proprio-standard […]