Questions for Research Directions on DreamerV3

I’m researching in Model-bases RL. I implement DreamerV3 and train on DeepMind Control Suite. I benchmark on 4 environments. I try some research directions like representation collapse, compounding error/stability, adaptive imagination horizon, reconstruction-free imagination quality, prior-rollout reward-overestimation. But it failed with 3 reasons:

  1. Variance swamps small effects. Two near-identical configs, same seed, differed 2–4× at a checkpoint on a small (size-1m) model. 10–30% sample-eff gains are basically unmeasurable here without many-seed sweeps I can’t afford everywhere.

  2. The proprio-standard regime is crowded / low-headroom.

  3. Phenomena are scale-dependent. E.g. the prior-rollout reward-overestimation from Biased Dreams (link) didn’t reproduce at classes=4 (it under-estimated), and was just noise across seeds at classes=32.

For rigorous empirical world-model work on a modest budget, what kinds of questions/contributions actually survive high run-to-run variance?

Two smaller ones if anyone has pointers:
(a) any latent-imagination phenomenon that’s scale-robust (shows up even on small models) and still under-explored?
(b) is careful characterization/diagnosis (not need to beat SOTA) still valued at solid venues?

Thanks!

submitted by /u/khoanhat
[link] [comments]

Liked Liked