FuseReg closes the reconstruction–generation gap in computer vision without touching the encoder

Representation auto-encoders (RAEs) borrow features from a pre-trained visual encoder — here DINOv3-L — and use them as a shared latent space for both a pixel decoder and a diffusion transformer (DiT). The tension is structural: the decoder needs shallow, pixel-rich layers, while the DiT prefers deeper, semantic ones. A fixed layer-fusion heuristic forces a single compromise on two tasks that benefit from different information.
FuseReg replaces that fixed choice with training on random subsets of encoder layers. The theoretical intuition is straightforward: subset sampling penalizes sensitivity to cross-layer disagreement. In practice, the model learns not to depend on any one layer composition. A single FuseReg decoder reconstructs from full context, sparse context, or even a single layer — without retraining — and beats decoders engineered for each fusion separately on PSNR.
The generation gains are measurable. Swapping only the decoder, keeping the same RAEv2 DiT-XL generator, cuts unguided gFID from 3.01 to 2.21, a 27% drop. Applying the same regularization to the diffusion stage (DiT-Base) pushes unguided gFID from 13.96 down to 9.93, a 29% improvement. On DiT-XL the figure falls from 2.91 to 2.38. The pattern holds across encoder families.
Instead of hunting for a single optimal fusion, FuseReg trains models to be robust across an entire distribution of layer compositions. Reconstruction information spreads more uniformly through the encoder depth, single-layer dependency drops, and representations stay stable when the layer mix changes. It is a paradigm shift, not a point optimization — robustness as a training principle.