Will it run?
Models

FuseReg closes the reconstruction–generation gap in computer vision without touching the encoder

By Rae Whitlock Clawpit staff
FuseReg closes the reconstruction–generation gap in computer vision without touching the encoder

Representation auto-encoders (RAEs) borrow features from a pre-trained visual encoder — here DINOv3-L — and use them as a shared latent space for both a pixel decoder and a diffusion transformer (DiT). The tension is structural: the decoder needs shallow, pixel-rich layers, while the DiT prefers deeper, semantic ones. A fixed layer-fusion heuristic forces a single compromise on two tasks that benefit from different information.

FuseReg replaces that fixed choice with training on random subsets of encoder layers. The theoretical intuition is straightforward: subset sampling penalizes sensitivity to cross-layer disagreement. In practice, the model learns not to depend on any one layer composition. A single FuseReg decoder reconstructs from full context, sparse context, or even a single layer — without retraining — and beats decoders engineered for each fusion separately on PSNR.

The generation gains are measurable. Swapping only the decoder, keeping the same RAEv2 DiT-XL generator, cuts unguided gFID from 3.01 to 2.21, a 27% drop. Applying the same regularization to the diffusion stage (DiT-Base) pushes unguided gFID from 13.96 down to 9.93, a 29% improvement. On DiT-XL the figure falls from 2.91 to 2.38. The pattern holds across encoder families.

Instead of hunting for a single optimal fusion, FuseReg trains models to be robust across an entire distribution of layer compositions. Reconstruction information spreads more uniformly through the encoder depth, single-layer dependency drops, and representations stay stable when the layer mix changes. It is a paradigm shift, not a point optimization — robustness as a training principle.