Learning Through Energy Refinement and Manifold Projection: A Cooperative EBM-AE Framework
Ryad Zemouri
Abstract
Energy-Based Models (EBMs) provide a flexible framework for generative modeling by learning an energy landscape that assigns low energy values to realistic samples and higher energies to unlikely observations. Despite their theoretical appeal, training EBMs remains challenging due to the computational cost of Langevin sampling and the difficulty of efficiently exploring the learned data manifold. In this work, we propose a cooperative Energy-Based Model and Autoencoder (EBM-AE) framework that combines energy-based refinement with manifold projection. The proposed approach jointly trains an EBM with a denoising autoencoder and introduces an iterative EBM$\rightarrow$AE$\rightarrow$EBM sampling procedure in which Langevin dynamics and autoencoder projection alternately refine generated samples. Within this framework, the autoencoder acts as a manifold projection operator that regularizes sampling trajectories, while the EBM performs energy-based refinement toward low-energy regions of the learned distribution. Extensive experiments conducted on the MNIST dataset demonstrate that joint EBM-AE training substantially improves generation quality compared with a conventional autoencoder. Beyond unconditional generation, we evaluate the proposed framework on image inpainting tasks involving structured and random masks. The results show that manifold projection provides the majority of the reconstruction capability, whereas the final energy-based refinement becomes increasingly beneficial as the reconstruction problem becomes more challenging. Taken together, the results indicate that combining manifold projection and energy minimization provides an effective and interpretable framework for generation, reconstruction, and out-of-distribution detection, while offering new insights into the complementary roles of energy-based modeling and representation learning.