The GAN is Dead; Long Live the GAN! A Modern GAN Baseline (2025)

Open in webOpen in zoteroOpen pdf

1 Abstract

There is a widely-spread claim that GANs are difficult to train, and GAN architectures in the literature are littered with empirical tricks. We provide evidence against this claim and build a modern GAN baseline in a more principled manner. First, we derive a well-behaved regularized relativistic GAN loss that addresses issues of mode dropping and non-convergence that were previously tackled via a bag of ad-hoc tricks. We analyze our loss mathematically and prove that it admits local convergence guarantees, unlike most existing relativistic losses. Second, our new loss allows us to discard all ad-hoc tricks and replace outdated backbones used in common GANs with modern architectures. Using StyleGAN2 as an example, we present a road map of simplification and modernization that results in a new minimalist baseline – R3GAN. Despite being simple, our approach surpasses StyleGAN2 on FFHQ, ImageNet, CIFAR, and Stacked MNIST datasets, and compares favorably against state-of-the-art GANs and diffusion models.

2 NOTES

This is an awesome paper on how to improve a complete GAN architecture, taking an already strong one (StyleGAN2), stripping it from a lot of its “techniques”, then building new ones upon it, focusing on simplification.

There are a couple of things that are relevant to speak here:

  1. The losses RpGAN + + : the first one is named relativistic pairing GAN (RpGAN) and is meant to address mode dropping and as can be seen, it has a clear difference from the usual minimax loss, where is applied separately to each term. The idea is that in RpGAN a fake sample is critiqued by its realness relative toa real sample, which effectively maintains a decision boundary in the neighborhood of each real sample and hence forbids mode dropping.
  2. However, RpGAN does not always converge, therefore a gradient penalty, with a similar idea to the GP penalty in WGAN-GP is applied to it. However, in here they are zero-centered, where the two most commonly-used 0-GP are and : penalizes the gradient norm of on real data, and penalizes the gradient of on fake data.

The table below shows all the configurations they tried for their network, after stripping StyleGAN2 from a ton of techniques (some that were only clear when looking at the code).

But the ones that stood out the most for me were:

  • dimension of being only 64
  • extremely small learning rate of
  • Adam due to the small learning rate
  • removal of batch normalization, as they say it is incompatible with , , or RpGAN.
  • much simpler architecture (as can be seem below)

Overall, their results show that they achieved better or almost equal FID as StyleGAN and other models. And in the cases were StyleGAN2 was better, it is argued that they had feature leakage, so it can’t be properly compared. A great paper, the loss is great if it really works so well as they said. Will have to give it a test and will come back here with feedback.