- Access by Xinjiang University
Stabilizing generative adversarial networks with adaptive noise injection
Phys. Rev. E 114, 035304 – Published 16 September, 2026
DOI: https://doi.org/10.1103/ch8n-wrv6
Abstract
In conventional generative adversarial networks (GANs), the discriminator typically employs a sigmoid activation to map features to probabilities. However, this activation suffers from the vanishing-gradient problem, which can lead to training instability and mode collapse. This work reformulates the final-layer activation of the discriminator as cumulative distribution functions (CDFs) of random variables, thereby introducing a learnable noise-scale parameter. We show analytically that, for CDF families with bounded reversed hazard rates in the negative tail, a noise-scale value smaller than unity scales the gradient of both the saturating and nonsaturating generator losses by a factor proportional to its reciprocal. Rather than fully resolving the vanishing-gradient problem, this mechanism provides more informative gradient signals to the generator when the discriminator is close to its optimum. Experiments on a two-dimensional Gaussian mixture and on MNIST show that the proposed CDF-based activation effectively improves mode coverage and training stability. Meanwhile, experiments on CIFAR-10 using a deep convolutional GAN (DCGAN) indicate that, when gradient flow is already stabilized by architectural features, the choice of final-layer activation plays a more limited role. An ablation study on MNIST further distinguishes the benefits of the proposed method from that of conventional noise regularization techniques. These results demonstrate that the proposed CDF-based activation also contributes a meaningful scheme for stabilizing adversarial training.
Physics Subject Headings (PhySH)
Article Text
References (50)
- C. M. Bishop, Training with noise is equivalent to Tikhonov regularization, Neural Comput. 7, 108 (1995).
- N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, Dropout: A simple way to prevent neural networks from overfitting, J. Mach. Learn. Res. 15, 1929 (2014).
- M. Ferianc, O. Bohdal, T. Hospedales, and M. Rodrigues, Impact of noise on calibration and generalisation of neural networks, in The Fortieth International Conference on Machine Learning, Workshop on Spurious Correlations, Invariance and Stability, Honolulu, Hawaii, edited by A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (ICML, 2023), pp. 1–7.
- S. Bai, F. Duan, F. Chapeau-Blondeau, and D. Abbott, Generalization of stochastic-resonance-based threshold networks with Tikhonov regularization, Phys. Rev. E 106, L012101 (2022).
- R. Benzi, A. Sutera, and A. Vulpiani, The mechanism of stochastic resonance, J. Phys. A: Math. Gen. 14, L453 (1981).
- N. G. Stocks, Suprathreshold stochastic resonance in multilevel threshold systems, Phys. Rev. Lett. 84, 2310 (2000).
- Z. Liao, Z. Shi, M. S. Sarker, and H. Tabata, Robust QRS detection based on simulated degenerate optical parametric oscillator-assisted neural network, Heliyon 10, e28903 (2024).
- F. Duan, F. Chapeau-Blondeau, and D. Abbott, Optimized injection of noise in activation functions to improve generalization of neural networks, Chaos Solitons Fractals 178, 114363 (2024).
- B. Kosko, K. Audhkhasi, and O. Osoba, Noise can speed backpropagation learning and deep bidirectional pretraining, Neural Netw. 129, 359 (2020).
- S. Ikemoto, F. Dallalibera, and K. Hosoda, Noise-modulated neural networks as an application of stochastic resonance, Neurocomputing 277, 29 (2018).
- W. McCulloch and W. Pitts, A logical calculus of the ideas immanent in nervous activity, Bull. Math. Biophys. 5, 115 (1943).
- Z. Xu, S. Lu, Y. Kang, and J. Jiang, Intelligent fault classification exploration inspired by suprathreshold stochastic resonance, IEEE Trans. Instrum. Meas. 74, 1 (2025).
- V. Nair and G. E. Hinton, Rectified linear units improve restricted Boltzmann machines, in Proceedings of the 27th International Conference on Machine Learning (ICML), edited by J. Fürnkranz and T. Joachims (Omni Press, Haifa, Israel, 2010), pp. 807–814.
- D. Hendrycks and K. Gimpel, Gaussian error linear units (GELUs), arXiv:1606.08415.
- Y. Ren, F. Duan, F. Chapeau-Blondeau, and D. Abbott, Self-gating stochastic-resonance-based autoencoder for unsupervised learning, Phys. Rev. E 110, 014107 (2024).
- I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, Generative adversarial nets, in Proceedings of the 28th International Conference on Neural Information Processing Systems, edited by Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger (MIT Press, Cambridge, MA, 2014), pp. 2672–2680.
- M. Arjovsky and L. Bottou, Towards principled methods for training generative adversarial networks, arXiv:1701.04862.
- T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, Improved techniques for training GANs, in Proceedings of the 30th International Conference on Neural Information Processing Systems, Barcelona, Spain, edited by D. D. Lee, U. von Luxburg, R. Garnett, M. Sugiyama, and I. Guyon (Curran Associates Inc., Red Hook, NY, 2016), pp. 2234–2242.
- H. Thanh-Tung and T. Tran, Catastrophic forgetting and mode collapse in GANs, in Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN) (IEEE, Glasgow, 2020), pp. 1–10.
- D. Saxena and J. Cao, Generative adversarial networks (GANs): Challenges, solutions, and future directions, ACM Comput. Surv. 54, 1 (2022).
- M. Wiatrak, S. V. Albrecht, and A. Nystrom, Stabilizing generative adversarial networks: A survey, arXiv:1910.00927.
- L. Mescheder, A. Geiger, and S. Nowozin, Which training methods for GANs do actually converge? in Proceedings of the 35th International Conference on Machine Learning, edited by J. Dy and A. Krause, Proceedings of Machine Learning Research, Vol. 80 (PMLR, Stockholm, 2018), pp. 3481–3490.
- T. Karras, S. Laine, and T. Aila, A style-based generator architecture for generative adversarial networks, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), edited by L. Davis, P. Torr, and S. Zhu (IEEE/CVF, Long Beach, CA, 2019), pp. 4401–4410.
- R. Feng, D. Zhao, and Z. Zha, Understanding noise injection in GANs, in Proceedings of the 38th International Conference on Machine Learning, edited by M. Meila and T. Zhang, Proceedings of Machine Learning Research, Vol. 139 (PMLR, Cambridge, MA, 2021), pp. 3284–3293.
- S. H. Hong and J. W. Jeong, Dynamic noise injection for facial expression recognition in-the-wild, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), edited by J. Hoffman and J. Sivic (IEEE/CVF, Vancouver, Canada, 2023), pp. 5709–5715.
- M. Arjovsky, S. Chintala, and L. Bottou, Wasserstein generative adversarial networks, in Proceedings of the 34th International Conference on Machine Learning (ICML), edited by D. Precup and Y. W. Teh, Proceedings of Machine Learning Research, Vol. 70 (PMLR, Sydney, 2017), pp. 214–223.
- I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. Courville, Improved training of Wasserstein GANs, in Advances in Neural Information Processing Systems 30 (NeurIPS), edited by I. Guyon, U. von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett (Curran Associates, Long Beach, CA, 2017), pp. 5767–5777.
- N. Kodali, J. Abernethy, J. Hays, and Z. Kira, On convergence and stability of GANs, arXiv:1705.07215.
- T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, Spectral normalization for generative adversarial networks, in 6th International Conference on Learning Representations (ICLR), edited by I. Murray, M. Ranzato, and O. Vinyals (OpenReview.net, Vancouver, Canada, 2018).
- G. Hinton, O. Vinyals, and J. Dean, Distilling the knowledge in a neural network, arXiv:1503.02531.
- C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, On calibration of modern neural networks, in International Conference on Machine Learning, edited by D. Precup and Y. W. Teh (PMLR, Sydney, 2017), Vol. 70, pp. 1321–1330.
- J. Zhang, Y. Dai, X. Yu, M. Harandi, N. Barnes, and R. Hartley, Uncertainty-aware deep calibrated salient object detection, arXiv:2012.06020.
- N. Papernot, A. Thakurta, S. Song, S. Chien, and Ú. Erlingsson, Tempered sigmoid activations for deep learning with differential privacy, in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI Press, Palo Alto, CA, 2021), Vol. 35, pp. 9047–9055.
- D. Page, K. Lin, and J. E. Wulsin, Do better backbones lead to better architectures? in Proceedings of the European Conference on Computer Vision (ECCV 2020), edited by A. Vedaldi, H. Bischof, T. Brox, and J. M. Frahm, Lecture Notes in Computer Science Vol. 12346 (Springer, Glasgow, 2020), pp. 577–592.
- H. Asatryan, H. Gottschalk, M. Lippert, and M. Rottmann, A convenient infinite dimensional framework for generative adversarial learning, Electron. J. Stat. 17, 391 (2023).
- V. Nagarajan and J. Z. Kolter, Gradient descent GAN optimization is locally stable, in Advances in Neural Information Processing Systems 30 (NeurIPS), edited by I. Guyon, U. von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Curran Associates, Long Beach, CA, 2017), pp. 5585–5595.
- S. Arora, R. Ge, Y. Liang, T. Ma, and Y. Zhang, Generalization and equilibrium in generative adversarial nets (GANs), in Proceedings of the 34th International Conference on Machine Learning (ICML), edited by D. Precup and Y. W. Teh Proceedings of Machine Learning Research, Vol. 70 (PMLR, Sydney, 2017), pp. 224–232.
- S. Boyd and L. Vandenberghe, Convex Optimization (Cambridge University Press, Cambridge, 2004).
- W. Feller, An Introduction to Probability Theory and its Applications, 3rd ed. (John Wiley & Sons, New York, 1968), Vol. I.
- Á. Baricz, Mills' ratio: Monotonicity patterns and functional inequalities, J. Math. Anal. Appl. 340, 1362 (2008).
- D. P. Kingma and J. Ba, Adam: A method for stochastic optimization, in International Conference on Learning Representations (ICLR) 2015, Poster Track, edited by Y. Bengio and Y. LeCun (ICLR, San Diego, CA, 2015), pp. 1–13.
- F. Duan, Source code, 2026, https://github.com/resonance-dfb/learning-noise-injection.
- Y. LeCun, C. Cortes, and C. J. C. Burges, The MNIST database of handwritten digits (1998), https://yann.lecun.org/exdb/mnist/.
- M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, GANs trained by a two time-scale update rule converge to a local Nash equilibrium, in Advances in Neural Information Processing Systems 30 (NeurIPS 2017), edited by I. Guyon, U. von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett (Curran Associates, Long Beach, CA, 2017), pp. 6626–6637.
- S. Barratt and R. Sharma, A note on the inception score, arXiv:1801.01973.
- C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, Rethinking the inception architecture for computer vision, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, Las Vegas, NV, 2016), pp. 2818–2826.
- A. Radford, L. Metz, and S. Chintala, Unsupervised representation learning with deep convolutional generative adversarial networks, arXiv:1511.06434.
- A. Krizhevsky, Learning multiple layers of features from tiny images, Technical Report TR-2009, University of Toronto, 2009.
- D. P. Kingma and M. Welling, Auto-encoding variational Bayes, in 2nd International Conference on Learning Representations (ICLR 2014), edited by Y. Bengio and Y. LeCun (OpenReview.net, Banff, AB, Canada, 2014).
- J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli, Deep unsupervised learning using nonequilibrium thermodynamics, in Proceedings of the 32nd International Conference on Machine Learning, edited by F. Bach and D. Blei, Proceedings of Machine Learning Research, Vol. 37 (PMLR, Lille, France, 2015), pp. 2256–2265.