Optimal Initialization Scale for Neural Networks With Locally Quadratic Loss Landscapes: An SGD Dynamics Perspective
Abstract: Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence. As the quasi-steady-state distribution depends on…
An artificial neural network control system for spacecraft attitude stabilization by Segura, Clement M.. Public domain
Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence.
In addition, we experimentally confirm our theoretical results by using classical SGD to train shallow fully connected neural networks on the MNIST dataset and a two-layer CNN model on the CIFAR-10 dataset. Our work thus supplies a mathematically grounded indicator for choosing the initialization variance and clarifies its physical meaning in terms of the parameter dynamics in neural network models.
- Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies.
- Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence.
- As the quasi-steady-state distribution depends on the expected cost function, the KL divergence eventually reveals the connection between the expected cost function and the initialization distribution.
Optimal Initialization Scale for Neural Networks With Locally Quadratic Loss Landscapes: An SGD Dynamics Perspective: The results show that, for the simple network model, if the variance of the initialization distribution satisfies our theoretical optimal condition, then the corresponding network achieves lower final training loss and higher test accuracy than the conventional He-normal initialization.
The rundown
As the quasi-steady-state distribution depends on the expected cost function, the KL divergence eventually reveals the connection between the expected cost function and the initialization distribution. Under the assumptions of a locally quadratic neural-network loss landscape and a quasi-stationary SGD regime, we derive an explicit upper bound for the expected loss in terms of the initialization variance.
Sources
- Peer-reviewedIEEE Transactions on Neural Networks and Learning Systems2026-08-17
How should this claim be treated?
ace
The debate