TRV-2026-0842Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-0842 version: 1 kind: certified reason: Certified into the record timestamp: 2026-08-21T06:03:08.715732Z status: published lens: g_space sector: science headline: Optimal Initialization Scale for Neural Networks With Locally Quadratic Loss Landscapes: An SGD Dynamics Perspective dek: Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence. As the quasi-steady-state distribution depends on… gain_title: Optimal Initialization Scale for Neural Networks With Locally Quadratic Loss Landscapes: An SGD Dynamics Perspective: The results show that, for the simple network model, if the variance of the initialization distribution satisfies our theoretical optimal condition, then the corresponding network achieves lower final training loss and higher test accuracy than the conventional He-normal initialization. problem_title: (none) trace_subject: (none) gain_reading: Optimal Initialization Scale for Neural Networks With Locally Quadratic Loss Landscapes: An SGD Dynamics Perspective: The results show that, for the simple network model, if the variance of the initialization distribution satisfies our theoretical optimal condition, then the corresponding network achieves lower final training loss and higher test accuracy than the conventional He-normal initialization. gain_evidence: (none) problem_reading: (none) problem_evidence: (none) quick_read: Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence. In addition, we experimentally confirm our theoretical results by using classical SGD to train shallow fully connected neural networks on the MNIST dataset and a two-layer CNN model on the CIFAR-10 dataset. Our work thus supplies a mathematically grounded indicator for choosing the initialization variance and clarifies its physical meaning in terms of the parameter dynamics in neural network models. limitation: tag: Evidence-backed gain key_points: Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. | Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence. | As the quasi-steady-state distribution depends on the expected cost function, the KL divergence eventually reveals the connection between the expected cost function and the initialization distribution. rundown: Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence. As the quasi-steady-state distribution depends on the expected cost function, the KL divergence eventually reveals the connection between the expected cost function and the initialization distribution. Under the assumptions of a locally quadratic neural-network loss landscape and a quasi-stationary SGD regime, we derive an explicit upper bound for the expected loss in terms of the initialization variance. sources: - peer_reviewed | IEEE Transactions on Neural Networks and Learning Systems | https://doi.org/10.1109/tnnls.2026.3721153 | 2026-08-17 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- 401a8af1288042afe7a1f3cc5df1adbfdecd1a9f8701c660d35faa1c2c9e9945
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0842 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace