Skip to main navigation Skip to search Skip to main content

The Vanishing Empirical Variance in Randomly Initialized Deep ReLU Networks

Research output: Chapter in Book/Conference proceedings/Edited volumeConference contributionScientificpeer-review

3 Downloads (Pure)

Abstract

Neural networks are typically initialized such that the hidden pre-activations’ theoretical variance remains constant to avoid the vanishing and exploding gradient problem. This condition is necessary to train very deep networks, but numerous analyses show this to be insufficient. We explain this behavior by analyzing the empirical variance, which is more meaningful in the practical setting that deals with data sets of finite size. We demonstrate its discrepancy with the theoretical variance, which grows with depth. We study the output distribution of neural networks at initialization and find that its kurtosis grows to infinity with increasing depth, even if the theoretical variance stays constant. As a result, the empirical variance vanishes: its asymptotic distribution converges in probability to zero. Our analysis focuses on fully connected ReLU networks with He-initialization, but we hypothesize that many more random weight initialization methods suffer from vanishing or exploding empirical variance. We support this hypothesis experimentally and demonstrate the failure of state-of-the-art random initialization methods in very deep regimes.

Original languageEnglish
Title of host publicationMachine Learning and Knowledge Discovery in Databases. Research Track - European Conference, ECML PKDD 2025, Proceedings
EditorsRita P. Ribeiro, Bernhard Pfahringer, Nathalie Japkowicz, Pedro Larrañaga, Alípio M. Jorge, Carlos Soares, Pedro H. Abreu, João Gama
PublisherSpringer
Pages362-379
Number of pages18
ISBN (Print)9783032060778
DOIs
Publication statusPublished - 2026
EventEuropean Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2025 - Porto, Portugal
Duration: 15 Sept 202519 Sept 2025

Publication series

NameLecture Notes in Computer Science
Volume16016 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

ConferenceEuropean Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2025
Country/TerritoryPortugal
CityPorto
Period15/09/2519/09/25

Keywords

  • empirical variance
  • kurtosis
  • ReLU
  • vanishing gradient

Fingerprint

Dive into the research topics of 'The Vanishing Empirical Variance in Randomly Initialized Deep ReLU Networks'. Together they form a unique fingerprint.

Cite this