Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks

Дата публикации: 17-08-2026 20:26:00


We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher--student setup. We focus on the extensive-width regime, where the teacher and student network widths scale proportionally with the input dimension, and the sample size grows quadratically. This scaling aims to describe overparameterized neural networks in which feature learning still plays a central role. In the high-dimensional limit, we derive a dynamical characterization of the gradient flow, in the spirit of dynamical mean-field theory (DMFT). Under $\ell_2$-regularization, we analyze these equations at long times and characterize the performance and spectral properties of the resulting estimator. This result provides a quantitative understanding of the effect of overparameterization on learning and generalization, and reveals a double descent phenomenon in the presence of label noise, where generalization improves beyond interpolation. In the small regularization limit, we obtain an exact expression for the perfect recovery threshold as a function of the network widths, providing a precise characterization of how overparameterization influences recovery.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks 010.9717-08-2026
2 Optimization and Generalization of Gradient Descent for Shallow ReLU Networks with Minimal Width 03.8417-08-2026
3 Statistical Learning Theory for Neural Operators 010.2117-08-2026
4 A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization 09.8217-08-2026
5 Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data 04.3617-08-2026
6 Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection 04.2317-08-2026
7 Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective 010.9417-08-2026
8 Reparameterized Complex-valued Neurons Can Efficiently Learn More than Real-valued Neurons via Gradient Descent 011.8617-08-2026
9 Error Analysis for Deep ReLU Feedforward Density-Ratio Estimation with Bregman Divergence 08.7817-08-2026
10 Gradient Span Algorithms Make Predictable Progress in High Dimension 06.3817-08-2026

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 8.7. Источник: jmlr.org.