stochastic gradient descent

How Biased Stochastic Gradients Still Generalize Well.

How Biased Stochastic Gradients Still Generalize Well

Analysis by the aitrendblend editorial team · Optimization Theory and Mathematical Foundations · 14 min read stochastic gradient descent algorithmic stability generalization bounds Zeroth-order SGD Clipped-SGD excess risk Two families of biased gradient methods, one built from function values only and one built from clipped gradients, now share a single stability proof. A graduate student […]

How Biased Stochastic Gradients Still Generalize Well Read More »

Why Batch Size Changes What Your Neural Network Learns

Why Batch Size Changes What Your Neural Network Learns

Analysis by the aitrendblend editorial team January 2025 Machine Learning Research Optimization Feature Learning GD (left) settles near a dense interior minimum; SGD with b=1 (right) escapes to a single datapoint on the boundary. From Ghosh et al., JMLR 2025. Pick any mainstream guide to training neural networks and you will read the same advice:

Why Batch Size Changes What Your Neural Network Learns Read More »