Optimization & Learning Theory

The mathematics under the hood: optimization methods, convergence guarantees, and the theory that explains when and why deep learning works. Explainers that take proofs seriously without losing the reader.

An Axiomatic Definition of Hierarchical Clustering: Why Hartigan Was Right All Along.

An Axiomatic Definition of Hierarchical Clustering: Why Hartigan Was Right All Along

Ery Arias-Castro and Elizabeth Coda from UC San Diego take an approach inspired by Lebesgue integration — starting from first principles — to show that three intuitive axioms about what a cluster ought to be uniquely recover Hartigan’s cluster tree,…

An Axiomatic Definition of Hierarchical Clustering: Why Hartigan Was Right All Along Read More »

Bayes Meets Bernstein: Why Meta-Learning Finally Gets Fast — A PAC-Bayes Breakthrough.

Bayes Meets Bernstein: Why Meta-Learning Finally Gets Fast — A PAC-Bayes Breakthrough

A research team from three leading institutions has proven something nobody expected: when learning a prior across many tasks using the Gibbs algorithm, fast convergence rates are always guaranteed at the meta level — even when individual tasks refuse to…

Bayes Meets Bernstein: Why Meta-Learning Finally Gets Fast — A PAC-Bayes Breakthrough Read More »

Escaping Saddle Points in Bilevel Optimization — A Breakthrough in Machine Learning Theory.

Escaping Saddle Points in Bilevel Optimization — A Breakthrough in Machine Learning Theory

A research team from three major U.S. universities has proven for the first time that bilevel optimization algorithms can reliably escape saddle points and find true local minima — solving a problem that has quietly undermined meta-learning, hyperparameter tuning, and…

Escaping Saddle Points in Bilevel Optimization — A Breakthrough in Machine Learning Theory Read More »

TTGDA: Two-Timescale Gradient Descent Ascent for Nonconvex Minimax Optimization.

TTGDA: Two-Timescale Gradient Descent Ascent for Nonconvex Minimax Optimization

Tianyi Lin, Chi Jin, and Michael I. Jordan from Columbia, Princeton, and UC Berkeley deliver the first rigorous nonasymptotic analysis of two-timescale gradient descent ascent — the algorithm that has powered GAN training in practice for years — establishing tight…

TTGDA: Two-Timescale Gradient Descent Ascent for Nonconvex Minimax Optimization Read More »

How to Tune a Robust Regression Model Without Knowing the Noise: Adaptive Error Estimation for Unregularized M-Estimators.

How to Tune a Robust Regression Model Without Knowing the Noise: Adaptive Error Estimation for Unregularized M-Estimators

Pierre C. Bellec (Rutgers) and Takuya Koriyama (Chicago Booth) prove that a simple, fully observable formula consistently estimates the prediction error of any Huber-class M-estimator in high dimensions — and that minimizing it over a grid of scale parameters automatically…

How to Tune a Robust Regression Model Without Knowing the Noise: Adaptive Error Estimation for Unregularized M-Estimators Read More »

Why Your AI Says It's Confident When It Shouldn't Be — And How MaxWEnt Fixes It.

Why Your AI Says It’s Confident When It Shouldn’t Be — And How MaxWEnt Fixes It

Researchers from Michelin and the Centre Borelli at ENS Paris-Saclay have built a smarter way to make neural networks say “I don’t know” — by maximizing the diversity of possible network weights rather than relying on standard regularization that quietly…

Why Your AI Says It’s Confident When It Shouldn’t Be — And How MaxWEnt Fixes It Read More »