blt
4 hours ago
Every few years, a derivative-free neural network optimization algorithm gets some hype. I'd bet my life savings that none of them ever make an impact.
Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.
The gradient is useful. Instead of trying random directions and hoping that one of them is an improvement, it tells you where to go. The more parameters you have, the more useful it becomes.
A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization has been suspected for decades and recently proved (using AI, [2]). Neural net optimization is nonconvex, but not radically different.
IMO, a more promising direction is gradient-based optimizers specialized to the neural network structure, like Muon [3].
[1] https://arxiv.org/abs/2202.00817
syntacticsalt
3 hours ago
Even for non-smooth or discontinuous objectives, I'd still reach for methods that use gradient-like information over zeroth-order methods. For non-smooth objectives, Clarke-generalized subdifferentials have been pretty effective outside of ML, and have been used in automatic differentiation contexts at least 10-15 years ago. A carelessly quick literature search suggested conservative gradients, too.
For discontinuous objectives, I know there's been work on using envelope approximations, but the little I'm aware of in that work was in low-dimensional settings where the structure of the discontinuity was known explicitly. On the other extreme, lack of continuity comes up all the time in infinite-dimensional, PDE-constrained optimization, and some methods rely on tangent cones or various generalized notions of subdifferentiability (e.g., Mordukhovich, Bouligand) to demonstrate convergence. Admittedly, that work was somewhat outside my area of expertise, so I may be getting the details there slightly wrong, but the broad point stands that even in those settings, some directional information can be obtained and used profitably without resorting to zeroth-order methods.
vatsachak
4 hours ago
Random selection seems to favor post training