Chapter Four · failure evidence
What Gradient Descent & Backprop got wrong, from 51 dissertations
Across these doctoral research records, gradient descent and backpropagation frequently struggle with non-convex landscapes, ill-conditioned learning dynamics, and non-smooth or discrete objective functions. Practitioners also encounter severe obstacles when computing gradients through numerical approximations, saturated layers, or hierarchical bilevel formulations. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Gradient descent gets trapped in local minima across non-convex and multi-modal landscapes
Optimization frequently terminates prematurely in poor local optima when minimizing non-convex functions across robotics, parameter identification, and structural estimation. Without global initialization or basin-of-attraction knowledge, algorithms settle on suboptimal solutions or suffer diversity collapse.
Tried and failed
gradient-based optimization of parameterized geometry applied to electromagnetic inverse defect reconstruction. Outcome: did not converge. Reason: local minima trapping and breakdown of linear relationship for deep defect inversion
Iterative algorithms for electromagnetic NDE signal inversion · Iowa State
Considered and rejected
Considered and rejected: Rejected standard Backpropagation/Gradient Descent for LSTM training due to susceptibility to local minima entrapment.
Caching in VANETs for Social Networking · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected standard backpropagation/gradient-descent due to slow training speed, learning rate tuning issues, and risk of local minima.
Efficient and Accurate Neural Network Based Internal Combustion Engine Modeling and Prediction · Scholarship at UWindsor Institutional Repository
Considered and rejected
Considered and rejected: Rejected standard gradient descent and Newton's method due to getting trapped in local minima.
Singularity distance computations of parallel manipulators of Stewart-Gough type · DSpace-CRIS at TU Wien
Tried and failed
gradient descent pose optimization without initialization applied to slice-to-volume medical image registration. Outcome: did not converge. Reason: optimization gets trapped in local minima under severe motion without a good initial pose estimate
A Robust and Efficient Framework for Slice-to-Volume Reconstruction: Application to Fetal MRI · MIT
Tried and failed
gradient descent on relative 3D rotation applied to relative orientation estimation. Outcome: worse than baseline. Reason: optimization directly over relative orientation frequently gets trapped in local optima
Image-based Pose Estimation for Previously-Unseen Objects · EPFL
Tried and failed
direct target-space loss functions for projection models applied to density-density response function prediction. Outcome: worse than baseline. Reason: gradient descent got trapped in poor local minima compared to direct coefficient regression
Machine Learning static RPA response properties for accelerating GW calculations · Imperial
Tried and failed
gradient descent with zero initialization applied to 6-DOF pose estimation from shadow corners. Outcome: did not converge. Reason: gets trapped in poor local minima
Tried and failed
greedy descent local search metaheuristic applied to dynamic orienteering problem. Outcome: worse than baseline. Reason: rapidly converged to suboptimal local extrema yielding poor solution fitness
Practically optimal UAV mission planning under uncertainty · Imperial
Tried and failed
gradient-based parameter identification from random initialization applied to inverse parametric PDE problems. Outcome: did not converge. Reason: Non-convex optimization landscape trapped gradient descent in local minima without basin-of-attraction knowledge or global pre-initialization.
Error Assessment for Finite Elements/Neural Networks Methods Applied to Parametric PDEs · EPFL
Lost to a baseline
Gradient descent in parameter space was less accurate, showed higher run-to-run variance, and converged slower (e.g., 25.6-40.9 s vs 8.77-10.9 s) than Particle Swarm Optimization (PSO) when solving the 4D inverse parameter identification problem.
Error Assessment for Finite Elements/Neural Networks Methods Applied to Parametric PDEs · EPFL
Considered and rejected
Considered and rejected: Rejected gradient descent for human-in-the-loop multi-parameter optimization because the required perturbation points scale linearly with dimension and it risks getting trapped in local minima.
Soft Exosuits for Improved Walking Efficiency and Community Based Post-Stroke Gait Rehabilitation · Harvard
Considered and rejected
Considered and rejected: Rejected direct search and gradient-based optimization algorithms because direct search lacks precision for multi-objective trade-offs and gradient methods are prone to trapping in local extrema
DIGITAL CODESIGN OF A POWER MODULE WITH INTEGRATED THERMAL MANAGEMENT · Georgia Tech
Considered and rejected
Considered and rejected: Rejected gradient descent optimization for trial wavefunction parameter optimization because it gets trapped in local minima; adopted dual annealing global optimization instead.
Electronic quantum fluids in graphene and quantum Hall systems · Oxford
Considered and rejected
Considered and rejected: Gradient-based optimization (steepest descent, conjugate gradient) for FQ solvent parameter fitting, rejected due to the non-convex parameter space with multiple local minima
Polarizable QM/MM Approaches for Molecular Excited States in Complex Environments · IRIS - SNS - prod
Considered and rejected
Considered and rejected: Rejected descent methods, evolutionary algorithms, and statistical sampling methods for global parameter optimization due to premature termination at local minima, non-convergence, or inefficiency in higher dimensions.
Efficient Biomolecular Computations Towards Applications in Drug Discovery · Virginia Tech
Considered and rejected
Considered and rejected: Rejected pure gradient-based adversarial optimization for failure generation because local optima cause diversity collapse into a single failure mode.
Breaking things so you don’t have to: risk assessment and failure prediction for cyber-physical AI · MIT
Considered and rejected
Considered and rejected: Gradient-Descent optimization algorithms rejected for JPEG quantization table optimization due to settling on local minima in multi-modal rate-distortion space (Evolutionary / Particle Swarm Optimization used instead).
Digital CMOS ISFET architectures and algorithmic methods for point-of-care diagnostics · Imperial
Improper step sizes and learning rate dynamics cause oscillation, divergence, or drift
Gradient updates exhibit severe oscillations or diverge when learning rates are fixed, excessively large, or improperly scaled across uneven parameter magnitudes. Vanishing step sizes prevent adaptation to drift, while large steps in edge-of-stability regimes induce harmful coordinate drift.
Tried and failed
gradient descent without non-linear activations applied to linear neural network training. Outcome: did not converge. Reason: linear networks lack bounded activations that stabilize optimization in large learning rate regimes
Lost to a baseline
Gradient descent on quadratic forms lagged behind Conjugate Gradient due to iterate oscillation between subspaces and poor convergence.
MICROWAVE IMAGING FOR WALK-WHILE-SCAN SECURITY SCREENING · DukeSpace
Considered and rejected
Considered and rejected: Rejected using non-vanishing learning rates (ηt -> 0) in stochastic gradient descent because a vanishing step-size prevents the model from adapting to non-stationary drift.
Adaptive estimation and change detection of correlation and quantiles for evolving data streams · Imperial
Tried and failed
gradient descent with large learning rates applied to optimization of functions with poor regularity. Reason: insufficient objective regularity prevents edge of stability, causing premature one-sided stability bounds instead
Quantitative convergence analysis of dynamical processes in machine learning · Georgia Tech
Tried and failed
gradient descent with uniform parameter weighting applied to joint pose and shape optimization. Outcome: did not converge. Reason: gradient magnitudes for rotation and shape dominated translation, causing severe oscillations and optimization failure
Integrated 3D anatomical model for myocardial segmentation in cardiac CT imagery · Georgia Tech
Tried and failed
gradient descent with fixed step size applied to non-standard smooth non-convex optimization. Outcome: did not converge. Reason: fixed step sizes cause divergence or sublinear convergence under generalized smoothness bounds
Optimization Theory and Machine Learning Practice: Mind the Gap · MIT
Lost to a baseline
Gradient descent with large stepsizes in the Edge of Stability regime loses to small-stepsize gradient flow / SGD for sparse recovery, exhibiting high test error due to coordinate drift towards zero.
Deep Learning Theory Through the Lens of Diagonal Linear Networks · EPFL
Lost to a baseline
Hypergradient descent sometimes got stuck for too small initial learning rates (struggled on F-MNIST experiments) due to sensitivity of the hyper learning rate across architectures.
Probabilistic Linear Algebra for Stochastic Optimization · Publikationssystem UB Tuebingen
Lost to a baseline
Using a single constant learning rate (0.0005) for SGD across all regularization levels without validation grid search performed poorly compared to liblinear coordinate descent.
Gradient calculations break down on non-smooth, discontinuous, or discrete objectives
Standard gradient descent and automatic differentiation fail when applied to problems involving discrete logic, tight obstacle constraints, or non-smooth path norms. In these settings, derivative computations become invalid or subgradient calculations struggle to guide optimization progress.
Considered and rejected
Considered and rejected: Rejected standard gradient descent/backpropagation for symbolic reasoning due to inability to represent discrete logical decisions and out-of-distribution failures.
The Capabilities of Neural Systems Depend on a Hierarchically Structured World · DSpace at UTSWMED
Tried and failed
gradient descent with automatic differentiation applied to non-smooth path regularized optimization. Outcome: too slow. Reason: subgradient computation via automatic differentiation struggles with non-smooth path-norm objectives
Predicting in Uncertain Environments: Methods for Robust Machine Learning · EPFL
Tried and failed
sequential path stepping gradient descent optimization applied to constrained robotic trajectory planning. Outcome: did not converge. Reason: algorithm failed in environments with tight obstacles and long paths
Remote robotic manipulation task execution using affordance primitives · UT Austin
Considered and rejected
Considered and rejected: Rejected classical gradient-based and sub-gradient optimization methods for decentralized robot control optimization because the objective function is discontinuous, non-convex, and multimodal.
HOLISTIC SCENE PERCEPTION FOR COLLABORATIVE HUMAN-ROBOT TEAMS · Cornell
Considered and rejected
Considered and rejected: Standard gradient-descent optimization algorithms (e.g., Newton-Raphson) were rejected for discrete/mixed socio-demographic enrichment due to lack of smoothness and derivative breakdown.
Inverse discrete choice modelling: a framework for socio-demographic enrichment of big data · Imperial
Optimization degrades when relying on noisy, numerical, or inaccurate gradient estimates
Finite-difference approximations and numerical local sampling introduce significant noise and computational overhead that degrade optimization accuracy. Furthermore, optimizing against linearized approximations or inaccurate learned forward models causes real-world errors to diverge.
Tried and failed
influence functions and gradient-based data attribution applied to deep neural network predictions. Outcome: no signal. Reason: linearized gradient approximations fail to accurately predict leave-one-out counterfactual model behavior
Considered and rejected
Considered and rejected: Rejected Finite-Difference Gradient Descent (FDGD) for latent space optimization because it failed completely under biologically realistic stochastic evaluation noise.
Optimization of neural response in the primate dorsal visual pathway Optimierung der neuronalen Antwort im dorsalen visuellen Pfad von Primaten · open_UMR Marburg DSpace 10.0
Tried and failed
numerical gradient descent based local sampling applied to low-dimensional robot motion planning. Outcome: too slow. Reason: numerical gradient computation overhead outweighed the benefit of informed local sampling in low dimensions
INFORMED EXPLORATION ALGORITHMS FOR ROBOT MOTION PLANNING AND LEARNING · Georgia Tech
Tried and failed
gradient-based input synthesis with learned forward models applied to iterative plant control waveform generation. Outcome: unstable. Reason: excessive offline gradient optimization against inaccurate forward models caused real-world response errors to diverge
A novel Adaptive Filtering approach to Drive File Identification for Service Environment Replication · Virginia Tech
Considered and rejected
Considered and rejected: Rejected gradient-based local optimizers (fmincon) for the outer loop in parametric coil optimization because finite-difference parameter gradients were noisy and caused tuning failures; switched to global genetic algorithms.
Integral Equation-Based Inverse Scattering and Coil Optimization in Magnetic Resonance Imaging · MIT
Vanishing gradients and saturation prevent effective parameter updates
Network weights stall during optimization when gradients vanish across softmax layers, inactive neurons, or bottleneck architectures. Direct use of margin losses was similarly discarded because vanishing gradients obstructed progress compared to smooth approximations.
Tried and failed
Standard backpropagation on bottleneck autoencoder applied to Dimensionality reduction for data reconstruction. Outcome: did not converge. Reason: Standard gradient descent struggled with bottleneck convergence compared to second-order optimization methods
Using auto associative neural networks for replacing missing sensor data · Iowa State
Tried and failed
gradient-based optimization of model interpolation weights applied to graph neural network model merging. Outcome: worse than baseline. Reason: vanishing gradients in softmax prevented zeroing out poor-performing model ingredients
Improved and domain specific model soups · Iowa State
Tried and failed
gradient descent without smooth activation relaxation applied to locally constant neural networks. Outcome: did not converge. Reason: vanished gradients occur everywhere except boundaries due to zero gradients when neurons are inactive
Considered and rejected
Considered and rejected: Directly using the margin loss in stochastic gradient descent was rejected due to gradient vanishing, using cross-entropy approximations instead.
Multi-Domain Text Classification with Adversarial Training · Carleton University Institutional Repository
Standard descent fails to capture parameter dependencies in hierarchical and constrained formulations
Simultaneous gradient descent fails in bilevel optimization because it neglects the dependencies of lower-level parameters on upper-level variables. Direct one-stage formulations and overly conservative descent probability thresholds also stall optimization or underperform alternative embedding approaches.
Considered and rejected
Considered and rejected: Rejected one-stage direct optimization loss L(beta; y, A) via gradient descent due to severe empirical underperformance compared to spectral-embedding-based approaches
REGRESSION IN SINGLE AND MULTILAYER NETWORKS WITH UNKNOWN LATENT MANIFOLD STRUCTURE · JScholarship
Considered and rejected
Considered and rejected: Rejected simultaneous gradient descent for solving bilevel optimization problems due to failure to account for parameter dependencies on upper-level variables.
Tried and failed
conservative descent probability thresholds in local optimization applied to high-dimensional robotic continuous control. Outcome: did not converge. Reason: overly strict probability constraints caused the optimizer to terminate steps prematurely and stall progress
Sample Efficient Bayesian Optimization: From Local Search to Preference Learning · Penn
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop a theoretical framework explaining why neural networks learn linear functions before non-linear functions during stochastic gradient descent dynamics. Blocker: No specific mathematical formulation, assumptions, or analytical approach are provided
On Scaling Dynamics in Deep Learning · Harvard
Left open
Extend the MedImpute formulation and coordinate descent algorithm to incorporate support vector machines and decision tree objective functions. Blocker: None
Novel Machine Learning Algorithms for Personalized Medicine and Insurance · MIT
Left open
Develop an online dynamic MRI subspace tracking method using stochastic gradient descent for temporal subspace updates without mini-batch initialization. Blocker: None
Generalizable low-latency accelerated dynamic MRI · Iowa State
Left open
Extend the convergence rate analysis of gradient descent-ascent dynamics to broader non-convexity conditions such as weaker gradient dominance or local convexity. Blocker: None
Convergence Rates of Gradient Descent-ascent Dynamics under Computation Constraints in Solving Min-max Optimization · Virginia Tech
Left open
Apply harmless diversity descent to find diverse policies for robotics planning and reinforcement learning. Blocker: No specific robotics benchmark, environment, or evaluation metrics are provided
Adding prior into generative model · UT Austin
Left open
Develop a generalized step-size selection algorithm for time and control updates in adjoint gradient-based time-optimal control without example-specific heuristics. Blocker: Lack of specific target algorithmic approach or convergence criteria
The adjoint gradient method for time-optimal control of multibody systems Die adjungierte Gradientenmethode für zeitoptimale Steuerung von Mehrkörpersystemen · DSpace-CRIS at TU Wien
Left open
Develop low-dimensional dynamic programming coordinate descent algorithms over subgraphs for discrete planning problems. Blocker: Lacks detailed algorithmic specification and concrete benchmark targets for the discrete planning subgraph formulation
Efficient Learning and Inference for High-dimensional Lagrangian Systems · Penn
Left open
Extend the convergence analysis of gradient descent-ascent dynamics to stochastic and adversarial min-max settings with gradient noise or perturbations. Blocker: None
Convergence Rates of Gradient Descent-ascent Dynamics under Computation Constraints in Solving Min-max Optimization · Virginia Tech
Left open
Prove stepsize-independent lower bounds for gradient descent under directional smoothness conditions. Blocker: None
Smoothness and Adaptivity in Nonlinear Optimization for Machine Learning Applications · MIT
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.