Chapter Four · failure evidence
What Variational Inference got wrong, from 67 dissertations
The records evaluate variational inference methods across various statistical and machine learning domains, highlighting frequent failures in uncertainty quantification, optimization stability, and posterior expressiveness. Practitioners regularly find that variational approximations suffer from undercoverage, posterior collapse, or inferior performance compared to sampling baselines. These records come from PhD theses at 16 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Mean-field assumptions underestimate posterior variance and fail to capture parameter correlations
Factorized mean-field approximations systematically underestimate posterior uncertainty and yield artificially narrow credible intervals when latent variables are correlated. Authors frequently rejected or replaced mean-field formulations because assuming complete parameter independence distorts inference and reduces coverage.
Tried and failed
stochastic variational inference for point processes applied to hierarchical temporal point process models. Reason: systematically underestimates posterior variance leading to poor credible interval coverage
Tried and failed
mean-field stochastic variational inference applied to Bayesian models with correlated parameters. Reason: Factorized posterior distributions inherently underestimate uncertainty when latent parameters exhibit strong correlations
Tried and failed
mean-field stochastic variational inference applied to joint manifold and velocity estimation. Reason: mean-field approximation failed to capture strong posterior correlations between latent parameters, underestimating uncertainty
Tried and failed
mean-field variational inference applied to correlated sparse linear regression. Reason: mean-field assumption underestimates posterior variance under feature correlation, causing severe empirical undercoverage
Theoretical and methodological advances in Bayesian semiparametrics with variational inference · Imperial
Tried and failed
mean-field variational inference for neural ODEs applied to uncertainty quantification in dynamic systems. Outcome: did not generalise. Reason: mean-field approximation failed to capture posterior correlations needed for accurate forecasting compared to MCMC
Lost to a baseline
SVB marginal credible interval coverage for non-zero coefficients (0.770) was lower than MCMC (0.928) on survival analysis setting 1 with c=0.25 due to variational variance underestimation.
Variational bayes for high-dimensional linear models · Imperial
Considered and rejected
Considered and rejected: Variational Bayes ('meanfield') was rejected for final epidemic parameter inference due to ignoring posterior correlations and underestimating credible intervals relative to HMC.
MCMC methods: graph samplers, invariance tests and epidemic models · Imperial
Considered and rejected
Considered and rejected: Rejected mean-field variational inference (SVI) for velocity-learning because it yielded artificially narrow credible intervals and failed to capture joint parameter correlations between degradation rate and angular speed.
Considered and rejected
Considered and rejected: Rejected mean-field variational posterior factorization in ADVI for ChronoStrain in favor of a fully joint Gaussian approximation to preserve time-series coherence
Algorithms for Reconstructing Biological History from Genomic Data · MIT
Tried and failed
mean-field variational Bayes applied to PDE parameter estimation. Reason: Severely underestimated posterior variance and uncertainty due to assuming complete independence between parameters
Interpretable models for spatially dependent and heterogeneous phenomena · Imperial
Considered and rejected
Considered and rejected: Rejected mean-field SVI variational family for velocity learning because it assumes parameter independence, leading to overconfident posterior uncertainty bounds.
Considered and rejected
Considered and rejected: rejected standard mean-field variational inference in deep generative OSBM in favor of structured stochastic variational inference (SSVI)
Towards Efficient Continual Learning in Deep Neural Networks · DukeSpace
Variational autoencoders struggle with temporal dynamics and complex structured representations
Variational autoencoders frequently fail to capture multimodal trajectory distributions, continuous physical topologies, or accurate downstream predictions. In multiple settings, latent spaces caused blurry rollouts, overfitted to source structures, or lacked tractable likelihood evaluation.
Tried and failed
variational autoencoders for sequential modeling applied to multimodal agent state estimation. Outcome: did not generalise. Reason: failed to capture multimodal distributions under sparse and partial observations
Trajectory Modeling using Generative Approaches for Scheduling, Planning, and Multi-Agent Systems · Georgia Tech
Tried and failed
continuous conditional variational autoencoder applied to recursive trajectory and motion generation. Outcome: did not generalise. Reason: Continuous latent space caused blurry generations and failed to follow conditioning velocity commands over recursive rollouts
Generative Latent Motion Planning and Reinforcement Learning for Legged Locomotion · MIT
Tried and failed
variational auto-encoders for counterfactual generation applied to anomaly repair and explanation. Outcome: worse than baseline. Reason: struggled to generate high-quality repairs compared to diffusion models
Reliable Anomaly Detection with Explanation and Feedback · Penn
Tried and failed
L1 or L2 regularization on posterior parameters applied to variational autoencoder latent representations. Outcome: worse than baseline. Reason: Failed to increase representation sparsity compared to vanilla variational autoencoder baseline.
Injecting Inductive Biases into Distributed Representations of Text · Cambridge
Tried and failed
deep convolutional variational autoencoder applied to multivariate time series feature extraction. Outcome: did not converge. Reason: gradient backpropagation degradation in overly deep architectures led to significantly higher validation loss
Investigating the Brain States Behind FMRI Temporal Dynamics Using Frame-based Analysis Methods and Variational Autoencoders · Georgia Tech
Tried and failed
variational autoencoders for continuous dynamical systems applied to phase space trajectory modeling. Reason: failed to capture global topology and caused discontinuities across consecutive time steps
Conservation laws as inductive biases · Imperial
Tried and failed
MLP-based variational autoencoder applied to cross-dataset structural anomaly detection. Outcome: did not generalise. Reason: model overfitted to source structure distributions and failed on unseen structural data
Machine learning tools for identifying structural artifacts in data · Imperial
Lost to a baseline
Variational Autoencoders (VAEs) lose to standard Autoencoders (AEs) when training SNR and testing SNR perfectly match due to gaps in the latent space
Model and data driven approaches to wireless image transmission · Imperial
Considered and rejected
Considered and rejected: Rejected standard Variational Autoencoder (VAE) architecture for turbofan RUL estimation because decoders fail on temporal state projections; replaced decoder with a regressor network (RVE).
Uncertainty quantification of faults in rotating machines · Texas Tech
Considered and rejected
Considered and rejected: Rejected Variational Autoencoders (VAEs) for conditional human mesh recovery because they do not permit direct and tractable likelihood evaluation for downstream tasks
Tried and failed
surrogate modeling on variational autoencoder latent space applied to property prediction from spatial representations. Outcome: worse than baseline. Reason: nonlinear latent embedding degraded forward predictions and increased uncertainty compared to linear dimensionality reduction
Neural Inverse Microstructure Design with Bayesian Scale-Bridging · Georgia Tech
Tried and failed
variational autoencoder on covariance-filtered geometric features applied to single-cell chromatin dispersion clustering. Outcome: worse than baseline. Reason: removing correlated features discarded informative signals needed to separate characteristics beyond simple cluster counts
Stochastic optimization encounters numerical instability, divergence, and local minima
Variational inference training frequently suffers from gradient instabilities, divergence at scale, or noisy loss landscapes caused by stochastic sampling. Without delicate tuning, optimization can become trapped in local minima or trigger unstable feedback loops that destabilize inference.
Tried and failed
variational Bayes LDA topic modelling applied to sparse multi-decade longitudinal text corpora. Outcome: unstable. Reason: computationally unstable and unrepeatable compared to Gibbs sampling on sparse corpora
Tried and failed
variational mutual information estimators applied to multimodal cross-modal representation learning. Outcome: unstable. Reason: estimators underestimate due to batch size limits or diverge in high mutual information regimes
Multimodal Representation Learning for Medical Image Analysis · MIT
Tried and failed
joint end-to-end backpropagation through amortized variational inference applied to variational dropout in deep neural networks. Outcome: unstable. Reason: gradients from the inference network backpropagating into the generative decoder adversely affected training stability
Deep neural networks with contextual probabilistic units · UT Austin
Considered and rejected
Considered and rejected: Standard Variational AutoEncoder (VAE) with KL-divergence regularization on a single latent variable was rejected for inverse continuous localization because stochastic sampling caused noisy loss landscapes.
Condition monitoring for dry cask storage using helical guided ultrsonic waves · UT Austin
Considered and rejected
Considered and rejected: Rejected standard variational mutual information approximation for attention control because estimating posterior p(lt|st) without ground-truth state created unstable feedback loops and mode collapse.
Task Generalized MDPs for Multi-Task Reinforcement Learning · Georgia Tech
Tried and failed
annealing smoothing variance during variational inference applied to latent space inverse problem solving. Reason: sample quality failed to improve at very small smoothing variance values
Applications of deep generative models : inverse problems, compression, and beyond · UT Austin
Tried and failed
Stein variational gradient descent with large particle count applied to Bayesian inference for chaotic dynamical systems. Outcome: did not converge. Reason: Excessive computation time and gradient instability under strict convergence tolerances
Tried and failed
stochastic variational inference with normalizing flows applied to sequential surrogate-based parameter estimation. Outcome: unstable. Reason: lack of temperature annealing caused mode collapse, overconfident posteriors, and optimization divergence
Neural network based surrogates for scalable Bayesian inference on a complex malaria model · Imperial
Tried and failed
stochastic variational inference applied to high-dimensional parameter estimation. Outcome: unstable. Reason: optimization instability increases significantly as the number of simultaneous input parameters scales up
Tried and failed
variational inference with mean-field Gaussian prior applied to full-waveform inversion with poor initialisation. Outcome: did not converge. Reason: Severe non-convexity and local minima when starting from a flat homogeneous prior
Variational models underperform sampling methods and simpler baselines on convergence and predictive metrics
Variational methods often yield lower topic coherence, poor fitting, or slow convergence relative to Markov Chain Monte Carlo and standard baselines. Several studies rejected variational estimators because sampling produced more dependable uncertainty quantification and avoided aggregation bias.
Tried and failed
neural topic modeling using variational autoencoders applied to long academic documents. Outcome: worse than baseline. Reason: consistently produced negative normalized pointwise mutual information and low topic coherence scores
Topic Modeling for Heterogeneous Digital Libraries: Tailored Approaches Using Large Language Models · Virginia Tech
Tried and failed
variational inference with laplace approximation applied to bayesian hierarchical spatiotemporal models. Outcome: did not converge. Reason: failed to converge to sensible parameter values without heavy customisation, requiring MCMC instead
Spatiotemporal modelling of all-cause and cause-specific mortality in England · Imperial
Tried and failed
non-amortized stochastic variational inference applied to latent variable sequence modeling. Outcome: worse than baseline. Reason: insufficient local parameter updates during optimization compared to amortized inference networks
Structure Modeling for Language Models · Harvard
Lost to a baseline
Mean field variational Bayes underperformed the joint Kronecker variational model under misspecified low-rank perturbations, requiring 5946 iterations for r=1 versus 303 for the joint model.
Geometric Methods in MCMC and Variational Bayes for Multiway Data · Cornell
Lost to a baseline
Markov Chain Monte Carlo (MCMC) yielded models that better fit observed well data compared to variational inference stochastic gradient optimiser (Wingate et al., 2016).
Simulation of geothermal reservoirs with data assimilation and reduced order modelling · Imperial
Considered and rejected
Considered and rejected: Variational inference / expectation propagation for fitting Bayesian state-space models, rejected due to unsuccessful convergence/fitting compared to MCMC (NUTS)
Towards a data-driven personalised management of Atopic Dermatitis severity · Imperial
Considered and rejected
Considered and rejected: Rejected frequentist linear mixed models (e.g., lme4) and variational inference due to intractable or unreliable uncertainty estimation and p-value derivation in high-dimensional hierarchies.
Studying the tissue-specificity of cancer driver genes through KRAS and genetic dependency screens. · Harvard
Considered and rejected
Considered and rejected: Rejected Gensim Variational Bayes LDA implementation in favor of Mallet Gibbs sampling due to higher topic overlap and lower coherence.
The Satisfaction Levels of Services Provided by Transportation Network Companies (TNCs): A Semantic and Spatial Examination Based on Social Media Data · DSpace at SUNY Buffalo
Considered and rejected
Considered and rejected: Rejected variational inference in favor of Markov Chain Monte Carlo (MCMC) due to aggregation and risk of propagating bias.
Tried and failed
variational free energy for hidden Markov model selection applied to time-series state inference. Reason: models with equivalent free energy yielded divergent posterior state dynamics and inconsistent downstream interpretations
Multiscale neural dynamics and task-dependent states in human MEG via time-delay embedding · EPFL
Unimodal and symmetric parametric variational families fail on complex posteriors
Approximating posteriors with Gaussian or simplified distributions proves too restrictive for multimodal, discrete, or asymmetric target spaces. These crude variational families oversmooth derivatives, underfit complex architectures, and fail to capture diverse posterior modes.
Tried and failed
weight-space mean-field variational inference applied to transformer neural networks. Outcome: worse than baseline. Reason: the variational approximation severely underfits across predictive benchmarks
Probabilistic learning and generation in deep sequence models · Imperial
Considered and rejected
Considered and rejected: MC dropout variational Bayes for NN-GLS uncertainty quantification was rejected because binary variational approximation of continuous weight posteriors is too crude to yield valid interval estimates.
SPATIAL METHODS WITH GEOPHYSICAL AND GENOMIC APPLICATIONS · JScholarship
Considered and rejected
Considered and rejected: Point-wise variational inference / inducing point approximations due to oversmoothing and loss of uncertainty quantification
Modernizing Latent Gaussian Process Inference for Non-Gaussian Responses · Virginia Tech
Considered and rejected
Considered and rejected: Rejected Variational Bayesian Inference (VBI) and Laplace approximation due to reliance on unimodal/Gaussian parametric approximations that underestimate uncertainty in non-linear, multimodal ODE posteriors.
Bayesian Machine Learning for Gender-Stratified Mechanistic Epidemiological Models Using Public Health Data. · Carleton University Institutional Repository
Tried and failed
mean-field variational inference with diagonal covariance applied to spatial transformation parameter estimation. Outcome: worse than baseline. Reason: diagonal covariance variational posterior was too restrictive to capture dependencies for accurate registration
Considered and rejected
Considered and rejected: Variational inference / simplified posterior distributions for BNN PDE discovery, because it yields less informative derivative uncertainty quantification despite computational savings
Considered and rejected
Considered and rejected: Decided against Variational Inference for the scale stability model because VI approximates posteriors symmetrically with Gaussians, which is inappropriate for psychological space.
Machine Learning for Psychophysical Scaling with Ordinal Comparisons · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Rejected continuous variational approximations (e.g., Gaussian families) for reference priors because true reference priors are discrete and variational approximations fail to discover diverse models
Variational autoencoders suffer from posterior collapse and excessive regularization
Latent variables in variational autoencoders often collapse completely, ignoring latent codes or reducing mixture weights to a single component. Excessive regularization heavily penalizes the primary reconstruction objective, leading to blurry samples and distorted representations.
Tried and failed
autoregressive latent covariance structure applied to variational autoencoders. Reason: did not prevent posterior collapse when it occurred in standard VAE baseline
Tried and failed
information bottleneck and variational representation learning applied to out-of-distribution visual policy learning. Outcome: did not generalise. Reason: Representations collapsed into action-only shortcuts or encoded high-variance visual distractors.
From Pixels to Partners: A Hierarchical Approach to Adaptive Human-Robot Collaboration · Virginia Tech
Tried and failed
variational autoencoder with decoder word dropout applied to text generation and reconstruction. Outcome: no signal. Reason: severe posterior collapse leading to poor reconstruction quality despite high dropout rates
Tried and failed
variational autoencoders with pixel-wise reconstruction loss applied to image generation. Reason: Gaussian prior assumptions and pixel-wise loss cause blurry outputs and posterior collapse
Tried and failed
variational autoencoder for latent trajectory modeling applied to cell shape dynamics over time. Outcome: worse than baseline. Reason: prior regularisation warped latent dynamical trajectories compared to standard autoencoders
Computing Interpretable Representations of Cell Morphodynamics · Imperial
Tried and failed
Mixture-of-Gaussians approximate posterior with Dirichlet weights applied to variational autoencoder latent space. Reason: Optimization collapsed mixture weights to a single component, failing to improve posterior expressiveness
Human-controllable and structured deep generative models · Imperial
Tried and failed
high auxiliary regularization weight in variational autoencoder applied to image generation and representation learning. Outcome: worse than baseline. Reason: overly strong regularization heavily penalized the primary reconstruction objective, degrading sample generation quality
Capsule Networks: Framework and Application to Disentanglement for Generative Models · Virginia Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Extend LDML estimation and inference for estimand-dependent nuisances to non-smooth settings with multi-valued parameter regimes. Blocker: Lacks specific theoretical approach or algorithmic formulation for handling non-smooth multi-valued parameter regimes
Left open
Extend nested sampling Bayesian inference to simultaneously model and fit multiple cognitive tasks across experimental paradigms. Blocker: The unfinished work lacks specific target cognitive tasks, shared parameter structures, or concrete datasets to implement.
Left open
Develop and compare methods such as variational MAP estimation, Bayes factors, or reversible jump MCMC to summarize binary indicator matrix MCMC posterior samples. Blocker: None
Left open
Extend supervised LDA to time-varying parameters using dynamic modeling or rolling windows and compare variational inference with MCMC/Gibbs sampling. Blocker: None
Left open
Develop statistical inference frameworks to estimate cluster variance and confidence intervals for UMAP embedding outputs. Blocker: Lacks a concrete theoretical framework or formulation to guide the development of statistical inference on model-free embeddings
Ensemble Methods for Latent Structure Detection from Heterogeneous Genomic and Phenotypic Data · Harvard
Left open
Explore alternative sampling and normalization methods to optimize the Bayesian Inverse Reinforcement Learning inference process across different interaction settings. Blocker: Vague task specification with no defined alternatives or evaluation targets
Inferring the Human's Objective in Human Robot Interaction · Virginia Tech
Left open
Implement variational inference for the bespoke Bayesian hierarchical Gaussian mixture model to improve computational scalability over MCMC. Blocker: None
Clustering of cardiometabolic and renal risk factors · Imperial
Left open
Reduce inference computational cost during structured approximate cross-validation for very large graphical models. Blocker: No specific algorithmic approach or target architecture is defined
Faster and easier: cross-validation and model robustness checks · MIT
Left open
Develop robust statistical inference methods for arbitrarily weak and invalid instrumental variables with high-dimensional covariates. Blocker: Lacks a specific mathematical approach or formulation for handling arbitrary weakness alongside invalidity
Statistical Inference For High-Dimensional Linear Models · Penn
Left open
Combine optimization methods like Neural ODEs with Markov basis MCMC in a naive Bayes scheme for origin-destination inference. Blocker: High-level algorithmic concept lacks concrete architectural specification and defined target metrics
Table inference for combinatorial origin‐destination choices in agent‐based population synthesis · Cambridge
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.