Chapter Four · failure evidence

What Autoencoders & Generative Modeling got wrong, from 106 dissertations

The records document widespread failure points across autoencoder and generative modeling architectures, highlighting recurring difficulties in optimization stability, representation quality, and benchmark competitiveness. Across diverse domains, these models frequently struggle with mode and posterior collapse, severe overfitting on source distributions, and underperformance relative to simpler baselines. These records come from PhD theses at 27 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Unsupervised compression and reconstruction fail to preserve task-relevant features

24 theses · 13 institutions

Standard autoencoders frequently optimize general reconstruction error at the expense of statistical interpretability or task-specific performance, sometimes treating critical anomalies as reconstructible noise. In addition, architectural bottlenecks and unconstrained representations often suffer from local optima, trivial identity mappings, or non-invertible feature spaces that undermine downstream utility.

Tried and failed

surrogate modeling on variational autoencoder latent space applied to property prediction from spatial representations. Outcome: worse than baseline. Reason: nonlinear latent embedding degraded forward predictions and increased uncertainty compared to linear dimensionality reduction

Neural Inverse Microstructure Design with Bayesian Scale-Bridging · Georgia Tech

Tried and failed

conditional variational autoencoder generating high-dimensional intermediate features applied to deep neural network feature activations. Outcome: did not converge. Reason: High feature dimensionality and non-linear classifier structure prevented effective training of small generative models

Enabling on-device domain adaptation of convolutional neural networks · Imperial

Tried and failed

Unimodal variational autoencoder applied to multimodal content representation learning. Outcome: worse than baseline. Reason: Missing cross-modal correlations reduces representation quality and reconstruction accuracy

THEORETICAL AND EMPIRICAL EXPLORATIONS OF INFLUENCER MARKETING · Penn

Tried and failed

variational autoencoder on covariance-filtered geometric features applied to single-cell chromatin dispersion clustering. Outcome: worse than baseline. Reason: removing correlated features discarded informative signals needed to separate characteristics beyond simple cluster counts

Categorizing the dispersion of open chromatin usage patterns of Treg related genes using an variational autoencoder based algorithm · Harvard

Considered and rejected

Considered and rejected: Rejected using Autoencoders for primary dimension identification because latent representations optimize reconstruction error rather than statistical interpretability

Towards soundscape fingerprinting: development, analysis and assessment of underlying acoustic dimensions to describe acoustic environments · Leibniz Universität Hannover Repository

Considered and rejected

Considered and rejected: Rejected an autoencoder architecture with symmetric 4-layer ConvLSTM encoder/decoder predicting autoregressively across multiple stochastic steps due to loss convergence failure and severe overfitting

Integrating Machine Learning Techniques for Streamlined Predictive Modeling in Cosmological Applications · Georgia Tech

Considered and rejected

Considered and rejected: Rejected standard Variational Autoencoder (VAE) architecture for turbofan RUL estimation because decoders fail on temporal state projections; replaced decoder with a regressor network (RVE).

Uncertainty quantification of faults in rotating machines · Texas Tech

Considered and rejected

Considered and rejected: Rejected non-convex dimensionality reduction (deep autoencoders) due to susceptibility to local optima, lack of robustness, and high trial-and-error training cost.

A Reduced Order Modeling Methodology for the Multidisciplinary Design Analysis of Hypersonic Aerial Systems · Georgia Tech

Considered and rejected

Considered and rejected: Rejected standard under-parameterized autoencoders with bottleneck layers because they cannot interpolate training examples or store them as exact fixed points.

Foundations of Machine Learning: Over-parameterization and Feature Learning · MIT

Considered and rejected

Considered and rejected: Rejected sigmoid activation functions because autoencoders failed to train effectively when handling fluctuating fields containing both positive and negative values.

Scientific machine learning for the analysis and reconstruction of turbulent flows · Imperial

Considered and rejected

Considered and rejected: Rejected training the autoencoder jointly from scratch with the decoder or using a single linear projection layer, as both lagged significantly behind pre-trained autoencoders in ROUGE metrics.

Efficient and Enhanced Text Summarization by Compressing and Data Augmentation for Transformers-Based Models · Scholarship at UWindsor Institutional Repository

Tried and failed

batch normalization applied to autoencoder anomaly detection. Reason: did not improve anomaly detection performance during model development

Developing AI Systems for Monitoring Heterogeneous Mental Health Disorders · Cornell

Tried and failed

denoising autoencoder confounder correction applied to detecting high-frequency biological outliers. Reason: autoencoder erroneously corrected true synthetic aberrations by treating them as technical noise or outliers

ON BIOINFORMATIC METHODS FOR THE DETECTION OF ALTERNATIVELY SPLICED VARIANTS FOR CLINICAL DIAGNOSTICS · Penn

Tried and failed

nested autoencoder architecture applied to time series reconstruction. Reason: matched standard autoencoder reconstruction without providing significant accuracy gains

Deep Time: Deep Learning Extensions to Time Series Factor Analysis with Applications to Uncertainty Quantification in Economic and Financial Modeling · Virginia Tech

Tried and failed

L2 regularization on autoencoder dimensionality reduction applied to unsupervised anomaly detection. Reason: higher regularization penalties caused underfitting and unpredictable performance degradation

On the Effectiveness of Dimensionality Reduction for Unsupervised Structural Health Monitoring Anomaly Detection · Virginia Tech

Tried and failed

autoencoder-based anomaly detection on textured backgrounds applied to surface anomaly detection. Outcome: no signal. Reason: The network acts as a general compressor, reconstructing textured surface variations indistinguishably from actual anomalies.

Detecting Anomalies and Obstacles in Road Scenes · EPFL

Tried and failed

autoencoder feature extraction in bidirectional GAN applied to microbiome data imputation. Outcome: worse than baseline. Reason: autoencoder failed to capture phylogenetic spatial relationships as effectively as convolutional layers

Deep Learning for Enhancing Human and Environmental Health · Virginia Tech

Considered and rejected

Considered and rejected: Rejected standard autoencoders (predicting neurons from themselves) to discard non-shared private variability.

High-dimensional neuronal activity from low-dimensional latent dynamics: a solvable model · EPFL

Considered and rejected

Considered and rejected: Rejected autoencoder latent-space partition for expert model selection because autoencoders learn generic input distributions rather than task-specific model fitness.

Continuous Learning for Lightweight Machine Learning Inference at the Edge · MIT

Considered and rejected

Considered and rejected: Rejected autoencoder-based bottlenecks for PVN internal state prediction because compression limited prediction accuracy and risked trivial identity transformations without bottlenecks.

Foundations of a New Learning Paradigm in AI Grounded in the Principles of Evolutionary Developmental Biology · EPFL

Considered and rejected

Considered and rejected: Rejected backpropagation/autoencoders for feature compression because autoencoders can reconstruct original sensitive data and require decoders.

An Efficient Multidimensional k-Anonymisation Strategy Using Self-Organising Maps · De Montfort Open Research Archive (DORA)

Considered and rejected

Considered and rejected: Rejected relying directly on raw noisy autoencoder reconstruction error for moving thresholds, adding an alternating binary convolutional filter for denoising.

A Hitchhikers Guide to Anomaly Detection: Machine Learning-based approaches to anomaly detection and Failure diagnosis in spacecraft avionics systems using low power devices & flight ready devices · Research Repository UCD

Considered and rejected

Considered and rejected: Rejected autoencoder neural network optimizers for LDSC design because the feature extraction process is non-invertible, preventing gradient backpropagation.

From Vulnerability to Resilience: Securing Signal Transmission Against Jamming and Spoofing Attacks · Virginia Tech

Considered and rejected

Considered and rejected: Rejected fully unconstrained deep autoencoders for unsupervised dictionary learning because untying weights broke recovery of generative physical filters.

Generative models for neural time series with structured domain priors · MIT

Complex generative models and autoencoders underperform simpler baselines

22 theses · 14 institutions

Variational and standard autoencoders regularly trail traditional machine learning methods such as decision trees, random forests, one-class support vector machines, and raw feature inputs across classification and anomaly detection tasks. In several image and signal domains, simpler supervised networks or matched filters consistently outperform deep generative formulations while avoiding high training overhead.

Tried and failed

variational autoencoder applied to tabular financial earnings metrics. Outcome: worse than baseline. Reason: standard autoencoders offered simpler training and direct reconstruction without probabilistic complexity

Machine learning in finance: from earnings data analysis to algorithmic trading · Imperial

Tried and failed

semi-supervised variational autoencoders applied to protein fitness prediction with combinatorial mutations. Outcome: worse than baseline. Reason: jointly trained models failed to outperform simpler supervised baselines like linear regression

Optimizing Protein Fitness and Function with Sparse Experimental Data · Harvard

Tried and failed

sequence-to-sequence variational autoencoder applied to sketch to latent space mapping. Outcome: worse than baseline. Reason: failed to show significant improvement over simpler vector KNN regression despite higher training complexity

Computational Gestural Making: A framework for exploring the creative potential of gestures, materials, and computational tools · MIT

Lost to a baseline

Vanilla autoencoders, adversarial autoencoders (1xAAE and 4xAAE), and PCA achieved lower clustering homogeneity scores (by 6% to 9%) compared to the variational autoencoder in UPSIDE.

Understanding Cell State Transitions in Development and Disease · ResearchWorks

Lost to a baseline

Autoencoder (F2=0.8156, precision=0.469) was beaten by OCSVM (F2=0.9914, precision=0.958).

Data Analytics and Machine Learning Applications in Fermentation Processes and Molecular Property Prediction · Virginia Tech

Lost to a baseline

single-layer linear autoencoder lost to VAE in field classification (scored 0.771 vs 0.992 on 3-field drive, and 0.3934 vs 0.829 on 5-field drive)

Non-equilibrium physics: from spin glasses to machine and neural learning · MIT

Lost to a baseline

CNN autoencoder feature clustering (0.36 accuracy, 0.22 NMI) was outperformed by ViT encoder feature clustering (0.78 accuracy, 0.61 NMI)

Deep Learning and Augmented Reality for 3D human-machine interaction · IRIS - POLITO - prod

Lost to a baseline

Variational Autoencoders (VAEs) lose to standard Autoencoders (AEs) when training SNR and testing SNR perfectly match due to gaps in the latent space

Model and data driven approaches to wireless image transmission · Imperial

Considered and rejected

Considered and rejected: Rejected Variational Autoencoders (VAEs) for high-resolution semantic image synthesis of climate impacts because they generated less realistic images than GANs.

Deep Learning Emulators for Accessible Climate Projections · MIT

Tried and failed

latent diffusion models applied to high-resolution structured grayscale medical images. Outcome: worse than baseline. Reason: struggled with fine structures, failing to match generative adversarial network fidelity

Controllable synthetic algorithms and evaluations for annotated clinical image synthesis · Imperial

Tried and failed

adversarial training without generative modeling applied to semi-supervised image segmentation. Outcome: worse than baseline. Reason: distorted the training signal compared to supervised and generative alternatives

Label-efficient medical image segmentation: the benefits and limitations of semi-supervised generative models · Imperial

Lost to a baseline

Autoencoder-based models (mDSC <= 0.08) lost to supervised U-Net (mDSC 0.50) and GMM (mDSC 0.17) on ATLAS-T1w lesion segmentation

Probabilistic and causal reasoning in deep learning for imaging · Imperial

Lost to a baseline

On public Cartographer dataset, Autoencoder Network (w/o skips) achieved SSIM of 0.904, marginally outperforming the thesis method's 0.903 (though thesis achieved higher PSNR 16.907 vs 16.459).

Integrating Perception, Prediction and Control for Adaptive Mobile Navigation · JScholarship

Lost to a baseline

1D CNN architectures (Kachuee et al. and Acharya et al.) achieved higher sensitivity on the minority Fusion (F) class under FGSM, BIM, and HSJ adversarial attacks compared to ECG-ATK-GAN.

Robust and Efficient AI-models for Medical Image Reconstruction, Segmentation, and Multimodal Knowledge Distillation · unevada

Lost to a baseline

Random Forest beat Autoencoders on pointwise anomaly detection accuracy (RF: 0.963 vs AE: 0.937) and F1-score (RF: 0.967 vs AE: 0.953)

An AI and data-driven approach to unwanted network traffic inspection · IRIS - POLITO - prod

Lost to a baseline

Autoencoder fault identification accuracy (0.9873) was beaten by Decision Tree (0.9898) and Neural Network (0.9897)

Fault diagnosis in aircraft fuel system components with machine learning algorithms · Cranfield

Lost to a baseline

Autoencoder and PCA dimensionality reduction were both beaten by raw feature representation (no dimensionality reduction) across all anomaly detection models (e.g., IF Top 100: 10.8 raw vs 6.6 autoencoder vs 0.2 PCA).

THE EFFECTIVENESS OF MACHINE LEARNING-BASED ANOMALY DETECTION ALGORITHMS APPLIED TO DEFENSE CONTRACT FINANCIAL DATA · Calhoun

Lost to a baseline

Matched filter alone slightly outperformed the U-Net autoencoder on pure QPSK + AWGN due to non-linear distortion introduced by the autoencoder

On the Use of Deep Learning Models for Interference Detection and Mitigation · Virginia Tech

Lost to a baseline

Attention autoencoder was outperformed by CNN autoencoder on non-sparse TEC reconstruction (R=0.9480 vs R=0.9604).

Employing Machine Learning Techniques to Increase the Quality of Ionospheric Modeling · Georgia Tech

Lost to a baseline

Autoencoder + MC-Dropout achieved 86.73 ± 6.02% balanced accuracy and 80.46 ± 12.50% sensitivity, losing to standalone Autoencoder (88.37 ± 3.96% balanced accuracy, 82.57 ± 9.97% sensitivity).

Self-supervised learning and uncertainty estimation for surgical margin detection with mass spectrometry · Queens University Institutional Repository

Considered and rejected

Considered and rejected: Rejected deep learning models (e.g., autoencoders) for unknown attack detection because they require large training datasets and performed worse than tree-based models.

Diagnosis and Mitigation of Evolving Threats for Sustainable Security · Research Repository UCD

Lost to a baseline

SDVAE and SOS-DVAE had higher reconstruction loss (0.086–0.090) compared to unconstrained generative VAE (0.078) on the TST dataset when removing task condition.

Relating Traits to Electrophysiology using Factor Models · DukeSpace

Adversarial training suffers from mode collapse and optimization instability

18 theses · 11 institutions

Generative adversarial networks repeatedly encounter severe mode collapse, vanishing gradients, and training divergence when balancing generator and discriminator objectives. These persistent instabilities often cause models to produce low-diversity outputs or collapse into noise, leading researchers to discard adversarial training in favor of collaborative or non-adversarial losses.

Considered and rejected

Considered and rejected: Decided against using GANs for generative spectral modeling because of training instability, choosing Variational Autoencoders (VAEs) instead.

Information Content and Analysis of X-ray Absorption Spectroscopy and X-ray Emission Spectroscopy Using Machine Learning · ResearchWorks

Tried and failed

generative adversarial network without variational autoencoder applied to 3D voxel shape generation. Outcome: unstable. Reason: mode collapse prevented learning the latent probability distribution, generating only a few shapes

Deep Learning Based Manufacturing Capability Modeling for Process Planning Automation · Georgia Tech

Tried and failed

multi-scale gradient input to discriminator applied to generative adversarial network image synthesis. Outcome: worse than baseline. Reason: broke the generator-discriminator balance, degrading output quality

Efficient Deep Learning Computing: From TinyML to LargeLM · MIT

Tried and failed

Classical generative adversarial network applied to decay histogram distribution modeling. Outcome: unstable. Reason: Mode collapse and vanishing gradients during training on histogram data

Development of the single-molecule tracking and fluorescence lifetime techniques for biophysical measurements and biomedical applications · UT Austin

Tried and failed

training generator ensemble without discriminator applied to generative adversarial network training. Outcome: did not converge. Reason: generator suffered severe mode collapse and converged to a single low-variance mode without adversarial feedback

Memory of Motion for Initializing Optimization in Robotics · EPFL

Tried and failed

factored latent codes for scene decomposition applied to generative image synthesis. Outcome: unstable. Reason: foreground latent code collapsed into background representation during training

Unsupervised learning of human movement from images · Imperial

Tried and failed

training GAN generator without discriminator applied to binary image generation. Outcome: did not converge. Reason: removing the adversarial feedback caused generated patterns to collapse into random noise

Algorithmic design of photonic structures with deep learning · Georgia Tech

Tried and failed

single-stage physics-constrained generative adversarial network applied to structural topology optimization design generation. Outcome: unstable. Reason: Mode collapse and severe constraint violation from combining complex physics loss with adversarial training in one stage.

Generative Design Using Deep Learning Methods for Functionality and Manufacturability · Georgia Tech

Tried and failed

differentially private generative adversarial networks applied to tabular synthetic data generation. Outcome: unstable. Reason: mode collapse caused failure to match one-way marginals across most variables

Causal Inference Under Privacy Constraints · MIT

Considered and rejected

Considered and rejected: Rejected Generative Adversarial Networks (GANs) and standard INNs/MDNs due to mode collapse, vanishing gradients, lower precision, and training instabilities on lower-dimensional sub-manifolds.

Towards Efficient and Robust Robot Planning · DukeSpace

Considered and rejected

Considered and rejected: Rejected generative networks (GANs/VAEs) for creating high-uncertainty adversarial active learning candidates due to high compute, instability, and noisy unrealistic medical artifacts, opting instead for gradient/attack-based perturbation.

Weakly-supervised Learning for Cost-Effective Medical Image Analysis Weakly-Supervised learning for Label-Effective Medical Image Analysis · Research Repository UCD

Considered and rejected

Considered and rejected: Rejected Generative Adversarial Networks (GANs) for structure generation due to training instability, mode collapse, and lack of explicit likelihood/interpretable latent space.

On-Demand Inverse Design of Phononic Metamaterials via Deep Learning · Leibniz Universität Hannover Repository

Considered and rejected

Considered and rejected: Rejected pure gradient-based adversarial optimization for failure generation because local optima cause diversity collapse into a single failure mode.

Breaking things so you don’t have to: risk assessment and failure prediction for cyber-physical AI · MIT

Considered and rejected

Considered and rejected: GAN adversarial discriminator training in favor of collaborative L2 loss among Noise2Noise generators to reduce complexity

Denoising Low-Dose CT Images using Multi-frame techniques · HARVEST

Considered and rejected

Considered and rejected: Rejected standard GAN adversarial training for video prediction, opting for pixel-wise MSE/BCE losses paired with transformational state architecture.

Learning, Moving, And Predicting With Global Motion Representations · Penn

Tried and failed

conditional recurrent neural network generative model applied to property-conditioned molecular sequence generation. Reason: mode collapse during generation leading to low uniqueness in generated outputs

Designing Macromolecules using Machine Learning and Simulations · MIT

Considered and rejected

Considered and rejected: Rejected generative adversarial networks (GANs) for inverse CTLE design due to training instability and inability to systematically model multi-modal parameter distributions.

High-speed Channel Analysis and Design using Polynomial Chaos Theory and Machine Learning · Georgia Tech

Considered and rejected

Considered and rejected: Rejected standard Conditional Generative Adversarial Networks (CGANs) for residential load scenario forecasting because training is unstable when conditioning on continuous vector-valued historical time series.

Learning and Optimization for Efficient and Optimal Operations in Sustainable Power Systems · ResearchWorks

Variational autoencoders experience posterior collapse and latent space distortion

15 theses · 10 institutions

Imposing prior regularization and Kullback-Leibler divergence terms frequently warps latent trajectories and triggers posterior collapse instead of learning organized representations. Standard Gaussian assumptions also struggle to represent non-Gaussian data statistics, leading to noisy loss landscapes, blurry outputs, and loose variational bounds.

Tried and failed

autoregressive latent covariance structure applied to variational autoencoders. Reason: did not prevent posterior collapse when it occurred in standard VAE baseline

Deep Time: Deep Learning Extensions to Time Series Factor Analysis with Applications to Uncertainty Quantification in Economic and Financial Modeling · Virginia Tech

Tried and failed

L1 or L2 regularization on posterior parameters applied to variational autoencoder latent representations. Outcome: worse than baseline. Reason: Failed to increase representation sparsity compared to vanilla variational autoencoder baseline.

Injecting Inductive Biases into Distributed Representations of Text · Cambridge

Tried and failed

variational autoencoder for sparse latent representation applied to text sentence embeddings. Outcome: unstable. Reason: struggled to achieve steady and consistent sparsity across hyperparameter configurations

Injecting Inductive Biases into Distributed Representations of Text · Cambridge

Tried and failed

variational autoencoder for latent trajectory modeling applied to cell shape dynamics over time. Outcome: worse than baseline. Reason: prior regularisation warped latent dynamical trajectories compared to standard autoencoders

Computing Interpretable Representations of Cell Morphodynamics · Imperial

Tried and failed

variational autoencoder without auxiliary predictor network applied to structured latent space performance prediction. Outcome: no signal. Reason: unsupervised reconstruction objective failed to organize the latent space with respect to the downstream target metric

Incorporating prior knowledge to efficiently design deep learning accelerators · UT Austin

Tried and failed

variational autoencoder with explicit mutual information maximization applied to image representation learning. Outcome: worse than baseline. Reason: leads to very high KL divergence, failing to match the marginalized posterior to the prior

Novel approaches for learning representations · UT Austin

Tried and failed

neural topic modeling using variational autoencoders applied to long academic documents. Outcome: worse than baseline. Reason: consistently produced negative normalized pointwise mutual information and low topic coherence scores

Topic Modeling for Heterogeneous Digital Libraries: Tailored Approaches Using Large Language Models · Virginia Tech

Considered and rejected

Considered and rejected: Rejected standard autoencoders (AE) in favor of variational autoencoders (VAE) because standard AE latent spaces are non-continuous and hinder smooth interpolation and balanced code generation

Gestalt Perception of Biological motion with a Generative Artificial Neural Network Model · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Standard Variational AutoEncoder (VAE) with KL-divergence regularization on a single latent variable was rejected for inverse continuous localization because stochastic sampling caused noisy loss landscapes.

Condition monitoring for dry cask storage using helical guided ultrsonic waves · UT Austin

Considered and rejected

Considered and rejected: Rejected Variational Autoencoders (VAEs) for conditional human mesh recovery because they do not permit direct and tractable likelihood evaluation for downstream tasks

Reconstructing 3d Humans From Images · Penn

Considered and rejected

Considered and rejected: Rejected standard VAE in favor of deterministic autoencoder because generative latent distribution modeling was unnecessary for embedding compression.

Neural Compression for Scalable Question-Answer Retrieval · DalSpace

Tried and failed

imposing latent space regularisation constraints applied to out-of-distribution detection via reconstruction error. Outcome: worse than baseline. Reason: latent constraints degrade reconstruction fidelity compared to unconstrained autoencoders, reducing anomaly detection discriminability

Learning Representations Toward the Understanding of Out-of-Distribution for Neural Networks · Georgia Tech

Tried and failed

conditional variational autoencoder with factorised latent space applied to 3D multi-object scene completion. Reason: failed to learn a structured prior, yielding incomplete reconstructions when sampling from Gaussian noise

Scene understanding for 3D multi-object scenes: labelling, reasoning and decomposing · Imperial

Tried and failed

variational autoencoders with Gaussian priors applied to systems with strongly non-Gaussian statistics. Outcome: did not generalise. Reason: standard Gaussian latent priors poorly represent strongly non-Gaussian statistics

Physics-Driven Machine Learning for Applications in Geophysical Fluid Dynamics · MIT

Tried and failed

i.i.d. latent sampling in variational autoencoders applied to set and graph generation. Reason: It optimizes an excessively loose lower bound on the true evidence lower bound.

Equivariant Neural Architectures for Representing and Generating Graphs · EPFL

Considered and rejected

Considered and rejected: Rejected likelihood-based and flow-based generative models (VAEs, PixelRNN) due to blurry outputs, slow generation, or excessive parameter counts

Deep Learning for Localized-Haptic Feedback in Tactile Surfaces · EPFL

Generative models fail to generalize beyond source training distributions

16 theses · 7 institutions

Autoencoders and generative models frequently overfit to nominal training distributions, leading to high false-alarm rates or outright failure when applied to unseen test structures and corrupted inputs. Under sparse data, occlusions, or complex physical constraints, these architectures struggle to extrapolate valid conformational landscapes or motion trajectories.

Tried and failed

MLP-based variational autoencoder applied to cross-dataset structural anomaly detection. Outcome: did not generalise. Reason: model overfitted to source structure distributions and failed on unseen structural data

Machine learning tools for identifying structural artifacts in data · Imperial

Tried and failed

generative neural network for data augmentation applied to imbalanced 3D medical image segmentation. Outcome: did not generalise. Reason: generated augmentations caused heavy overfitting to validation data and failed to generalize to unseen test data

Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial

Tried and failed

standard generative adversarial network inversion applied to image restoration under gross corruptions. Outcome: did not generalise. Reason: Restored outputs deviate heavily toward corrupted regions without prior knowledge of anomaly masks.

Robust Learning for Fine-Grained Anomaly Detection in a Data-Rich but Label-Rare Environment · Georgia Tech

Tried and failed

adversarial bias field augmentation applied to medical image segmentation domain generalization. Outcome: did not generalise. Reason: increased vulnerability to out-of-distribution spike noise artifacts, degrading segmentation performance

Improving the domain generalization and robustness of neural networks for medical imaging · Imperial

Tried and failed

compositional generative mixture models for classification applied to instance segmentation under occlusion. Outcome: did not generalise. Reason: overlapping objects with similar textures caused ambiguous boundaries and merged instance detections

Addressing Occlusion in Panoptic Segmentation · Virginia Tech

Tried and failed

variational autoencoder anomaly detection applied to semiconductor fabrication process monitoring. Outcome: overfit. Reason: overfitting on nominal training data caused high reconstruction error on unseen nominal runs, degrading classification

Applications of Probabilistic Machine Learning Models to Semiconductor Fabrication · MIT

Tried and failed

reconstruction-based neural autoencoders applied to time series anomaly detection. Outcome: did not generalise. Reason: models over-smoothed extreme values or erroneously fit out-of-range outlier points

Machine Learning Systems for Unsupervised Time Series Anomaly Detection · MIT

Tried and failed

One-Class SVM and feedforward autoencoders applied to sequential anomaly detection in logs. Outcome: did not generalise. Reason: Unable to reliably detect anomalies without supervised attack knowledge for tuning hyperparameters

Data-driven Algorithms for Critical Detection Problems: From Healthcare to Cybersecurity Defenses · Virginia Tech

Tried and failed

validation loss for autoencoder model selection applied to sequential autoencoders on sparse count data. Outcome: overfit. Reason: loss does not penalize learning the identity mapping instead of underlying latent dynamics

General and interpretable models for inferring dynamical computation in biological neural networks · Georgia Tech

Tried and failed

continuous conditional variational autoencoder applied to recursive trajectory and motion generation. Outcome: did not generalise. Reason: Continuous latent space caused blurry generations and failed to follow conditioning velocity commands over recursive rollouts

Generative Latent Motion Planning and Reinforcement Learning for Legged Locomotion · MIT

Tried and failed

variational autoencoder with normalizing flow prior applied to biomolecular conformation generation. Outcome: did not generalise. Reason: coarse-grained coordinate priors failed to capture local dihedral torsional flexibility, enforcing overly planar backbone geometries

Physically Interpretable Biomolecular Conformation Generation with A Deep Probabilistic Framework · Harvard

Tried and failed

generalist generative models for molecular dynamics applied to protein conformational landscape sampling. Outcome: did not generalise. Reason: failed to accurately sample target conformational energy landscapes compared to ground truth

Learning on Graphs with Long-Range Dependencies: Methods and Applications · EPFL

Tried and failed

variational autoencoders for sequential modeling applied to multimodal agent state estimation. Outcome: did not generalise. Reason: failed to capture multimodal distributions under sparse and partial observations

Trajectory Modeling using Generative Approaches for Scheduling, Planning, and Multi-Agent Systems · Georgia Tech

Tried and failed

transfer learning with conditional recurrent neural networks applied to generative sequence design of polymers. Outcome: data insufficient. Reason: limited labeled training data prevented meaningful sequence reconstruction and led to invalid string generation

Designing Macromolecules using Machine Learning and Simulations · MIT

Considered and rejected

Considered and rejected: Rejected conditional variational autoencoders (cVAEs) for cross-subject neural decoding mapping because directed graphical models are inflexible when adapting to new destination subjects and require deep neural networks that are difficult to fine-tune on limited data.

Score-based Approach to Analysis of Unnormalized Models and Applications · DukeSpace

Considered and rejected

Considered and rejected: Rejected using Trajectory DBSCAN and WebPPL Bayesian Programming as the core generative simulation engines because generated paths were constrained copies of training data that failed to generalize to unseen sites

Machine Learning Simulation of Pedestrians Exploring the Built Environment · MIT

Autoregressive generative forecasting accumulates compounding temporal errors

4 theses · 4 institutions

Applying generative models to multi-step recursive forecasting results in rapid error accumulation across extended prediction horizons. These temporal modeling attempts suffer from unstable rollouts, severe overfitting on time-series inputs, and inaccurate imputation.

Tried and failed

standard diffusion sampling on corrupted data models applied to generative image reconstruction. Outcome: worse than baseline. Reason: produced repetitive, low-diversity generations lacking detail without momentum adjustment

Learning generative models from corrupted data · UT Austin

Tried and failed

latent space GAN autoregressive multi-step forecasting applied to deseasonalized time series prediction. Outcome: worse than baseline. Reason: severe error accumulation during recursive forecasting in the autoencoder latent space

Utilizing Recurrent Neural Networks for Temporal Data Generation and Prediction · Virginia Tech

Tried and failed

deep generative autoregressive forecasting without state estimation applied to spatiotemporal fluid and atmospheric dynamics. Outcome: unstable. Reason: rapid accumulation of autoregressive errors over long prediction horizons

Development of generative adversarial networks for spatiotemporal fluid flow, atmospheric and flood predictions · Imperial

Tried and failed

direct time series generative adversarial network forecasting applied to time series prediction. Reason: mathematical formulation issues and implementation discrepancies in original code

Utilizing Recurrent Neural Networks for Temporal Data Generation and Prediction · Virginia Tech

Considered and rejected

Considered and rejected: Rejected using Generative AI (standard GANs/GAIN) for time-based production data due to severe overfitting and inaccurate time-series imputation.

Enhanced Oil Field Data-Wrangling using Machine Learning · Texas Tech

Left open by the authors

Problems the authors named and did not get to.

Left open

Perform additional error analysis on the variational autoencoder component of the SG-UREVA relation extraction model. Blocker: Lack of specific hypotheses, error categories, or targeted evaluation metrics to investigate

N-ary Cross-sentence Relation Extraction: From Supervised to Unsupervised Learning · Virginia Tech

Left open

Apply minimax Pareto fairness optimization to variational autoencoders to reduce reconstruction error disparities across demographic groups. Blocker: None

Minimax Fairness in Machine Learning · DukeSpace

Left open

Improve CNN autoencoder surrogate model generalization to unseen geometries and features for finite element error indicator prediction. Blocker: Lacks specific target geometries, performance metrics, and a concrete architectural approach for out-of-distribution generalization

Development of Surrogate Model for FEM Error Prediction using Deep Learning · Virginia Tech

Left open

Validate AutoEKF on synthetic data and benchmark against particle filter methods, or replace linearization with variational autoencoders. Blocker: None

TOWARDS DATA DRIVEN NETWORK EPIDEMIC MODELING. · Penn

Left open

Evaluate alternative continuous-mapping models in place of convolutional autoencoders for GAN-based synthetic EHR generation. Blocker: EHR datasets such as MIMIC typically require credentialed data access agreements

Synthetic Electronic Medical Record Generation using Generative Adversarial Networks · Virginia Tech

Left open

Optimize hyperparameters for the input stress and output finite element error autoencoder architectures. Blocker: Lack of the original synthetic FEM stress/error training datasets and detailed baseline hyperparameter search space

Development of Surrogate Model for FEM Error Prediction using Deep Learning · Virginia Tech

Left open

Replace linear experts in the Generalized Variational Autoencoder with multilayer perceptrons or partially linear mixture-of-experts formulations. Blocker: None

Supervised Variational Autoencoders for Structural Learning and Statistical Inference with Heterogeneous Data · Virginia Tech

Left open

Develop uncertainty propagation from state to value uncertainty in partially observable environments using variational autoencoders for ADFQ. Blocker: Lacks specific analytical formulation or target POMDP environments for evaluating VAE-based uncertainty propagation.

Off-Policy Temporal Difference Learning For Robotics And Autonomous Systems · Penn

Left open

Extend the CRsAE model to handle non-concentrating posteriors by integrating variational autoencoders (VAEs) for uncertainty quantification. Blocker: None

Deep Learning for Inverse Problems in Engineering and Science · Harvard

Left open

Benchmark and systematically evaluate sparse autoencoders against clustering-based concept extraction methods across standard vision and video representation learning benchmarks. Blocker: None

Disentangling Visual Concepts Across Space and Time: From Image Hierarchies to Video Dynamics · YorkSpace

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.