Chapter Four · failure evidence

What End-to-End Deep Learning got wrong, from 51 dissertations

Across diverse problem domains, end-to-end deep learning frequently suffers from severe sample inefficiency, optimization convergence issues, and poor generalization under distribution shifts. Consequently, practitioners often reject or replace end-to-end systems with modular pipelines, physics-informed architectures, and classical machine learning baselines that achieve superior performance. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

End-to-end reinforcement learning and policy models suffer from extreme sample inefficiency and long-horizon control failures

13 theses · 9 institutions

Direct sensor-to-control and policy models struggle with high sample complexity, intractable simulation requirements, and compounding imitation errors over extended horizons. They frequently converge to poor local optima, overfit simulated dynamics, and fail to transfer reliably to physical systems.

Tried and failed

direct end-to-end feedforward neural network policies applied to contact-rich robotic manipulation control. Outcome: worse than baseline. Reason: fell into local minima and overfit without structured mechanics-based constraints

An Optimization Approach to Certified Manipulation · MIT

Considered and rejected

Considered and rejected: Rejected model-free reinforcement learning algorithms (e.g., A3C) for end-to-end driving due to poor sample efficiency compared to conditional imitation learning

Towards Recognition as a Regularizer in Autonomous Driving · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Rejected monolithic end-to-end vision-language-action policies for off-road autonomy due to lack of interpretability, vulnerability to compounding imitation learning errors over long horizons, and inability to train from data lacking low-level control labels

Off-Road Navigation Under Sensing Uncertainty · ResearchWorks

Considered and rejected

Considered and rejected: Rejected end-to-end DRL learning of guidance and low-level control directly from states to actuator forces/torques because policy overfits simulated dynamics and fails during sim-to-reality transfer.

Deep Reinforcement Learning as Guidance for Aerospace Robotics · Carleton University Institutional Repository

Tried and failed

end-to-end deep reinforcement learning applied to multi-agent sensor trajectory planning. Outcome: data insufficient. Reason: mapping high-dimensional system states directly to control actions required intractable amounts of training simulation data

Mobile sensors management algorithms for environmental monitoring · UT Austin

Tried and failed

end-to-end deep reinforcement learning applied to long-horizon sequential control tasks. Outcome: did not converge. Reason: Controllers got trapped in poor local optima across various model capacities.

SPECIFICATION-GUIDED REINFORCEMENT LEARNING · Penn

Considered and rejected

Considered and rejected: Rejected pure end-to-end policy learning (RL/BC), because credit assignment and sample complexity fail over long robotics planning horizons.

Reasoning over Hierarchical Abstractions for Long-Horizon Planning in Robotics · MIT

Considered and rejected

Considered and rejected: Rejected direct end-to-end RL on raw tactile images due to sample inefficiency requiring 8 hours of physical robot wear and poor out-of-distribution generalization

TEXterity: Tactile Extrinsic deXterity · MIT

Considered and rejected

Considered and rejected: Rejected end-to-end scaling of foundation models directly to robot control due to high inference latency, real-time control constraints, and data bottlenecks.

Physical AI via Hierarchical Decision Processes · JScholarship

Considered and rejected

Considered and rejected: Rejected end-to-end continuous action RL for redundant manipulator priority control in favor of discrete priority permutation selection to fix sample inefficiency

Adaptive Control in Robotics via Deep Reinforcement Learning: From Autonomous Task Management to Zero-Shot Human Assistance · Carleton University Institutional Repository

Considered and rejected

Considered and rejected: Rejected end-to-end sensor-to-control models (e.g. imitation/reinforcement learning) because lack of structural modularity prevents reuse and requires excessive data.

Neural scene representations for dense-semantic SLAM · Imperial

Considered and rejected

Considered and rejected: End-to-end learning directly from raw camera frames was rejected to avoid sim-to-real transfer failure, maintain multi-UAV generalizability, and improve sample efficiency.

Energy-aware 3D path planning via Reinforcement Learning for aerial object detection and mapping · Iowa State

Considered and rejected

Considered and rejected: Rejected end-to-end neural policy learning for multimodal demonstration understanding in favor of modular neuro-symbolic program synthesis

Scaling Human Supervision for Robotic Manipulation · ResearchWorks

Modular pipelines and simpler statistical baselines consistently outperform end-to-end models

10 theses · 8 institutions

End-to-end architectures frequently underperform classical algorithms such as gradient-boosted decision trees, random forests, linear regression, and modular pipelines with differentiable components. In multiple evaluation settings, end-to-end models yield higher prediction error, worse convergence rates, or lower classification accuracy than simpler baselines.

Tried and failed

end-to-end deep neural networks applied to markerless relative pose estimation. Outcome: worse than baseline. Reason: produced unacceptably high relative orientation and yaw errors

Distributed Predictive Formation Control of Autonomous Rotary-Wing Micro Aerial Vehicles · EPFL

Tried and failed

end-to-end adaptive decoder training applied to neural signal decoding for gait. Outcome: worse than baseline. Reason: end-to-end optimization reduced decoding performance compared to modular training with differentiable components

TOWARDS EFFICIENT AND SCALABLE MACHINE LEARNING FOR FUTURE NEURAL INTERFACES · Cornell

Tried and failed

end-to-end decision-focused learning applied to contextual linear optimization. Outcome: worse than baseline. Reason: achieves slower regret convergence rates than estimate-then-optimize under low dual degeneracy

Machine Learning Methods for Data-driven Decision Making: Contextual Optimization, Causal Inference, and Algorithmic Fairness · Cornell

Lost to a baseline

Data-driven neural networks trained end-to-end on neural recordings were outperformed by task-driven neural network models.

Reverse engineering primate sensorimotor control with machine learning · EPFL

Lost to a baseline

When all task-relevant features are fully known a priori, feature-based Bayesian inference baselines (Coactive, FERL, DemPref) match or outperform end-to-end unstructured reward learning

Robot See, Robot Do: On the Development of Robust and Adaptive Imitation Learning for Robots · Virginia Tech

Lost to a baseline

End-to-End deep superpixel models (Pipelines F-I, 77.3%-83.2% mIoU) performed worse than standalone RGB baseline (88.8% mIoU) on 512x512 gLitter

Deep Learning Using Tiny Domain-Specific Datasets with Sparse Labels · University of Nottingham Repository

Lost to a baseline

End-to-end Neural Network with GNN embeddings underperformed Random Forest and XGBoost baselines across all metrics (Accuracy 0.5836, F1 0.4873, AUC-ROC 0.5405, AUC-PR 0.4672)

Predicting psychological treatment dropout using graph neural networks · OpenBU

Considered and rejected

Considered and rejected: Rejected using the GNN as an end-to-end prediction model, using it instead as an embedding preprocessor because GBDTs outperform deep learning on tabular data

Predicting psychological treatment dropout using graph neural networks · OpenBU

Considered and rejected

Considered and rejected: Rejected end-to-end multimodal deep networks in favor of modular 2-stage feature extraction + gradient-boosted trees for usability and data efficiency.

Multimodality: Models, Algorithms, and Applications · MIT

Lost to a baseline

End-to-end softmax dense layers achieved lower accuracy (94.8%) compared to back-end PLDA scoring (96.0%) on 2-second x-vector voice quality identification.

Intra-speaker Voice Quality Recognition for Voice Therapy · Georgia Tech

Considered and rejected

Considered and rejected: Rejected training an end-to-end WER-optimizing neural network or complex fusion model for fusing MiDaS depth and MediaPipe z-axis to avoid overfitting and alignment complexity, choosing linear regression instead.

Sign Language Recognition Using Wearable Motion Sensing and Video Co-Training · DSpace at SUNY Buffalo

Bypassing intermediate representations and domain physics causes models to plateau and overfit

10 theses · 8 institutions

Mapping raw sensor inputs directly to final targets without intermediate physical states or domain constraints forces networks to require massive datasets and violates geometric inductive biases. Relying purely on end-to-end loss functions fails to guarantee constraint satisfaction and impedes interpretability.

Tried and failed

task-guided end-to-end image-to-image translation applied to cross-domain depth completion. Outcome: worse than baseline. Reason: downstream task feedback provided insufficient constraint on image translation quality

Domain adaptation for semantic and 3D tasks · Imperial

Considered and rejected

Considered and rejected: Rejected pure end-to-end deep learning from raw symbols/samples because it requires massive datasets for every modem operating mode, fails to generalize, and is difficult to debug compared to observable physical transductions.

Noise Metrology in Optical Communication Systems · Cambridge

Considered and rejected

Considered and rejected: Rejected pure end-to-end data-driven neural regression models in favor of task-driven transfer models due to poor out-of-distribution neural explainability.

Reverse engineering primate sensorimotor control with machine learning · EPFL

Tried and failed

end-to-end training without intermediate representation supervision applied to multi-stage 3D object detection. Outcome: did not generalise. Reason: optimizing downstream loss alone creates arbitrary representations that violate geometric inductive bias

Pseudo-LiDAR: Camera-based 3D object detection for autonomous driving · Cornell

Considered and rejected

Considered and rejected: Rejected end-to-end deep learning from raw multiplexed interferograms without crude phase estimation due to needing significantly more training data and epochs to learn wave propagation

Single-shot quantitative interferometric microscopy for imaging high-speed dynamics · MIT

Considered and rejected

Considered and rejected: Rejected 'end-to-end' machine learning models lacking rheological domain knowledge because imposing physical laws via loss functions increases training time without guaranteeing invariance or constraint satisfaction on test data

Mathematics, Methods, and Models for Data-Driven Rheology · MIT

Considered and rejected

Considered and rejected: Rejected end-to-end deep learning from raw fluoroscopic images to head parameters in favor of ML (ResNet-50 + XGBoost) due to limited dataset size (~hundreds of images).

A Method to Determine Patient Eye-Lens Dose During Fluoroscopically-Guided Neuro-Interventional Procedures · DSpace at SUNY Buffalo

Tried and failed

direct end-to-end regression without intermediate state representations applied to composite failure prediction. Outcome: did not generalise. Reason: direct input-output mapping lacks sufficient intermediate physical information compared to indirect prediction, causing performance to plateau

Machine learning for predictive virtual testing of composite airframes · Imperial

Considered and rejected

Considered and rejected: End-to-end raw time-series machine learning models were rejected in favor of physically-interpretable low-dimensional feature pairs to prevent overfitting

Hemodynamics of Native and Bioprosthetic Aortic Valves: Insights from a Reduced Degree-of-Freedom Model · JScholarship

Considered and rejected

Considered and rejected: Rejected using end-to-end raw audio spectrograms without domain-specific feature engineering (as done in VGGVox) due to lack of interpretability.

INTERPRETABILITY FOR ARTIFICIAL INTELLIGENCE IN SPEAKER RECOGNITION TASKS · Calhoun

End-to-end networks fail to generalize under environment shifts and dynamic conditions

7 theses · 6 institutions

End-to-end models trained on fixed conditions suffer severe performance degradation, aliasing, and hallucinations when deployed in novel environments or under dynamic channel shifts. Without explicit intermediate feature masking or data consistency mechanisms, models overfit to their training distributions.

Tried and failed

end-to-end learning from raw images applied to perception failure prediction. Outcome: did not generalise. Reason: failed to generalize failure prediction across novel deployment environments without explicit intermediate feature masking

Introspective perception for mobile robots · UT Austin

Tried and failed

end-to-end differentiable optical-digital design applied to extended depth-of-field imaging. Outcome: did not generalise. Reason: training exclusively on planar targets prevented reconstruction of sharp edges across multiple scene depths

Programmable Optics for Computational Photography · Publikationssystem UB Tuebingen

Tried and failed

end-to-end communication training under fixed channel conditions applied to transmission across dynamic physical channels. Outcome: did not generalise. Reason: models overfit to specific channel conditions and degraded under mismatched dynamic channel environments

Deep learning enabled semantic communications with speech recognition and synthesis · Imperial

Considered and rejected

Considered and rejected: Rejected pure image-domain end-to-end deep learning networks due to lack of data consistency layers, leading to hallucinations and poor generalizability.

Optimizing reconstruction and segmentation of free-breathing whole-heart CMR to enable clinical implementation · Georgia Tech

Considered and rejected

Considered and rejected: Rejected end-to-end supervised deep learning inversion models for MRI due to severe degradation and aliasing artifacts under test-time sampling pattern or anatomy shifts.

Compressed sensing using generative models : theory and applications · UT Austin

Considered and rejected

Considered and rejected: Rejected end-to-end trained black-box CNNs for pupil phase retrieval because they require massive labeled datasets, fail outside training distributions, and lack physical interpretability.

Adaptive optics for corrections of phase and polarisation state aberrations in microscopes · Oxford

Considered and rejected

Considered and rejected: Rejected direct end-to-end learning of task failure probability p(f|z) from raw images due to extreme sample scarcity and severe overfitting in novel environments.

Introspective perception for mobile robots · UT Austin

Lost to a baseline

End-to-end models with GRU units performed worse than simple RNNs under state sequence auxiliary supervision when transition distribution shift exceeded 0.8.

COMPOSITIONAL GENERALIZATION IN INSTRUCTION FOLLOWING TASKS · Penn

Mathematical and structural obstacles impede end-to-end gradient backpropagation and convergence

4 theses · 3 institutions

Direct end-to-end optimization can completely stall when layers encounter non-differentiable operations like argmax selection or pure noise from random measurement matrices. Furthermore, gradient projection phenomena and lack of curriculum scheduling cause training to collapse or trap optimization in suboptimal states.

Tried and failed

direct end-to-end training without curriculum learning applied to continuous sequence-to-sequence recognition. Outcome: worse than baseline. Reason: learning full-length continuous sequences directly from scratch caused optimization difficulties and performance degradation

Deep audio-visual speech recognition · Imperial

Tried and failed

end-to-end learning with optimization problem layers applied to decision-making systems. Outcome: did not converge. Reason: gradient projection phenomenon impedes effective gradient backpropagation during training

Machine Learning in Decision-Making Systems: Fairness, Robustness, and Data Bias · EPFL

Considered and rejected

Considered and rejected: Rejected end-to-end trained models with random measurement matrices for each image because the network receives pure noise and fails to learn.

Compressed sensing using generative models : theory and applications · UT Austin

Tried and failed

differentiable greedy decoding via argmax selection applied to end-to-end speech and language pipelines. Reason: argmax probability selection and token mapping operations are inherently non-differentiable

Deep learning enabled semantic communications with speech recognition and synthesis · Imperial

Left open by the authors

Problems the authors named and did not get to.

Left open

Optimize hyperparameters end-to-end across the entire bidirectional CycleGAN domain adaptation network instead of tuning individual sub-networks separately. Blocker: Requires the proprietary exoskeleton sensor datasets and specific simulation-to-real training pipeline developed in the thesis.

Enabling Scalable, Versatile, and Robust Control for Robotic Exoskeletons · Georgia Tech

Left open

Incorporate end-to-end deep learning tracking architectures without separate detection heads for traffic signal operations. Blocker: None

Machine learning application powering automation of efficient traffic operation · Iowa State

Left open

Develop an end-to-end deep neural network that predicts 3D depth maps directly from fringe phase maps without post-processing reconstruction. Blocker: None

High speed 3D photomechanics testing via additional temporal sampling · Iowa State

Left open

Learn transport mappings directly via neural optimal transport to unify matching and posterior inference into an end-to-end process. Blocker: None

Informed machine learning models for advancing cardiac disease prognosis · EPFL

Left open

Train an end-to-end deep learning model to predict air pollution metrics directly from raw street view images. Blocker: None

Leveraging Street View and Remote Sensing Imagery to Enhance Air Quality Modeling through Computer Vision and Machine Learning · Virginia Tech

Left open

Train end-to-end deep learning models directly on crack pattern images using expanded experimental or numerical simulation datasets. Blocker: Requires further experimental testing apparatus or complex high-fidelity numerical simulation data not yet generated

Damage Assessment of Stone Masonry Piers Using Imaged Surface Cracks · EPFL

Left open

Combine explicit depth estimation with end-to-end learning for bird's-eye view 3D object detection and map prediction. Blocker: None

Learning Birds-Eye View Representations for Autonomous Driving · Cambridge

Left open

Extend the Implicit AutoEncoder to jointly learn trainable implicit representations end-to-end within the autoencoder architecture. Blocker: None

Representation learning for point cloud understanding · UT Austin

Left open

Develop fully automated end-to-end whole slide image analysis pipelines that generalize across clinical datasets without model drift over time. Blocker: The goal is a broad research direction without specific targets, datasets, or defined methodological approaches.

Tumor Profiling from Pathology Images using Deep Learning · Harvard

Left open

Evaluate end-to-end deep learning architectures across multimodal streams for learning-centered emotion classification. Blocker: No specific multimodal dataset or architecture is defined for this broad task

Metodología para la identificación de emociones en un ambiente educativo con aprendizaje computacional · Repositorio Institucional BUAP

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.