Chapter Four · failure evidence

What Parameter Optimization & Hyperparameter Tuning got wrong, from 47 dissertations

Doctoral researchers investigating parameter optimization and hyperparameter tuning frequently encountered issues with severe overfitting, search instability, and heavy computational demands. Across diverse applications, sophisticated tuning routines often failed to outperform untuned defaults or simpler search strategies while regularly becoming trapped in suboptimal configurations. These records come from PhD theses at 18 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Hyperparameter tuning and extended fine tuning cause severe overfitting and degrade out of sample generalization

13 theses · 10 institutions

Optimizing hyperparameters or extending fine tuning repeatedly caused models to memorize training data, overemphasize specific lag features, or suffer catastrophic forgetting. In several cases, achieving higher training likelihoods or near-perfect fits directly harmed out-of-sample prediction and validation performance.

Tried and failed

gradient-based hyperparameter optimization for sparse Gaussian processes applied to robotic manipulation dynamic modeling. Outcome: overfit. Reason: optimizers consistently overestimated signal variance, treating true functional features as noise

Decision-Making Architectures for Control of Uncertain Systems · Georgia Tech

Tried and failed

fine-tuning models solely on synthetic generated data applied to automatic speech recognition models. Outcome: overfit. Reason: training exclusively on synthetic data led to overfitting and degraded generalization to genuine speech

Voice conversion and text-to-speech for privacy protection applications · Imperial

Considered and rejected

Considered and rejected: Rejected manual divergence regularization parameter tuning (beta-NSF) in data-constrained inverse problems due to frequent severe underfitting or overfitting.

Breaking things so you don’t have to: risk assessment and failure prediction for cyber-physical AI · MIT

Considered and rejected

Considered and rejected: Rejected full unconstrained 10-sublayer free phase/amplitude parameter optimization due to severe parameter degeneracy and overfitting of noise.

Reflectivity ferromagnetic resonance for layer-resolved dynamic study of multi-layered systems · Oxford

Tried and failed

prolonged fine-tuning of pretrained diffusion models applied to medical image synthesis. Outcome: overfit. Reason: extended training caused catastrophic forgetting and overfitting, degrading reconstruction quality

Multi-contrast magnetic resonance imaging with deep generative learning · UT Austin

Tried and failed

fine-tuning pre-trained transformers with limited data applied to multivariate time series forecasting. Outcome: worse than baseline. Reason: overfitting to seasonal and occupancy patterns degraded performance below zero-shot baseline

Toward Transformer-based Large Energy Models for Smart Energy Management · Virginia Tech

Tried and failed

random forest hyperparameter tuning via grid search applied to time series error correction. Outcome: overfit. Reason: overemphasized lag-1 features, degrading downstream simulation accuracy and inflating variance

HYDRO-METEOROLOGICAL UNCERTAINTY QUANTIFICATION FOR WATER RESOURCES PLANNING AND MANAGEMENT: ADVANCES IN SYNTHETIC FORECASTING AND STOCHASTIC WATERSHED MODELS · Cornell

Considered and rejected

Considered and rejected: Rejected hyperparameter tuning on RF error correction via out-of-bag error in favor of default parameters because it over-emphasized lag-1 error and worsened simulation variance.

HYDRO-METEOROLOGICAL UNCERTAINTY QUANTIFICATION FOR WATER RESOURCES PLANNING AND MANAGEMENT: ADVANCES IN SYNTHETIC FORECASTING AND STOCHASTIC WATERSHED MODELS · Cornell

Considered and rejected

Considered and rejected: Rejected using a running-window prior ensemble because small prior variance would require tuning ad-hoc variance inflation parameters across time steps without enough proxy data to prevent overfitting.

Quantifying changes in climate and surface elevation of polar ice sheets during the last glacial-interglacial transition · ResearchWorks

Tried and failed

random restarts for marginal likelihood hyperparameter optimization applied to Gaussian process regression. Outcome: overfit. Reason: higher marginal likelihood yielded models with poorer actual predictive accuracy

Accelerating HLS Autotuning of Large, Highly-parameterized Reconfigurable SoC Mappings · Penn

Tried and failed

physics-constrained hyperparameter optimization without cross-validation applied to time series prediction in volatile regimes. Outcome: overfit. Reason: Enforcing exact energy conservation caused condensed learning with near-perfect fit that degraded out-of-sample prediction.

LSSVR + PSO = MFG · Iowa State

Tried and failed

fine-tuning singular values with extended training applied to generative adversarial network adaptation. Outcome: overfit. Reason: Extended optimization caused training set memorization despite achieving misleadingly improved quantitative evaluation metrics.

Data-Efficient Learning in Image Synthesis and Instance Segmentation · Virginia Tech

Considered and rejected

Considered and rejected: Fine-tuning generative step models for >1 epoch, rejected due to detrimental overfitting.

Equipping language models for systematic reasoning · UT Austin

Considered and rejected

Considered and rejected: Rejected STRidge and LASSO in PDE-READ due to hyperparameter fine-tuning sensitivity and bias towards over-fitted solutions

Learning Differential Equations from Noisy, Limited Data · Cornell

Sophisticated hyperparameter optimization strategies perform worse than untuned baselines or simpler random searches

6 theses · 5 institutions

Complex optimization routines such as grid search, genetic algorithms, and Bayesian methods frequently trailed untuned default model settings or simpler random sampling. Several investigations showed that untuned baseline models achieved superior or near-identical accuracy while avoiding substantial tuning effort and inflated variance.

Lost to a baseline

Kalman with GCV hyperparameter optimization performed worse than Savitzky-Golay gridsearched optima on several benchmark systems.

Open-Source Dynamical Systems Research, with a Side of (Francis) Bacon · ResearchWorks

Tried and failed

physics parameter domain randomisation applied to sim-to-real transfer in robot manipulation. Outcome: worse than baseline. Reason: did not outperform simpler random force injection despite requiring significantly higher tuning effort

Exploring sim-to-real transfer for learning-based robot manipulation · Imperial

Tried and failed

random forest Bayesian optimization for configuration tuning applied to runtime configuration optimization. Outcome: worse than baseline. Reason: direct flag-to-runtime mapping struggled to navigate high-dimensional configuration spaces effectively compared to random search

Accelerating regression testing through test environment tuning · UT Austin

Lost to a baseline

Random search with twice as many samples outperformed standard BO methods without gradient information on hyperparameter optimization tasks.

Advances in Sparse and Bayesian Optimization for Autonomous Scientific Discovery · Cornell

Considered and rejected

Considered and rejected: Rejected hyperparameter optimization of C and gamma for SVM in favor of default parameters because defaults achieved near-identical accuracy while preserving generalizability across multiple fusion schemes.

Fusion Approaches to Individual Tree Species Classification Using Multi-Source Remotely Sensed Data · YorkSpace

Lost to a baseline

In Experiment 2, Grid Search hyperparameter tuning of XGBoost (MSE 0.0078, R2 0.7602, MAPE 0.3912) and Genetic Algorithm tuning (MSE 0.0073, R2 0.7639, MAPE 0.3912) lost to the base untuned XGBoost model (MSE 0.0067, R2 0.7690, MAPE 0.3843).

Using Data Analytics and Machine Learning in Sustainable Forest Management from Remote Sensing Data · YorkSpace

Lost to a baseline

In Experiment 1, Grid Search hyperparameter tuning of XGBoost (R2 0.8600, MAPE 0.2647) and Genetic Algorithm tuning (R2 0.8635, MAPE 0.2689) lost to the base XGBoost model (R2 0.8738, MAPE 0.2525).

Using Data Analytics and Machine Learning in Sustainable Forest Management from Remote Sensing Data · YorkSpace

Lost to a baseline

Kalman smoothing with GCV hyperparameter optimization was outperformed by gridsearched Kalman smoothing and gridsearched Savitzky-Golay across multiple benchmark ODEs.

Open-Source Dynamical Systems Research, with a Side of (Francis) Bacon · ResearchWorks

Hyperparameter tuning suffers from instability and failure to generalize across varying data conditions

8 theses · 6 institutions

Tuning procedures frequently degraded when confronted with data nonidealities such as extreme outliers, leverage points, varying physical conditions, or heterogeneous textures. Authors often chose globally fixed calibrations or direct network parameterizations to avoid unstable per-scan or online adjustments that masked imaging defects.

Tried and failed

standard Gaussian process regression applied to data with leverage points and outliers. Outcome: did not generalise. Reason: mean hyperparameter optimization via weighted least squares causes severe bias from vertical outliers and bad leverage points

Robust and Data-Driven Uncertainty Quantification Methods as Real-Time Decision Support in Data-Driven Models · Virginia Tech

Tried and failed

physics-constrained kernel regression with heuristic hyperparameter optimization applied to stochastic volatility modeling across expiries. Outcome: did not generalise. Reason: Failed to handle varying expiry dates even with additional parameter calibration

LSSVR + PSO = MFG · Iowa State

Considered and rejected

Considered and rejected: Decided against application-specific hyperparameter fine-tuning to ensure strict generalizability across imaging modalities

Generalizable low-latency accelerated dynamic MRI · Iowa State

Considered and rejected

Considered and rejected: Rejected per-scan tuning/optimization of hyperparameter α across varying collimations and object positions in the feasibility study in favor of a globally fixed calibration to prevent masking imaging nonidealities.

Algorithms for Quantitative Imaging with Advanced Cone-Beam Computed Tomography Configurations · JScholarship

Considered and rejected

Considered and rejected: Treating strong convexity constants mu_i as tunable hyperparameters was rejected in favor of parametrizing a neural network mu_zeta(x, u) to eliminate manual per-datapoint hyperparameter tuning.

Engineering AI systems and AI for engineering: compositionality and physics in learning · UT Austin

Considered and rejected

Considered and rejected: Using DSSIM as a direct loss function for autoencoder training was rejected due to hyperparameter tuning instability in high-resolution, heterogeneous breast textures

Machine Learning Approaches to Improve Diagnosis and Management of Mammographic Calcifications · DukeSpace

Tried and failed

extended hyperparameter optimization applied to time series regression with extreme outliers. Reason: tuning network hyperparameters failed to eliminate or reduce extreme prediction outlier errors

Machine Learning-Based Predictive Health Model of Turbofan Engine · Virginia Tech

Considered and rejected

Considered and rejected: Rejected online test-time adaptation in favor of episodic per-core adaptation due to unstable hyperparameter tuning.

Towards High-Fidelity Prostate Tissue Characterization and Cancer Detection with Micro-Ultrasound and Deep Learning · Queens University Institutional Repository

High computational overhead and poor scalability make extensive hyperparameter tuning impractical

6 theses · 4 institutions

Exhaustive parameter searches and formal optimization algorithms encountered severe scalability bottlenecks across high-dimensional spaces. Practitioners frequently abandoned extensive tuning because the substantial increase in computation time was not justified by the marginal accuracy gains.

Considered and rejected

Considered and rejected: Rejected SMT/Max-Sat based parameter optimization for continuous control due to poor scalability, choosing Bayesian optimization instead

Programmatic reinforcement learning · UT Austin

Tried and failed

autoencoder on proper orthogonal decomposition coefficients applied to subsurface reservoir simulation. Outcome: worse than baseline. Reason: required extensive hyperparameter tuning, suffered poor consistency across scenarios, and ran slower than baselines

Fast modelling of gas reservoirs using non-intrusive reduced order modelling and machine learning · Imperial

Considered and rejected

Considered and rejected: Rejected relying solely on manual feature engineering and hyperparameter tuning in favor of AutoML using H2O library to eliminate expert bias and intensive labor.

Machine learning (ml) approaches to model interdependencies between dynamic loads and crack propagation · Cranfield

Considered and rejected

Considered and rejected: Rejected grid search for hyperparameter tuning because it was computationally infeasible across the extensive parameter space.

A DATA-DRIVEN APPROACH TO PREDICTING AUSTRALIAN BUSHFIRES · Calhoun

Considered and rejected

Considered and rejected: Rejected full hyper-parameter tuning of ML models because marginal accuracy gains were offset by large tuning time increases.

Program analysis for machine learning models · UT Austin

Considered and rejected

Considered and rejected: Neural networks were rejected for predicting phase behavior because they were less accurate than SVR/RF and required excessive hyperparameter tuning time.

Optimization of chemical enhanced oil recovery methods for naturally fractured carbonate reservoirs · UT Austin

Parameter optimization algorithms get trapped in suboptimal parameter regimes and local optima

5 theses · 5 institutions

Autonomous search routines, Bayesian optimizers, and iterative language model tuners regularly suffered from premature convergence due to low exploration or missing domain guidance. These methods became trapped in suboptimal subspaces and poor local optima instead of discovering effective global parameters.

Tried and failed

autonomous reinforcement learning parameter optimization without expert priors applied to chemical vapor deposition growth recipes. Outcome: did not converge. Reason: search frequently became trapped in poor local optima without human domain guidance

Synthesis and Applications of Large-Area Monolayer Graphene · MIT

Tried and failed

Bayesian optimization with low exploration parameter applied to neural stimulation parameter tuning. Outcome: did not converge. Reason: insufficient exploration led to premature convergence to suboptimal parameter regimes

A FRAMEWORK FOR DESIGNING DATA-DRIVEN OPTIMIZATION SYSTEMS FOR NEURAL MODULATION: WITH APPLICATIONS IN OPTOGENETIC AND ELECTRICAL BRAIN STIMULATION · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Bayesian Optimization for hyperparameter tuning due to its tendency to get trapped in local optima rather than global optima.

Optimization of Reverse Supply Chain For End-of-life Products · Cranfield

Tried and failed

LLM-driven iterative runtime hyperparameter optimization applied to large language model serving inference. Outcome: did not converge. Reason: LLM optimizer became trapped in a suboptimal parameter subspace during iterative search

LLM-Guided Run-Time Parameter Optimization for  Energy-Efficient Model Inference · Virginia Tech

Considered and rejected

Considered and rejected: Rejected Tree-structured Parzen Estimator (TPE) in favor of grid search for hyperparameter tuning due to inferior performance in initial testing.

Improving the Efficiency of Deep Reinforcement Learning based UAV Obstacle Avoidance with Edge AI · Research Repository UCD

Left open by the authors

Problems the authors named and did not get to.

Left open

Apply formal hyper-parameter optimization and overfitting safeguards to artificial neural networks used in adsorption process design. Blocker: No specific neural network architecture, baseline dataset, or optimization protocol is specified.

Computational design of multi-sorbent adsorption processes for post-combustion carbon capture · Imperial

Left open

Incorporate robust optimization techniques into simulated annealing optimal decision tree training and hyperparameter tuning to reduce overfitting to noisy data. Blocker: None

A Simulated Annealing Approach to Designing Optimal Decision Trees for Classification, Prescriptive, and Survival Analysis · MIT

Left open

Develop initialization and hyperparameter tuning algorithms for ensemble regression methods to prevent overfitting. Blocker: None

A Sequential Modeling Approach to Explain Complex Processes and Systems · Virginia Tech

Left open

Develop and implement methods for optimal hyperparameter tuning for the ENF-ADBEL neural network. Blocker: Lack of specific optimization target, baseline performance criteria, or designated optimization methodology.

DESIGN OF FUZZY LOGIC-BASED INTEGRATED ADAPTIVE DECAYED BRAIN EMOTIONAL LEARNING NETWORKS FOR ONLINE TIME SERIES PREDICTION · DalSpace

Left open

Implement automated hyperparameter optimization for the 3D-DDnet low-dose CT denoising architecture. Blocker: None

A 3D Deep Learning Architecture for Denoising Low-Dose CT Scans · Virginia Tech

Left open

Automate hyperparameter tuning for the ComputeCOVID19+ CT framework using IterML rather than perturb-and-observe methods. Blocker: None

Real-Time Computed Tomography-based Medical Diagnosis Using Deep Learning · Virginia Tech

Left open

Develop an automatic hyperparameter optimization approach for deep learning models predicting remaining useful life. Blocker: None

Optimisation of deep learning techniques on remaining useful life prediction of complex engineering systems · Cranfield

Left open

Develop an automated hyperparameter optimization method for automatic thresholding or SVD cutoff selection in 3D ultrasound super-resolution imaging. Blocker: None

Ultrafast 3-D ultrasound super-resolution imaging with a row-column array · Imperial

Left open

Incorporate Bayesian optimization, sequential experimental designs, and context information into hyperparameter tuning for image analysis models. Blocker: None

Statistical Methods for Performance Evaluation of Machine Learning and Artificial Intelligence Models · Virginia Tech

Left open

Implement automated hyperparameter tuning using grid search or Bayesian optimization for the CNN-Bi-LSTM prognostics model. Blocker: None

Intelligent Data-Driven Maintenance Planning for Marine Renewable Energy Systems · DalSpace

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.