Chapter Four · failure evidence
What Feature Extraction & Selection got wrong, from 70 dissertations
Across diverse domains, empirical feature selection and dimensionality reduction methods frequently underperform simpler baselines or unreduced full feature sets. Practitioners repeatedly observe issues ranging from the premature removal of subtle predictive signals to severe overfitting and representation degradation caused by handcrafted or transformed features. These records come from PhD theses at 25 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Aggressive feature selection discards critical predictive signals and hurts model accuracy
Pruning or filtering variables through univariate metrics, ranking thresholds, or selection algorithms frequently discarded subtle, multivariate, or minority-class signals. Consequently, models trained on the selected feature subsets consistently underperformed baselines trained on all raw features or larger sets.
Tried and failed
ANOVA feature selection applied to high-dimensional small-sample tabular data. Outcome: worse than baseline. Reason: filtering features via univariate ANOVA harmed model predictive accuracy across multiple algorithms
Predicting cancer patient response to chemotherapy using machine learning from small data · Imperial
Tried and failed
feature selection and dimensionality reduction techniques applied to FPGA compilation performance prediction. Outcome: worse than baseline. Reason: discarding features reduced the predictability performance across all models
ELECTRONIC DESIGN AUTOMATION FOR HIGH-PERFORMANCE AND RELIABLE 3D MEMORY CUBES AND PROCESSORS · Georgia Tech
Tried and failed
mRMR feature selection for dimensionality reduction applied to multivariate sensor time-series classification. Outcome: worse than baseline. Reason: aggressive feature pruning removed critical predictive signals compared to larger feature subsets or raw inputs
Evaluation of Markerless Motion Capture to Assess Physical Exposures During Material Handling Tasks · Virginia Tech
Tried and failed
statistical feature selection via bivariate correlation analysis applied to time-series state classification. Outcome: worse than baseline. Reason: selected features discarded subtle multivariate signals needed to detect complex cognitive states
Driver distraction detection using experimental methods and machine learning algorithms. · Cranfield
Tried and failed
restricting feature selection to shared platform variables applied to mass spectrometry tissue classification models. Outcome: worse than baseline. Reason: filtering for features present across both instruments discarded predictive information, lowering classification accuracy
Tried and failed
feature selection based on universal class correlation applied to multi-class tracking data classification. Outcome: worse than baseline. Reason: discarded discriminative class-specific features required to distinguish individual classes
Using Machine Learning to Differentiate Set Pieces in Football via Tracking Data · MIT
Lost to a baseline
Feature selection on gas-phase adiabatic ionization potential using random forest failed to improve performance and slightly worsened ionization potential predictions relative to full feature sets.
Lost to a baseline
All-features baseline beat all FS methods (KBest, mRMR, KGroups) on the TOX_171 dataset (91.15% vs 85.95% best FS).
Data Value Analytics for Feature Selection and Dataset Retrieval · Research Repository UCD
Lost to a baseline
Using all available features in Random Forest (RF alone) outperformed MOSS for three phenotypes, though by very small margins.
Metabolic phenotypes of marine heterotrophic bacteria · OpenBU
Lost to a baseline
Monolithic reference models using all features beat all cascade variations and stacking ensemble methods in F1 score by 3.23% to 4.69% across datasets.
Promoting the perception of emerging technologies in work environments through edge computing and hybrid intelligence · DeustoTeka
Lost to a baseline
Deterioration prediction using univariate filter selection (F1 60% ± 2%) and no feature selection (F1 38% ± 4%) lost to Recursive Feature Elimination (RFE) wrapper selection (F1 72% ± 2%).
Considered and rejected
Considered and rejected: Rejected performing feature selection on the 32 spectral/structural Random Forest predictors because reduced variable sets did not improve classification accuracy over the full model
Understory Composition, Seedling Growth, and Snag Frequency Across a Water-Limited Forest Landscape · oURspace
Tried and failed
feature selection and dimensionality reduction applied to error-related potential classification from EEG. Outcome: worse than baseline. Reason: yielded no significant performance increase over using the full feature set
Considered and rejected
Considered and rejected: Rejected feature selection dimensionality reduction (PCA, ANOVA, Fisher Score) for EEG ErrPs because unreduced covariance and correlation features performed better.
Tried and failed
information gain guided genetic algorithm feature selection applied to credit risk classification. Outcome: worse than baseline. Reason: ranking-biased genetic initialization discarded feature combinations required for non-linear nearest neighbor classification
Data mining in computational finance · Cranfield
Tried and failed
recursive feature elimination with cross validation applied to gradient boosted trees on imbalanced data. Outcome: worse than baseline. Reason: feature selection discarded weakly predictive features crucial for detecting the rare minority class
Handcrafted and engineered features underperform learned representations or simple baselines
Manually engineered features, domain heuristics, and hybrid feature sets often failed to capture complex contextual patterns and added redundant correlations. Across multiple modalities, direct raw inputs, purely data-driven models, or learned deep representations surpassed handcrafted feature sets.
Tried and failed
combining hand-engineered features with machine learning models applied to competitive predictive modeling task. Outcome: worse than baseline. Reason: hybrid feature set degraded the performance relative to purely data-driven model baselines
Essays on Open Science and Open Software in Firm Innovation · Harvard
Tried and failed
Logistic regression on TF-IDF features applied to text classification of discourse tone. Outcome: worse than baseline. Reason: TF-IDF bag-of-words failed to capture subtle stylistic and rhetorical context compared to pretrained transformers
Modeling Legal Constructs · Cornell
Tried and failed
predicting hand-crafted audio features in auxiliary tasks applied to visual speech recognition front-ends. Outcome: worse than baseline. Reason: hand-crafted spectral features provided weaker representations than self-supervised learned audio representations
Deep audio-visual speech recognition · Imperial
Tried and failed
raw operational parameters as ML features applied to aerodynamic flutter prediction. Outcome: worse than baseline. Reason: Derived physical and acoustic features provided stronger predictive signal than raw independent operating variables.
Data-driven modelling of compressor stall flutter · Imperial
Tried and failed
handcrafted feature extraction applied to interactive multidimensional scaling for images. Outcome: worse than baseline. Reason: color histograms and SIFT yielded less meaningful low-dimensional projections and visual explanations than CNN features
Explainable Interactive Projections for Image Data · Virginia Tech
Tried and failed
augmenting learned latent representations with handcrafted features applied to crop yield prediction. Outcome: worse than baseline. Reason: handcrafted features were redundant and highly cross-correlated with learned latent environment representations
Machine learning for mapping multimodal, multiscale phenotypic data to yield · Iowa State
Tried and failed
intermediate attribute modeling before downstream prediction applied to long-term outcome prediction from text. Outcome: worse than baseline. Reason: Intermediate human-annotated proxy features captured less predictive signal than end-to-end direct outcome modeling
EMPIRICAL EVIDENCE THAT USING AI TOOLS CAN ENHANCE HUMAN COGNITION · Penn
Tried and failed
handcrafted statistical features applied to nanopore signal classification. Outcome: worse than baseline. Reason: did not scale well with an increasing number of classes compared to recurrent neural networks
Tried and failed
augmenting pretrained language models with handcrafted features applied to source code vulnerability classification. Outcome: worse than baseline. Reason: supplementary metrics and multi-task learning did not provide additional useful signal over pretrained representations
Deep representation learning for software vulnerability detection · Imperial
Tried and failed
direct lexical word features in discrete choice applied to social media engagement prediction. Outcome: worse than baseline. Reason: yielded lower prediction accuracy and poorer interpretability than aggregate topic features
Tried and failed
classical machine learning with hand-engineered features applied to single-molecule fluorescence time-series classification. Outcome: worse than baseline. Reason: hand-crafted timing and peak features lost critical temporal patterns needed for accurate peptide identification
Tried and failed
support vector regression with domain-specific engineered features applied to time-series signal regression. Outcome: worse than baseline. Reason: adding morphological signal-specific features degraded regression accuracy compared to standard time-series features alone
Inkjet-Printed ECG Electrodes and AI-based Data Reliability for Smart-Health CPS · Texas Tech
Considered and rejected
Considered and rejected: Rejected discarding raw sensor inputs in favour of manual feature engineering (e.g., PCA, linear discriminant analysis), selecting end-to-end multi-layer feature extraction instead
Deep learning methods for improving diabetes management tools · Imperial
Tried and failed
genetic programming for automated feature engineering applied to occupancy forecasting. Outcome: worse than baseline. Reason: Failed to provide accuracy gains over manual feature extraction
The digitalization of energy systems: towards higher energy efficiency · EPFL
Feature selection algorithms overfit training data in high-dimensional or small-sample regimes
Techniques like genetic algorithms, unconstrained wrapper searches, and standard cross-validation frequently overfitted noise and unannotated artifacts in high-dimensional data. Selecting features without strong regularization or uncertainty constraints degraded generalisation on validation and test sets.
Tried and failed
genetic algorithm feature selection on concatenated multimodal data applied to high-dimensional omics classification. Outcome: overfit. Reason: selected non-biological artifacts, unannotated features, and exogenous confounders from the combined assay dataset
Metabolomics and Machine Learning for Early-Stage Cancer Diagnosis · Georgia Tech
Tried and failed
feature selection maximizing training accuracy without uncertainty constraints applied to radiogenomics regression models. Outcome: overfit. Reason: optimizing training accuracy selected too many features, increasing complexity and degrading cross-validation performance
UNCERTAINTY MITIGATION IN IMAGE-BASED MACHINE LEARNING MODELS FOR PRECISION MEDICINE · Georgia Tech
Considered and rejected
Considered and rejected: Rejected standard cross-validation approaches for feature selection in small-sample microarray settings due to severe classification instability.
Power enhanced gene differential expression analysis by incorporating gene network information · Iowa State
Considered and rejected
Considered and rejected: Rejected BIC and AIC feature selection for downscaling GAMs because they selected all candidate features and overfit noise without physical justification.
Climate and chronology: environmental variability and the archaeology of marine isotope stage 3 · Oxford
Tried and failed
random forest classification applied to high-dimensional sparse genomic variant features. Outcome: overfit. Reason: severe overfitting during training on high-dimensional genomic features compared to logistic regression baseline
Decoding Germline Genetic Influence on Cancer Somatic Mutation Acquisition · Harvard
Tried and failed
genetic algorithm feature selection with linear regression applied to high-dimensional genomic regression datasets. Outcome: overfit. Reason: linear regression without regularization severely overfits in high-dimensional settings during wrapper feature evaluation
Application of Optimization and Simulation Models in Genomic Prediction and Genomic Selection · Iowa State
Lost to a baseline
Random Forest classifier test accuracy dropped more severely (from 0.78 to 0.69) compared to Support Vector Machine (stayed at 0.71) when pruned to top four LIME-selected features due to overfitting on the full feature set.
Multiscale Integration of Cross-Modal Subsurface Data for Reservoir Characterization under Label-Constrained Environments · Georgia Tech
Unsupervised projection and dimensionality reduction degrade downstream prediction performance
Applying techniques like PCA, KPCA, or Isomap prior to modeling often overcomplicated data representations and obscured localized discriminative signals. As a result, dimensionality reduction transformations underperformed direct regression, raw untransformed features, or simple summary metrics.
Considered and rejected
Considered and rejected: Rejected Principal Component Analysis (PCA) because zero-variance features and categorical flag data caused standard PCA algorithms to fail.
An Analysis of Grey Theory for Personal Affective Computing · De Montfort Open Research Archive (DORA)
Considered and rejected
Considered and rejected: Rejected including all 17 embedded categorical features in PCA/SVD due to memory/virtual machine crashes and extensive NAs in sparse port embeddings.
Intruder Alert: Dimension Reduction and Density-Based Clustering for a Cybersecurity Application · Carleton University Institutional Repository
Tried and failed
Principal component analysis for feature extraction applied to microstructural image representations. Outcome: worse than baseline. Reason: overcomplicated data representation, degrading downstream surrogate model accuracy compared to simple summary metrics
Exploration of the Additive Manufacturing Process Development Space using High-throughput Mechanical Property Assays · Georgia Tech
Tried and failed
PCA dimensionality reduction before regression applied to tabular physical test data. Outcome: worse than baseline. Reason: PCA feature extraction did not systematically improve predictive performance compared to direct regression
Compacted Snow Testing Methodology and Instrumentation · Virginia Tech
Tried and failed
PCA dimensionality reduction before classification applied to high-dimensional tabular genomic data. Outcome: worse than baseline. Reason: Reduced classification performance compared to raw features and obstructed feature-level SHAP interpretability
Applications of Machine Learning in Source Attribution and Gene Function Prediction · Virginia Tech
Lost to a baseline
Kernel-PCA (KPCA) and Isomap feature extraction transformations were outperformed by raw untransformed power features in K-Means clustering (yielding lower Silhouette scores down to -0.15 and lower Calinski-Harabasz indices).
Lost to a baseline
PCA-based feature selection (Jaccard 0.9624) was beaten by random feature selection (Jaccard 0.985) in K-Means segmentation dimensionality reduction
MORPHOLOGICAL CLASSIFICATION OF SUBTYPES OF VOLUMETRIC PROJECTION NEURONS FROM MOUSE BRAIN SCANS · JScholarship
Tried and failed
weighted eigenvector features for classification applied to signal defect classification. Outcome: worse than baseline. Reason: weighted eigenvectors introduced noise or variance compared to using simpler eigenvalue magnitudes alone
Ultrasonic signal processing and classification using principal component analysis · Iowa State
Time-frequency transforms lose critical localized information or fail across models
Wavelet, spectral, and Fourier transforms frequently suffered from resolution degradation, blurring, or incompatibility with nonlinear models. These transforms degraded classification and regression performance compared to raw time-domain signals, shallow neural networks, or standard alternative features.
Tried and failed
cumulative energy spectral analysis on stacked signals applied to seismic reflection data. Outcome: worse than baseline. Reason: wavelet stretching during normal moveout correction blurred and degraded spectral resolution features
Interpretation of reflection seismic data by analysis of cumulative energy spectra · Virginia Tech
Tried and failed
GMM and SVM classifiers applied to wavelet-based neural signal decoding. Outcome: worse than baseline. Reason: could not match performance of shallow neural networks on RMS wavelet features
Neural speech decoding with magnetoencephalography · UT Austin
Tried and failed
Daubechies 12 wavelet feature extraction applied to speech phoneme recognition. Outcome: worse than baseline. Reason: None
Speech phoneme recognition using wavelets and artificial neural networks · Iowa State
Tried and failed
directional wave component separation for feature extraction applied to pulse waveform disease classification. Outcome: worse than baseline. Reason: separated forward and backward components did not improve SVM performance over unseparated signals
Improved pulse wave analysis for diagnosing and monitoring heart failure · Imperial
Tried and failed
continuous wavelet transform feature extraction applied to time series correlation and regression. Reason: inconsistent time resolution across frequency bins prevented systematic correlation sweeping
Developing acoustic emission techniques for condition monitoring of rubbing contacts · Imperial
Tried and failed
wavelet transform features for neural networks applied to time-series regression. Outcome: worse than baseline. Reason: wavelet features failed with nonlinear neural models despite performing well with linear regression
Tried and failed
autocorrelation shell transform denoising applied to non-smooth piecewise constant signals. Outcome: worse than baseline. Reason: standard wavelets better capture sharp discontinuities and localized non-smooth features
Topics on Multiresolution Signal Processing and Bayesian Modeling with Applications in Bioinformatics · Georgia Tech
Tried and failed
FFT magnitude features with PCA classification applied to aligned ultrasonic A-scan defect signals. Outcome: worse than baseline. Reason: Fourier magnitude spectrum lost phase and localized time-frequency defect details compared to wavelets or DCT
Ultrasonic signal processing and classification using principal component analysis · Iowa State
Considered and rejected
Considered and rejected: Discarded the use of wavelet features as input for sliding-window decoding because it performed significantly worse across most timesteps compared to raw time-domain signals.
Decoding non-invasive brain activity with novel deep-learning approaches · Oxford
Feature importance metrics exhibit systemic bias toward specific variable types
Random Forest feature importance metrics such as Mean Decrease Gini or Mean Decrease in Impurity exhibited structural ranking biases. These methods systematically favored continuous variables and high-cardinality categorical features over other predictive inputs.
Tried and failed
heuristic threshold-based feature selection from tree importances applied to high-dimensional tabular metagenomic abundance data. Reason: empirical threshold selection retained excessive non-informative features compared to Bayesian-optimized thresholding
Antibiotic Resistance Characterization in Human Fecal and Environmental Resistomes using Metagenomics and Machine Learning · Virginia Tech
Tried and failed
Mean Decrease Gini feature importance applied to tabular random forest feature selection. Reason: ranking bias towards continuous and high-cardinality categorical features
Retrofitting U.S. Suburbia in the Era of Shared Autonomous Vehicles · Georgia Tech
Tried and failed
random forest mean decrease in impurity applied to feature importance estimation. Reason: biased toward continuous variables and features with many categories
Feature selection assumptions fail under context dependence and distribution shifts
Selected feature subsets frequently failed to transfer when external factors masked risk components or when sequences varied in length. Assuming context independence or invariant feature relationships caused models to degrade severely across different experimental backgrounds.
Tried and failed
ANOVA feature selection applied to multivariate time series classification. Outcome: did not generalise. Reason: failed to yield stable or high performance across varying sequence lengths
Temporal learning techniques for vehicle cybersecurity · Texas Tech
Tried and failed
cross-feature-representation training and testing applied to linear classification of trial-level signals. Outcome: did not generalise. Reason: training on regularized mixed-effects features did not generalize to standard least-squares estimates at test time
Extracting Feature Vectors From Event-Related fMRI Data to Enable Machine Learning Analysis · Virginia Tech
Tried and failed
maximal correlation and mutual information regularization applied to domain generalization under spurious correlation. Outcome: did not generalise. Reason: regularization suppressed spurious features but failed to learn fully invariant representations, yielding random chance accuracy
Maximal Correlation Feature Selection and Suppression With Applications · MIT
Tried and failed
two-phase linear regression with stepwise feature selection applied to bond investment risk premium prediction. Outcome: did not generalise. Reason: external credit guarantees masked idiosyncratic risk factors, drastically reducing model explanatory power
Financial risk : charter school-specific risk factors and idiosyncratic risk · UT Austin
Tried and failed
gradient-ordered feature selection assuming context independence applied to genetic locus fitness across environments. Outcome: did not generalise. Reason: features exhibited strong background dependence, violating the assumption of hub independence across backgrounds
The structure of fitness landscapes across genotypes and environments · Harvard
Considered and rejected
Considered and rejected: Rejected global feature selection methods (Decision Trees, Random Forests, and L1-regularized Logistic Regression) because they do not capture instance-specific variability and lack runtime adaptability.
Left open by the authors
Problems the authors named and did not get to.
Left open
Perform systematic feature selection and reduction using techniques like RFE and PCA on the 17 identified search and instance features. Blocker: None
Automated design of population-based algorithms: a case study in vehicle routing · University of Nottingham Repository
Left open
Benchmark co-segregation-based Bayesian linear regression feature selection against standard methods like PCA and random forests for microbial phenotypic GWAS. Blocker: None
Probing Genomic, Metabolic, And Phenotypic Evolution In Microbes Using Comparative And Experimental Evolution Data · Georgia Tech
Left open
Implement alternative feature scoring metrics like PCA loadings, Katz centrality, and PageRank into the HARVEST microbiome feature selection pipeline. Blocker: None
A Modular Data Analytic Pipeline for Feature Selection in High Dimensional Microbial Data Sets · HARVEST
Left open
Extract additional meta-features and evaluate a larger, more diverse pool of feature selection methods within the meta-learning framework. Blocker: No specific meta-features or feature selection algorithms are specified to evaluate
An investigation of fuzzy methods and meta learning for feature selection · University of Nottingham Repository
Left open
Identify significant predictors of composite stress distributions across loading stages using feature extraction techniques on the U-Net model. Blocker: The goal is too vague regarding specific feature extraction methods and target metrics
A Deep Learning Approach to Predict Full-Field Stress Distribution in Composite Materials · Virginia Tech
Left open
Perform feature selection and incorporate hierarchical relationships between subcellular locations into CFMS protein localization classifiers. Blocker: The proposals (improve input quality, feature selection, location relationships) lack specific implementation details or defined mathematical approaches
Left open
Perform feature selection to reduce the 17 longitudinal EHR features to a minimal informative subset for delirium classification. Blocker: Requires the private clinical ICU EHR dataset used in the thesis.
Automatic delirium classification in intensive care using non-invasive eye tracking · Imperial
Left open
Evaluate the reduced feature subsets produced by the fuzzy feature selection methods using classifiers like SVM and Random Forest. Blocker: None
An investigation of fuzzy methods and meta learning for feature selection · University of Nottingham Repository
Left open
Reverse engineer features learned by weak RFML learners to identify causes of recurring misclassifications without training full ensembles. Blocker: None
Sensitivity Analysis of RFML-based SEI Algorithms · Virginia Tech
Left open
Evaluate alternative feature selection methods within the class-imbalance classification pipeline to optimize feature subsets. Blocker: None
Machine Learning Approaches for Healthcare Analysis · Scholarship at UWindsor Institutional Repository
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.