Chapter Four · failure evidence

What Feature Extraction & Selection got wrong, from 70 dissertations

Across diverse domains, empirical feature selection and dimensionality reduction methods frequently underperform simpler baselines or unreduced full feature sets. Practitioners repeatedly observe issues ranging from the premature removal of subtle predictive signals to severe overfitting and representation degradation caused by handcrafted or transformed features. These records come from PhD theses at 25 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Aggressive feature selection discards critical predictive signals and hurts model accuracy

15 theses · 11 institutions

Pruning or filtering variables through univariate metrics, ranking thresholds, or selection algorithms frequently discarded subtle, multivariate, or minority-class signals. Consequently, models trained on the selected feature subsets consistently underperformed baselines trained on all raw features or larger sets.

Tried and failed

ANOVA feature selection applied to high-dimensional small-sample tabular data. Outcome: worse than baseline. Reason: filtering features via univariate ANOVA harmed model predictive accuracy across multiple algorithms

Predicting cancer patient response to chemotherapy using machine learning from small data · Imperial

Tried and failed

feature selection and dimensionality reduction techniques applied to FPGA compilation performance prediction. Outcome: worse than baseline. Reason: discarding features reduced the predictability performance across all models

ELECTRONIC DESIGN AUTOMATION FOR HIGH-PERFORMANCE AND RELIABLE 3D MEMORY CUBES AND PROCESSORS · Georgia Tech

Tried and failed

mRMR feature selection for dimensionality reduction applied to multivariate sensor time-series classification. Outcome: worse than baseline. Reason: aggressive feature pruning removed critical predictive signals compared to larger feature subsets or raw inputs

Evaluation of Markerless Motion Capture to Assess Physical Exposures During Material Handling Tasks · Virginia Tech

Tried and failed

statistical feature selection via bivariate correlation analysis applied to time-series state classification. Outcome: worse than baseline. Reason: selected features discarded subtle multivariate signals needed to detect complex cognitive states

Driver distraction detection using experimental methods and machine learning algorithms. · Cranfield

Tried and failed

restricting feature selection to shared platform variables applied to mass spectrometry tissue classification models. Outcome: worse than baseline. Reason: filtering for features present across both instruments discarded predictive information, lowering classification accuracy

Advancements in ambient ionization and MALDI mass spectrometry imaging methods and their application in disease diagnosis and characterization · UT Austin

Tried and failed

feature selection based on universal class correlation applied to multi-class tracking data classification. Outcome: worse than baseline. Reason: discarded discriminative class-specific features required to distinguish individual classes

Using Machine Learning to Differentiate Set Pieces in Football via Tracking Data · MIT

Lost to a baseline

Feature selection on gas-phase adiabatic ionization potential using random forest failed to improve performance and slightly worsened ionization potential predictions relative to full feature sets.

Using Data-Driven Models to Understand Transition Metal Catalyst Energy Landscapes and Metal-Organic Framework Stability · MIT

Lost to a baseline

All-features baseline beat all FS methods (KBest, mRMR, KGroups) on the TOX_171 dataset (91.15% vs 85.95% best FS).

Data Value Analytics for Feature Selection and Dataset Retrieval · Research Repository UCD

Lost to a baseline

Using all available features in Random Forest (RF alone) outperformed MOSS for three phenotypes, though by very small margins.

Metabolic phenotypes of marine heterotrophic bacteria · OpenBU

Lost to a baseline

Monolithic reference models using all features beat all cascade variations and stacking ensemble methods in F1 score by 3.23% to 4.69% across datasets.

Promoting the perception of emerging technologies in work environments through edge computing and hybrid intelligence · DeustoTeka

Lost to a baseline

Deterioration prediction using univariate filter selection (F1 60% ± 2%) and no feature selection (F1 38% ± 4%) lost to Recursive Feature Elimination (RFE) wrapper selection (F1 72% ± 2%).

Wearable sensors and machine learning for tracking movement, deterioration, and recovery in acute in-patients · Imperial

Considered and rejected

Considered and rejected: Rejected performing feature selection on the 32 spectral/structural Random Forest predictors because reduced variable sets did not improve classification accuracy over the full model

Understory Composition, Seedling Growth, and Snag Frequency Across a Water-Limited Forest Landscape · oURspace

Tried and failed

feature selection and dimensionality reduction applied to error-related potential classification from EEG. Outcome: worse than baseline. Reason: yielded no significant performance increase over using the full feature set

Robots as Minions, Sidekicks, and Apprentices: Using Wearable Muscle, Brain, and Motion Sensors for Plug-and-Play Human-Robot Interaction · MIT

Considered and rejected

Considered and rejected: Rejected feature selection dimensionality reduction (PCA, ANOVA, Fisher Score) for EEG ErrPs because unreduced covariance and correlation features performed better.

Robots as Minions, Sidekicks, and Apprentices: Using Wearable Muscle, Brain, and Motion Sensors for Plug-and-Play Human-Robot Interaction · MIT

Tried and failed

information gain guided genetic algorithm feature selection applied to credit risk classification. Outcome: worse than baseline. Reason: ranking-biased genetic initialization discarded feature combinations required for non-linear nearest neighbor classification

Data mining in computational finance · Cranfield

Tried and failed

recursive feature elimination with cross validation applied to gradient boosted trees on imbalanced data. Outcome: worse than baseline. Reason: feature selection discarded weakly predictive features crucial for detecting the rare minority class

Optimizing Suicide Prevention in Adolescents: A Longitudinal Approach to Risk Modeling and Resource Allocation · Harvard

Handcrafted and engineered features underperform learned representations or simple baselines

14 theses · 8 institutions

Manually engineered features, domain heuristics, and hybrid feature sets often failed to capture complex contextual patterns and added redundant correlations. Across multiple modalities, direct raw inputs, purely data-driven models, or learned deep representations surpassed handcrafted feature sets.

Tried and failed

combining hand-engineered features with machine learning models applied to competitive predictive modeling task. Outcome: worse than baseline. Reason: hybrid feature set degraded the performance relative to purely data-driven model baselines

Essays on Open Science and Open Software in Firm Innovation · Harvard

Tried and failed

Logistic regression on TF-IDF features applied to text classification of discourse tone. Outcome: worse than baseline. Reason: TF-IDF bag-of-words failed to capture subtle stylistic and rhetorical context compared to pretrained transformers

Modeling Legal Constructs · Cornell

Tried and failed

predicting hand-crafted audio features in auxiliary tasks applied to visual speech recognition front-ends. Outcome: worse than baseline. Reason: hand-crafted spectral features provided weaker representations than self-supervised learned audio representations

Deep audio-visual speech recognition · Imperial

Tried and failed

raw operational parameters as ML features applied to aerodynamic flutter prediction. Outcome: worse than baseline. Reason: Derived physical and acoustic features provided stronger predictive signal than raw independent operating variables.

Data-driven modelling of compressor stall flutter · Imperial

Tried and failed

handcrafted feature extraction applied to interactive multidimensional scaling for images. Outcome: worse than baseline. Reason: color histograms and SIFT yielded less meaningful low-dimensional projections and visual explanations than CNN features

Explainable Interactive Projections for Image Data · Virginia Tech

Tried and failed

augmenting learned latent representations with handcrafted features applied to crop yield prediction. Outcome: worse than baseline. Reason: handcrafted features were redundant and highly cross-correlated with learned latent environment representations

Machine learning for mapping multimodal, multiscale phenotypic data to yield · Iowa State

Tried and failed

intermediate attribute modeling before downstream prediction applied to long-term outcome prediction from text. Outcome: worse than baseline. Reason: Intermediate human-annotated proxy features captured less predictive signal than end-to-end direct outcome modeling

EMPIRICAL EVIDENCE THAT USING AI TOOLS CAN ENHANCE HUMAN COGNITION · Penn

Tried and failed

handcrafted statistical features applied to nanopore signal classification. Outcome: worse than baseline. Reason: did not scale well with an increasing number of classes compared to recurrent neural networks

A Geometric Transformer for Structural Biology: Development and Applications of the Protein Structure Transformer · EPFL

Tried and failed

augmenting pretrained language models with handcrafted features applied to source code vulnerability classification. Outcome: worse than baseline. Reason: supplementary metrics and multi-task learning did not provide additional useful signal over pretrained representations

Deep representation learning for software vulnerability detection · Imperial

Tried and failed

direct lexical word features in discrete choice applied to social media engagement prediction. Outcome: worse than baseline. Reason: yielded lower prediction accuracy and poorer interpretability than aggregate topic features

How Words Move Hearts: Interpretable Machine Learning Models of Bias, Engagement, and Influence in Socio-Political Systems · EPFL

Tried and failed

classical machine learning with hand-engineered features applied to single-molecule fluorescence time-series classification. Outcome: worse than baseline. Reason: hand-crafted timing and peak features lost critical temporal patterns needed for accurate peptide identification

Development of tools and methods for protein identification: From single molecules to in vivo applications · EPFL

Tried and failed

support vector regression with domain-specific engineered features applied to time-series signal regression. Outcome: worse than baseline. Reason: adding morphological signal-specific features degraded regression accuracy compared to standard time-series features alone

Inkjet-Printed ECG Electrodes and AI-based Data Reliability for Smart-Health CPS · Texas Tech

Considered and rejected

Considered and rejected: Rejected discarding raw sensor inputs in favour of manual feature engineering (e.g., PCA, linear discriminant analysis), selecting end-to-end multi-layer feature extraction instead

Deep learning methods for improving diabetes management tools · Imperial

Tried and failed

genetic programming for automated feature engineering applied to occupancy forecasting. Outcome: worse than baseline. Reason: Failed to provide accuracy gains over manual feature extraction

The digitalization of energy systems: towards higher energy efficiency · EPFL

Feature selection algorithms overfit training data in high-dimensional or small-sample regimes

7 theses · 4 institutions

Techniques like genetic algorithms, unconstrained wrapper searches, and standard cross-validation frequently overfitted noise and unannotated artifacts in high-dimensional data. Selecting features without strong regularization or uncertainty constraints degraded generalisation on validation and test sets.

Tried and failed

genetic algorithm feature selection on concatenated multimodal data applied to high-dimensional omics classification. Outcome: overfit. Reason: selected non-biological artifacts, unannotated features, and exogenous confounders from the combined assay dataset

Metabolomics and Machine Learning for Early-Stage Cancer Diagnosis · Georgia Tech

Tried and failed

feature selection maximizing training accuracy without uncertainty constraints applied to radiogenomics regression models. Outcome: overfit. Reason: optimizing training accuracy selected too many features, increasing complexity and degrading cross-validation performance

UNCERTAINTY MITIGATION IN IMAGE-BASED MACHINE LEARNING MODELS FOR PRECISION MEDICINE · Georgia Tech

Considered and rejected

Considered and rejected: Rejected standard cross-validation approaches for feature selection in small-sample microarray settings due to severe classification instability.

Power enhanced gene differential expression analysis by incorporating gene network information · Iowa State

Considered and rejected

Considered and rejected: Rejected BIC and AIC feature selection for downscaling GAMs because they selected all candidate features and overfit noise without physical justification.

Climate and chronology: environmental variability and the archaeology of marine isotope stage 3 · Oxford

Tried and failed

random forest classification applied to high-dimensional sparse genomic variant features. Outcome: overfit. Reason: severe overfitting during training on high-dimensional genomic features compared to logistic regression baseline

Decoding Germline Genetic Influence on Cancer Somatic Mutation Acquisition · Harvard

Tried and failed

genetic algorithm feature selection with linear regression applied to high-dimensional genomic regression datasets. Outcome: overfit. Reason: linear regression without regularization severely overfits in high-dimensional settings during wrapper feature evaluation

Application of Optimization and Simulation Models in Genomic Prediction and Genomic Selection · Iowa State

Lost to a baseline

Random Forest classifier test accuracy dropped more severely (from 0.78 to 0.69) compared to Support Vector Machine (stayed at 0.71) when pruned to top four LIME-selected features due to overfitting on the full feature set.

Multiscale Integration of Cross-Modal Subsurface Data for Reservoir Characterization under Label-Constrained Environments · Georgia Tech

Unsupervised projection and dimensionality reduction degrade downstream prediction performance

8 theses · 7 institutions

Applying techniques like PCA, KPCA, or Isomap prior to modeling often overcomplicated data representations and obscured localized discriminative signals. As a result, dimensionality reduction transformations underperformed direct regression, raw untransformed features, or simple summary metrics.

Considered and rejected

Considered and rejected: Rejected Principal Component Analysis (PCA) because zero-variance features and categorical flag data caused standard PCA algorithms to fail.

An Analysis of Grey Theory for Personal Affective Computing · De Montfort Open Research Archive (DORA)

Considered and rejected

Considered and rejected: Rejected including all 17 embedded categorical features in PCA/SVD due to memory/virtual machine crashes and extensive NAs in sparse port embeddings.

Intruder Alert: Dimension Reduction and Density-Based Clustering for a Cybersecurity Application · Carleton University Institutional Repository

Tried and failed

Principal component analysis for feature extraction applied to microstructural image representations. Outcome: worse than baseline. Reason: overcomplicated data representation, degrading downstream surrogate model accuracy compared to simple summary metrics

Exploration of the Additive Manufacturing Process Development Space using High-throughput Mechanical Property Assays · Georgia Tech

Tried and failed

PCA dimensionality reduction before regression applied to tabular physical test data. Outcome: worse than baseline. Reason: PCA feature extraction did not systematically improve predictive performance compared to direct regression

Compacted Snow Testing Methodology and Instrumentation · Virginia Tech

Tried and failed

PCA dimensionality reduction before classification applied to high-dimensional tabular genomic data. Outcome: worse than baseline. Reason: Reduced classification performance compared to raw features and obstructed feature-level SHAP interpretability

Applications of Machine Learning in Source Attribution and Gene Function Prediction · Virginia Tech

Lost to a baseline

Kernel-PCA (KPCA) and Isomap feature extraction transformations were outperformed by raw untransformed power features in K-Means clustering (yielding lower Silhouette scores down to -0.15 and lower Calinski-Harabasz indices).

Study of retrofitted system for Intelligent Compaction Analyzer, a machine learning approach for Quality Control of Asphalt Pavement during Construction · unevada

Lost to a baseline

PCA-based feature selection (Jaccard 0.9624) was beaten by random feature selection (Jaccard 0.985) in K-Means segmentation dimensionality reduction

MORPHOLOGICAL CLASSIFICATION OF SUBTYPES OF VOLUMETRIC PROJECTION NEURONS FROM MOUSE BRAIN SCANS · JScholarship

Tried and failed

weighted eigenvector features for classification applied to signal defect classification. Outcome: worse than baseline. Reason: weighted eigenvectors introduced noise or variance compared to using simpler eigenvalue magnitudes alone

Ultrasonic signal processing and classification using principal component analysis · Iowa State

Time-frequency transforms lose critical localized information or fail across models

9 theses · 7 institutions

Wavelet, spectral, and Fourier transforms frequently suffered from resolution degradation, blurring, or incompatibility with nonlinear models. These transforms degraded classification and regression performance compared to raw time-domain signals, shallow neural networks, or standard alternative features.

Tried and failed

cumulative energy spectral analysis on stacked signals applied to seismic reflection data. Outcome: worse than baseline. Reason: wavelet stretching during normal moveout correction blurred and degraded spectral resolution features

Interpretation of reflection seismic data by analysis of cumulative energy spectra · Virginia Tech

Tried and failed

GMM and SVM classifiers applied to wavelet-based neural signal decoding. Outcome: worse than baseline. Reason: could not match performance of shallow neural networks on RMS wavelet features

Neural speech decoding with magnetoencephalography · UT Austin

Tried and failed

Daubechies 12 wavelet feature extraction applied to speech phoneme recognition. Outcome: worse than baseline. Reason: None

Speech phoneme recognition using wavelets and artificial neural networks · Iowa State

Tried and failed

directional wave component separation for feature extraction applied to pulse waveform disease classification. Outcome: worse than baseline. Reason: separated forward and backward components did not improve SVM performance over unseparated signals

Improved pulse wave analysis for diagnosing and monitoring heart failure · Imperial

Tried and failed

continuous wavelet transform feature extraction applied to time series correlation and regression. Reason: inconsistent time resolution across frequency bins prevented systematic correlation sweeping

Developing acoustic emission techniques for condition monitoring of rubbing contacts · Imperial

Tried and failed

wavelet transform features for neural networks applied to time-series regression. Outcome: worse than baseline. Reason: wavelet features failed with nonlinear neural models despite performing well with linear regression

Virtual metrology applied to milling process · EPFL

Tried and failed

autocorrelation shell transform denoising applied to non-smooth piecewise constant signals. Outcome: worse than baseline. Reason: standard wavelets better capture sharp discontinuities and localized non-smooth features

Topics on Multiresolution Signal Processing and Bayesian Modeling with Applications in Bioinformatics · Georgia Tech

Tried and failed

FFT magnitude features with PCA classification applied to aligned ultrasonic A-scan defect signals. Outcome: worse than baseline. Reason: Fourier magnitude spectrum lost phase and localized time-frequency defect details compared to wavelets or DCT

Ultrasonic signal processing and classification using principal component analysis · Iowa State

Considered and rejected

Considered and rejected: Discarded the use of wavelet features as input for sliding-window decoding because it performed significantly worse across most timesteps compared to raw time-domain signals.

Decoding non-invasive brain activity with novel deep-learning approaches · Oxford

Feature importance metrics exhibit systemic bias toward specific variable types

3 theses · 3 institutions

Random Forest feature importance metrics such as Mean Decrease Gini or Mean Decrease in Impurity exhibited structural ranking biases. These methods systematically favored continuous variables and high-cardinality categorical features over other predictive inputs.

Tried and failed

heuristic threshold-based feature selection from tree importances applied to high-dimensional tabular metagenomic abundance data. Reason: empirical threshold selection retained excessive non-informative features compared to Bayesian-optimized thresholding

Antibiotic Resistance Characterization in Human Fecal and Environmental Resistomes using Metagenomics and Machine Learning · Virginia Tech

Tried and failed

Mean Decrease Gini feature importance applied to tabular random forest feature selection. Reason: ranking bias towards continuous and high-cardinality categorical features

Retrofitting U.S. Suburbia in the Era of Shared Autonomous Vehicles · Georgia Tech

Tried and failed

random forest mean decrease in impurity applied to feature importance estimation. Reason: biased toward continuous variables and features with many categories

Automated methods to find high-redshift quasars · Imperial

Feature selection assumptions fail under context dependence and distribution shifts

6 theses · 6 institutions

Selected feature subsets frequently failed to transfer when external factors masked risk components or when sequences varied in length. Assuming context independence or invariant feature relationships caused models to degrade severely across different experimental backgrounds.

Tried and failed

ANOVA feature selection applied to multivariate time series classification. Outcome: did not generalise. Reason: failed to yield stable or high performance across varying sequence lengths

Temporal learning techniques for vehicle cybersecurity · Texas Tech

Tried and failed

cross-feature-representation training and testing applied to linear classification of trial-level signals. Outcome: did not generalise. Reason: training on regularized mixed-effects features did not generalize to standard least-squares estimates at test time

Extracting Feature Vectors From Event-Related fMRI Data to Enable Machine Learning Analysis · Virginia Tech

Tried and failed

maximal correlation and mutual information regularization applied to domain generalization under spurious correlation. Outcome: did not generalise. Reason: regularization suppressed spurious features but failed to learn fully invariant representations, yielding random chance accuracy

Maximal Correlation Feature Selection and Suppression With Applications · MIT

Tried and failed

two-phase linear regression with stepwise feature selection applied to bond investment risk premium prediction. Outcome: did not generalise. Reason: external credit guarantees masked idiosyncratic risk factors, drastically reducing model explanatory power

Financial risk : charter school-specific risk factors and idiosyncratic risk · UT Austin

Tried and failed

gradient-ordered feature selection assuming context independence applied to genetic locus fitness across environments. Outcome: did not generalise. Reason: features exhibited strong background dependence, violating the assumption of hub independence across backgrounds

The structure of fitness landscapes across genotypes and environments · Harvard

Considered and rejected

Considered and rejected: Rejected global feature selection methods (Decision Trees, Random Forests, and L1-regularized Logistic Regression) because they do not capture instance-specific variability and lack runtime adaptability.

An MQTT Application-Layer Traffic Analyzer for Interpretable Flow-Level Intrusion Detection and Zero-DayThreat Identification in IoT Environment Using TabNet · YorkSpace

Left open by the authors

Problems the authors named and did not get to.

Left open

Perform systematic feature selection and reduction using techniques like RFE and PCA on the 17 identified search and instance features. Blocker: None

Automated design of population-based algorithms: a case study in vehicle routing · University of Nottingham Repository

Left open

Benchmark co-segregation-based Bayesian linear regression feature selection against standard methods like PCA and random forests for microbial phenotypic GWAS. Blocker: None

Probing Genomic, Metabolic, And Phenotypic Evolution In Microbes Using Comparative And Experimental Evolution Data · Georgia Tech

Left open

Implement alternative feature scoring metrics like PCA loadings, Katz centrality, and PageRank into the HARVEST microbiome feature selection pipeline. Blocker: None

A Modular Data Analytic Pipeline for Feature Selection in High Dimensional Microbial Data Sets · HARVEST

Left open

Extract additional meta-features and evaluate a larger, more diverse pool of feature selection methods within the meta-learning framework. Blocker: No specific meta-features or feature selection algorithms are specified to evaluate

An investigation of fuzzy methods and meta learning for feature selection · University of Nottingham Repository

Left open

Identify significant predictors of composite stress distributions across loading stages using feature extraction techniques on the U-Net model. Blocker: The goal is too vague regarding specific feature extraction methods and target metrics

A Deep Learning Approach to Predict Full-Field Stress Distribution in Composite Materials · Virginia Tech

Left open

Perform feature selection and incorporate hierarchical relationships between subcellular locations into CFMS protein localization classifiers. Blocker: The proposals (improve input quality, feature selection, location relationships) lack specific implementation details or defined mathematical approaches

Predicting Subcellular Locations of Conserved Eukaryotic Protein Families with Co-Fractionation Mass Spectrometry Data · UT Austin

Left open

Perform feature selection to reduce the 17 longitudinal EHR features to a minimal informative subset for delirium classification. Blocker: Requires the private clinical ICU EHR dataset used in the thesis.

Automatic delirium classification in intensive care using non-invasive eye tracking · Imperial

Left open

Evaluate the reduced feature subsets produced by the fuzzy feature selection methods using classifiers like SVM and Random Forest. Blocker: None

An investigation of fuzzy methods and meta learning for feature selection · University of Nottingham Repository

Left open

Reverse engineer features learned by weak RFML learners to identify causes of recurring misclassifications without training full ensembles. Blocker: None

Sensitivity Analysis of RFML-based SEI Algorithms · Virginia Tech

Left open

Evaluate alternative feature selection methods within the class-imbalance classification pipeline to optimize feature subsets. Blocker: None

Machine Learning Approaches for Healthcare Analysis · Scholarship at UWindsor Institutional Repository

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.