Chapter Four · failure evidence
What Data Augmentation got wrong, from 94 dissertations
Across diverse machine learning domains, data augmentation strategies frequently degrade model performance or fail to improve upon unaugmented baselines. Major failure modes include corrupting domain specific semantics, generating low quality synthetic samples, inducing overfitting, and introducing harmful noise perturbations. These records come from PhD theses at 28 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Transformations corrupt domain semantics, temporal structures, and ground truth labels
Augmentations such as cropping, swapping, mixing, and temporal shifting frequently alter underlying class semantics and destroy critical structural signals. Across modalities including audio, text, time series, and medical imaging, these aggressive perturbations create invalid ground truth associations that degrade downstream accuracy.
Tried and failed
frame-independent data augmentation applied to video representation learning. Outcome: worse than baseline. Reason: disrupts temporal coherence across video frames, degrading action recognition performance
Action Recognition with Knowledge Transfer · Virginia Tech
Tried and failed
rule-based text data augmentation applied to intent classification. Outcome: worse than baseline. Reason: heuristic word-level perturbations alter sentence semantics and label alignment
Lifelong Machine Learning with Data Efficiency and Knowledge Retention · EPFL
Tried and failed
Naive semantics-preserving data augmentation applied to code vulnerability classification models. Outcome: worse than baseline. Reason: Uncurated uniform transformations introduced noise rather than meaningful invariant training signals.
Refactoring programs to improve the performance of deep learning for vulnerability detection · Iowa State
Considered and rejected
Considered and rejected: Rejected standard computer vision data augmentations (random cropping, horizontal flips, rotations) as they distort frequency and temporal representations in spectrograms.
Applications of Deep Convolutional Neural Networks to Passive Acoustic Monitoring of Baleen Whales · DalSpace
Tried and failed
mixup data augmentation applied to medical image segmentation across population shifts. Outcome: worse than baseline. Reason: linear interpolation created unrealistic images that degraded generalization across pathological groups
Improving the domain generalization and robustness of neural networks for medical imaging · Imperial
Tried and failed
data augmentation in contrastive learning applied to visual similarity representation learning. Outcome: did not generalise. Reason: augmentation corrupted fine-grained similarity signals needed for zero-shot and fine attribute discrimination
Considered and rejected
Considered and rejected: Rejected standard data augmentation (rotations, flips, color jitter) for vineyard disease detection because disease symptoms affect radiometric response without altering plant morphology.
Service robotics and machine learning for close-range remote sensing · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected standard pix2pix data augmentations because resizing/rotating narrow bacterial cells produced discontinuities and invalid ground-truth labels.
Bacterial Deepfakes: Generating Synthetic Microscopy Data to Improve Adaptability of Deep Learning-Based Segmentation Models · ResearchWorks
Considered and rejected
Considered and rejected: Discarded very small crops in data augmentation for UNO because cropping occludes critical image information and ruins pseudo-label quality.
Knowledge transfer and retention in deep neural networks · IRIS - UNITN - prod
Considered and rejected
Considered and rejected: Rejected standard image rotations/flips for IMC data augmentation because pixel intensity distributions within segmented cells remain invariant
A Computational Analysis Pipeline for Imaging Mass Cytometry Data for Cancer Research · Queens University Institutional Repository
Tried and failed
perturbation-based contrastive data augmentation applied to relational graph triples. Reason: standard data augmentations alter discrete semantic information in relational triples
Tried and failed
aggressive data augmentations in contrastive learning applied to image and video quality assessment. Reason: augmentations alter or destroy distortion information that defines perceptual quality labels
Learning variable frame rate and unsupervised video quality assessment · UT Austin
Tried and failed
synonym substitution and duplication data augmentation applied to relation classification. Outcome: did not generalise. Reason: None
Considered and rejected
Considered and rejected: Rejected Random Swap (RS) and Random Deletion (RD) data augmentation techniques because literature showed they fail to preserve dataset class labels after augmentation.
DATA MINING AND RE-IDENTIFICATION: ANALYSIS OF DATABASE QUERY PATTERNS THAT POSE A THREAT TO ANONYMISED INFORMATION · De Montfort Open Research Archive (DORA)
Considered and rejected
Considered and rejected: Rejected applying synthetic data expansion/augmentation on raw STTF sensor data to preserve actual recorded sensor characteristics.
From ODD Definition to Deployment: Weather-Resilient Perception in Autonomous Driving via Sensor Benchmarking, Adaptation, and Fusion · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Rejected applying data augmentation to PLC shape-based curve-fitting methods because it restricts degrees of freedom on actual data points
Optimizing Sales Forecasting, Inventory, Pricing and Sourcing Decisions · EPFL
Considered and rejected
Considered and rejected: Rejected shear and strain data augmentation on X-ray datasets because distortive transformations generate unrealistic synthetic fractures
Deep Learning and Augmented Reality for 3D human-machine interaction · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected semantic-relation augmentation (synonyms/antonyms) because lack of learner consensus, cross-association difficulty, formality variance, and need for human guidance negate automation.
Tried and failed
generic time-series data augmentation applied to intertwined multivariate time-series data. Outcome: did not generalise. Reason: altered the underlying semantic meaning of the intertwined multivariate dynamics
Exploring dispersion dynamics in agitated mixers via numerical simulations and machine learning · Imperial
Tried and failed
random time-translation data augmentation applied to time-series neural posterior estimation. Outcome: worse than baseline. Reason: enforcing time-invariance degraded parameter posterior estimation precision compared to fixing the feature alignment
Gravitational Waveform Modelling with Machine Learning and for Eccentric Binary Systems · Cornell
Tried and failed
Random temporal shift data augmentation applied to 1D CNN time-series regression. Outcome: worse than baseline. Reason: Random signal shifts degraded prediction accuracy instead of improving shift invariance.
Tried and failed
Cutout and intensity shifting data augmentations applied to self-supervised pretraining for 3D object detection. Outcome: worse than baseline. Reason: Augmentations destroyed inherent object information critical for downstream recognition
3D deep learning threat detection for real-time computed tomography baggage screening · Imperial
Considered and rejected
Considered and rejected: Rejected image augmentation using compressing and stretching because altering the height/width ratio distorted air-void circular morphology and confused them with noise.
Three-dimensional Segmentation of Air-void System in Hardened Concrete using Photometric Stereo and Artificial Intelligence Methods · TXST Digital Repository
Complex and adaptive augmentation techniques underperform simpler baselines or unaugmented models
Sophisticated augmentation policies, adaptive expansions, and automated transformation pipelines often achieve worse accuracy and loss than training without any augmentation. Simpler heuristic transformations or clean unaugmented datasets consistently outperform these complex methods across text, vision, and tabular benchmarks.
Lost to a baseline
Duplication and Entity Replacement data augmentation baselines produced higher test RMSE than original unaugmented training across all tensor datasets
Accurate and Trustworthy Recommender Systems: Algorithms and Findings · Georgia Tech
Lost to a baseline
x-vector fine-tuned on ADReSSo2021 with data augmentation achieved 0.6862 accuracy, performing worse than the un-augmented fine-tuned baseline (0.7163).
LEARNING UTTERANCE LEVEL REPRESENTATION FROM SPEECH · JScholarship
Lost to a baseline
On ImageNet test-time augmentation (10 samples), random crop (79.60%), AutoAugment (79.20%), and FastAutoAugment (79.28%) all degraded ResNet-50 performance compared to no augmentation (80.43%)
Incorporating inductive biases into machine learning algorithms · Oxford
Lost to a baseline
Baseline StarGAN-EVC achieved higher Macro-F1 (56.64% with no augmentation vs 55.41% augmented) in SER data augmentation experiments.
Enhancing speech intelligibility through paralinguistic features · Imperial
Lost to a baseline
On RawFooT D45 texture classification, AdaAug (75.27%) and Augerino (78.97%) achieved lower test accuracy than Random Augmentation (79.99%)
Incorporating inductive biases into machine learning algorithms · Oxford
Considered and rejected
Considered and rejected: Rejected class-specific augmentations conditioned only on labels because label-dependent transforms create a mismatch between training and testing that degrades performance below unaugmented or input-agnostic variants
Incorporating inductive biases into machine learning algorithms · Oxford
Considered and rejected
Considered and rejected: Sequence-level paraphrastic data augmentation via beam search or sampled paraphrasing was rejected in favor of greedy-search paraphrasing for the data-augmentation baseline.
Overcoming Data Challenges in Machine Translation · JScholarship
Considered and rejected
Considered and rejected: Rejected naive data augmentation for compositional generalization due to arbitrary heuristics, tuning overhead, and domain specificity.
Compositional Robot Learning for Generalizable Interactions · MIT
Lost to a baseline
Small rotation data augmentation (86.0% accuracy) was beaten by the no-augmentation baseline (86.7% accuracy).
IMAGE QUALITY ASSESSMENT OF ACTIVE SONAR IMAGES THROUGH BAYESIAN DEEP LEARNING · Calhoun
Considered and rejected
Considered and rejected: Decided against complex data augmentations (e.g., MixUp, CutMix, color jitter) for medical segmentation in favor of simple random rotations and flips.
Adaptive and weighted optimization for efficient and robust learning · UT Austin
Considered and rejected
Considered and rejected: Rejected hard augmentations (RandAug) for image consistency regularization in CoPrompt, as it caused severe feature divergence and lower accuracy (79.90%) vs simple augmentations (80.48%).
Representation Learning under Limited Supervision · Queens University Institutional Repository
Lost to a baseline
On Iris UCI dataset, fixed expansion (r=0.1) achieved lower adaptive robust loss (0.0783) than adaptive augmentation (0.0870).
Novel Examination of Interpretable Surrogates and Adversarial Robustness in Machine Learning · YorkSpace
Lost to a baseline
On VLCS, style-transfer data augmentation failed to improve upon the unstylized baseline (72.31% vs 72.49% average accuracy).
Addressing Distributional Shift challenges in Computer Vision for Real-World Applications · IRIS - POLITO - prod
Lost to a baseline
Default Augmentation baseline B (78.8% accuracy) beat MixUp (74.6%), CutMix (76.8%), and RandAugment (76.9%) on Kvasir dataset.
Training Strategy for Limited Labeled Data by Learning from Confusion · Iowa State
Lost to a baseline
On Heart Disease UCI dataset, original unaugmented model achieved lower adaptive robust loss (0.3465) than adaptive augmentation (0.3604).
Novel Examination of Interpretable Surrogates and Adversarial Robustness in Machine Learning · YorkSpace
Lost to a baseline
No Adaptation baseline achieved higher recall on data1 for ResNet (0.969 vs 0.918/0.928) and EfficientNet (0.962 vs 0.859) compared to geometric and DCGAN augmentations.
Learning Effectively from Medical Imaging Datasets · HARVEST
Tried and failed
individual data augmentations during policy training applied to robot imitation learning policies. Outcome: worse than baseline. Reason: individual visual or proprioceptive perturbations severely degraded in-distribution task performance compared to unaugmented baseline
A High Performance Robotics Data Augmentation Framework · Georgia Tech
Lost to a baseline
On Parkinsons UCI dataset, fixed expansion (r=0.1) achieved lower adaptive robust loss (0.1542) than adaptive augmentation (0.1627).
Novel Examination of Interpretable Surrogates and Adversarial Robustness in Machine Learning · YorkSpace
Generative and synthetic data augmentations introduce distributional shift and unrealistic artifacts
Synthetic samples produced by language models, generative adversarial networks, diffusion models, and speech synthesizers often introduce grammatical errors, domain shifts, and unrealistic features. As a result, synthetic samples fail to generalize to unseen test cohorts and fall behind models trained purely on real data.
Tried and failed
synthetic generative data augmentation applied to atypical speech recognition. Outcome: worse than baseline. Reason: synthetic audio failed to outperform simple addition of unperturbed out-of-domain control data
On matching data and model in LF-MMI-based dysarthric speech recognition · EPFL
Tried and failed
synthetic data augmentation using repurposed conditional embeddings applied to automatic speech recognition across accents. Outcome: worse than baseline. Reason: synthetic multi-accent speech generation did not improve downstream model performance over real data training
Voice conversion and text-to-speech for privacy protection applications · Imperial
Tried and failed
generative text data augmentation applied to text quantification training datasets. Outcome: worse than baseline. Reason: generated ungrammatical, nonsensical text with misspellings, degrading training data quality
Quantification learning with deep neural networks · Iowa State
Tried and failed
standard GAN data augmentation applied to photoplethysmography time-series signals. Outcome: worse than baseline. Reason: generated synthetic samples provided limited performance gain over simple baseline augmentation techniques
Toward accurate health monitoring through large-scale Photoplethysmography signal from wearable devices · Georgia Tech
Tried and failed
heavy synthetic data augmentation during pre-training without adaptation applied to semantic segmentation pre-training. Outcome: worse than baseline. Reason: severe domain shift and distortion introduced by augmentations without calibration layers hurt downstream performance
Label-Efficient Visual Understanding with Consistency Constraints · Virginia Tech
Lost to a baseline
ASR models fine-tuned with multi-accent TTS augmentation (FTaug: 7.88% WER, FMaug: 6.79% WER) were beaten by models fine-tuned on real data alone (FT: 7.66% WER, FM: 6.33% WER)
Voice conversion and text-to-speech for privacy protection applications · Imperial
Considered and rejected
Considered and rejected: Rejected Seq2Seq, GAN, and generative LM models (e.g. BART, PPLM) for text augmentation because generated text contained grammatical errors, misspellings, and lacked meaning.
Quantification learning with deep neural networks · Iowa State
Considered and rejected
Considered and rejected: Rejected adding instances via generative data augmentation to balance data diversity, due to introducing confounds from machine-generated text differing from human text.
Experiment Design for Hypotheses About How NLP Models Work · ResearchWorks
Considered and rejected
Considered and rejected: Rejected standard GANs for data augmentation because synthetic data may not accurately represent the true data distribution
A Novel Lightweight Convolutional Neural Network for Medical Image Classification and Segmentation · HARVEST
Tried and failed
generative neural network for data augmentation applied to imbalanced 3D medical image segmentation. Outcome: did not generalise. Reason: generated augmentations caused heavy overfitting to validation data and failed to generalize to unseen test data
Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial
Tried and failed
semi-supervised learning with synthetic blend augmentations applied to medical image outlier and class classification. Outcome: did not generalise. Reason: improves distinct class accuracy but increases confusion between highly similar or fine-grained classes
Machine learning for outlier detection in medical imaging · Imperial
Tried and failed
generative data augmentation for image classification applied to histopathology whole slide image classification. Outcome: did not generalise. Reason: boosted internal validation metrics but failed to improve or degraded performance on external cohort
Advancing Personalized Medicine Through Generative Artificial Intelligence · Georgia Tech
Tried and failed
diffusion-based synthetic data augmentation for classification applied to medical image classification. Outcome: worse than baseline. Reason: accuracy gains saturated and diminished as real training sample sizes increased
Augmenting medical image classifiers with synthetic data across populations · Harvard
Tried and failed
heuristic and back-translation data augmentation for questions applied to machine reading comprehension. Outcome: worse than baseline. Reason: generated synthetic question variations did not provide useful training signal to improve reading comprehension accuracy
Machine Reading Comprehension: Challenges and Approaches · Cornell
Tried and failed
paraphrasing and surface copy for data augmentation applied to style-constrained text dataset augmentation. Reason: generated text failed to match target domain style and stylistic conventions
Lost to a baseline
SemanticGAN-augmented models achieved lower pick-and-place accuracy (0.32 AC, 0.18 AG) than the baseline trained on 100 real samples with no augmentation (0.44 AC, 0.29 AG).
Data-Efficient Learning Frameworks for Adaptive Intelligent Robots in Human-Robot Collaboration Scenarios · IRIS - POLITO - prod
Tried and failed
pretraining with synthetic trajectory data augmentation applied to visuomotor policy learning. Outcome: worse than baseline. Reason: variable quality in synthetic trajectories degraded downstream performance compared to human-only data
Scaling robot learning with heterogeneous data from the real world, simulation, and the web · UT Austin
Data augmentation induces overfitting, feature redundancy, and demographic bias
Multiplying training data volume or repetitively applying fine-grained perturbations leads models to overfit to transformation artifacts and mask geometries. In addition, these expanded datasets can exacerbate subgroup disparities and amplify pre-existing dataset biases.
Tried and failed
asymmetric data augmentation across dataset streams applied to incremental learning with reference data. Outcome: worse than baseline. Reason: model exploited augmentation artifacts present only in the training stream, degrading representation quality
Referencing Unlabelled World Data to Prevent Catastrophic Forgetting in Class-incremental Learning · Virginia Tech
Tried and failed
heavy multi-sample data augmentation applied to certified robustness via randomized smoothing. Outcome: worse than baseline. Reason: increased class-wise accuracy disparity and degraded overall certified accuracy compared to standard single-sample augmentation
Understanding and Improving Representational Robustness of Machine Learning Models · MIT
Considered and rejected
Considered and rejected: Rejected data augmentation for industrial anomaly detection due to risks of overfitting and need for complex augmentation strategies.
Neuro-Symbolic Integration in Artificial Intelligence and its Applications · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected using a neural network parameterization for data augmentation transformations because it easily overfit validation proxies.
Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial
Tried and failed
naive data duplication for dataset augmentation applied to surgical scene segmentation. Outcome: overfit. Reason: overfitting occurred without providing added diversity to the dataset
Image synthesis with class-aware semantic diffusion models for surgical scene segmentation · Imperial
Tried and failed
standard random color jitter data augmentation applied to semantic image segmentation. Outcome: did not generalise. Reason: random perturbations increased subgroup demographic bias and induced excessive false positive predictions
Color Invariant Skin Segmentation · Virginia Tech
Considered and rejected
Considered and rejected: Decided against fixed-edge mask data augmentation during training because it overfit to mask geometry and failed to generalize as well to localized vector orientations.
Dimensionality reduction for validating engine flow simulations · Oxford
Considered and rejected
Considered and rejected: Rejected copy-pasting small object data augmentation due to high likelihood of overfitting from lack of background blending
Tracking and Measuring Objects in Obscure Image Scenarios Through the Lens of Shot Put in Track and Field · Virginia Tech
Considered and rejected
Considered and rejected: Rejected relying solely on post-hoc stain normalization and data augmentation for multi-site deployment, because models still learned residual site-specific signatures
Privacy-Preserving Federated Learning for Secure and Scalable Digital Pathology · DSpace at SUNY Buffalo
Tried and failed
noisy text augmentation for cross-validation applied to transformer language models on small datasets. Outcome: overfit. Reason: augmentations lacked sufficient diversity and heightened model sensitivity to noise, worsening overfitting on small splits
Tried and failed
data augmentation to mitigate bias applied to tabular recidivism prediction datasets. Reason: the dataset contained fundamental label bias and historical prejudice rather than simple sample representation imbalance
Multi-objective approaches towards trustworthy machine learning · UT Austin
Considered and rejected
Considered and rejected: Rejected synthetic data generation and data augmentation because small dataset size could lead to overfitting, bias amplification, and false validity.
Methods for Classifying Driver Engagement in Autonomous Vehicles Using Physiological Sensors · Carleton University Institutional Repository
Tried and failed
extending training duration with standard data augmentation applied to image classification models. Outcome: overfit. Reason: training baseline models for longer schedules caused overfitting instead of performance gains
Unified confusion-derived learning framework for image classification · Iowa State
Tried and failed
training dataset expansion via data augmentation applied to object detection model training. Outcome: overfit. Reason: tripling the dataset using data augmenters caused model overfitting, worsening error rates
Tried and failed
adversarial sample augmentation during iterative training applied to reward model training for alignment. Outcome: overfit. Reason: low adversarial data diversity caused the model to overfit when adding too many generated samples
Robust and Flexible Reward Modeling for LLM Alignment · Georgia Tech
Tried and failed
fine-grained rotational data augmentation applied to image classification models. Outcome: overfit. Reason: excessive rotation steps introduced data redundancy that degraded generalization performance
Development of a method to classify and analyse the composition of mixed waste materials in real-time · Cranfield
Improper pipeline placement, multi-stage compounding, and redundant augmentations yield diminishing returns
Sequentially cascading distinct augmentations, applying augmentations continuously throughout entire training schedules, or inserting them into evaluation steps degrades representation learning. Combining multiple transformations or augmenting already well-represented classes provides minimal benefit over individual techniques or single-stage training.
Tried and failed
combining heuristic augmentations with synthetic data applied to accented speech recognition. Outcome: worse than baseline. Reason: None
Voice conversion and text-to-speech for privacy protection applications · Imperial
Tried and failed
sequential cascading of multiple distinct data augmentations applied to video action recognition training. Outcome: worse than baseline. Reason: compounding intra-clip and cross-clip augmentations degraded visual representations compared to stochastic single-augmentation selection
Action Recognition with Knowledge Transfer · Virginia Tech
Tried and failed
data augmentation during clustering optimization applied to deep image clustering network input. Outcome: worse than baseline. Reason: feeding augmented rather than raw samples directly to the clustering head degraded cluster assignment quality
Yet another image clustering framework using deep learning · Iowa State
Considered and rejected
Considered and rejected: Rejected using full signature augmentation (lead-lag + basepoint) universally without testing simpler models, as it degraded performance on the timing group compared to basepoint alone
Tried and failed
synthetic data augmentation for majority classes applied to semantic image segmentation. Outcome: no signal. Reason: sufficient real features were already present, yielding minimal to diminishing performance returns
Image synthesis with class-aware semantic diffusion models for surgical scene segmentation · Imperial
Tried and failed
combining multiple data augmentation techniques applied to robotic manipulation out-of-distribution generalization. Outcome: did not generalise. Reason: simple pick-and-place tasks limited the utility of augmentations beyond baseline generation
A High Performance Robotics Data Augmentation Framework · Georgia Tech
Tried and failed
unfiltered auxiliary training data augmentation applied to image classification model training. Outcome: did not generalise. Reason: distribution shift and bias in auxiliary data degraded downstream test performance
Probing, Improving, and Verifying Machine Learning Model Robustness · MIT
Tried and failed
data augmentation with poorly conditioned transformations applied to multilayer perceptron training. Outcome: did not generalise. Reason: higher condition numbers in transformed training data increased test error on clean datasets
Computational Tradeoffs and Symmetry in Polynomial Nonnegativity · MIT
Tried and failed
data augmentation during fine-tuning applied to chemical reaction condition prediction. Outcome: worse than baseline. Reason: None
Tried and failed
data augmentation across multiple sequential training stages applied to few-shot object detection training. Reason: applying augmentation in both stages offered no compounding benefit over single-stage application
Data- and compute-efficient visual recognition and generation · UT Austin
Tried and failed
combining multiple photometric data augmentations applied to low-contrast object detection. Reason: yielded no considerable performance improvement over individual augmentations alone
Object detection for low-contrast complex background applications · Iowa State
Tried and failed
synthetic detection noise for track data augmentation applied to multi-object tracking models. Outcome: no signal. Reason: random object dropout and false positive injection yielded only marginal performance gains
Autonomous Vehicle Perception Quality Assessment · Virginia Tech
Tried and failed
data augmentation on small training subsets applied to keypoint detection in deep pose estimation. Reason: augmentation provided minimal performance gains over non-augmented training when sample size was very small
Tried and failed
contrastive data augmentation and textured mesh rendering applied to vision-language navigation. Outcome: worse than baseline. Reason: augmentations and synthetic mesh renders yielded no performance gains on unseen environments
Towards multi-modal AI systems with open-world cognition · Georgia Tech
Tried and failed
continuous mosaic data augmentation throughout training applied to object detection model training. Outcome: worse than baseline. Reason: applying heavy mosaic augmentation for the entire duration degraded final model performance compared to disabling it late in training
Hierarchical transfer learning for small object detection · Iowa State
Considered and rejected
Considered and rejected: Decided against data augmentation during the AIL reward inference evaluation step, as empirically it decreased performance.
On the use of expert data to imitate behavior and accelerate Reinforcement Learning · OpenBU
Noise and blur perturbations corrupt representations and degrade model robustness
Injecting synthetic Gaussian noise, blur, contrast perturbations, and spectral noise fails to improve environmental invariance and degrades task precision across multiple architectures. In addition, training with artificial noise distortions degrades adversarial robustness and increases vulnerability to out-of-distribution artifacts.
Tried and failed
additive white Gaussian noise data augmentation applied to audio speaker diarization. Outcome: worse than baseline. Reason: models showed poor noise robustness and performance degraded across all architectures
Speaker diarization: importance of the modulation spectrum and incorporating uncertainty modelling · Imperial
Tried and failed
data augmentation with noise and reverberation applied to speaker verification x-vector models. Outcome: worse than baseline. Reason: external noise and reverberation augmentation slightly increased equal error rate
Towards Automatic Analysis of Audio Recordings from Children with Autism Spectrum Disorder · Georgia Tech
Tried and failed
synthetic noise data augmentation during pre-training applied to aerial visual object detection models. Outcome: worse than baseline. Reason: degraded detection precision across multiple evaluation categories instead of improving robustness
Tried and failed
Gaussian noise and blur data augmentation applied to generative image and video models. Outcome: worse than baseline. Reason: Degraded generation and reconstruction performance instead of regularizing the models
Probabilistic learning and generation in deep sequence models · Imperial
Tried and failed
Gaussian noise data augmentation applied to adversarial robustness in deep learning. Outcome: worse than baseline. Reason: training on Gaussian-augmented data degraded robustness against adversarial perturbations
Robust Efficient Edge AI: New Principles and Frameworks for Empowering Artificial Intelligence on Edge Devices · Georgia Tech
Tried and failed
adversarial bias field augmentation applied to medical image segmentation domain generalization. Outcome: did not generalise. Reason: increased vulnerability to out-of-distribution spike noise artifacts, degrading segmentation performance
Improving the domain generalization and robustness of neural networks for medical imaging · Imperial
Tried and failed
Gaussian blur data augmentation for defocus robustness applied to microscopy image segmentation. Outcome: did not generalise. Reason: Synthetic Gaussian blur fails to capture real optical defocus characteristics and provides no benefit over unaugmented data
Single-cell Methods and Spatial Analysis for Highly Multiplexed Tissue Images · Harvard
Tried and failed
Random contrast data augmentation applied to low-contrast object detection. Outcome: worse than baseline. Reason: generated uninformative or misleading features that acted as negative training examples
Object detection for low-contrast complex background applications · Iowa State
Tried and failed
spectral noise data augmentation applied to hyperspectral image semantic segmentation. Outcome: worse than baseline. Reason: None
Hyperspectral Remote Sensing for UXO Detection and Damage Assessment on Airfield Pavements · MIT
Considered and rejected
Considered and rejected: Rejected random rotation and random cropping data augmentation because zero-padded image corners produced all-zero artifact patches.
A Unified Framework for Advancing Soil Erosion and Flood Assessment Through Deep Learning and Process-Based Modeling · Publikationssystem UB Tuebingen
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop multimodal data augmentation methods bridging text and audio by pairing text augmentations with text-to-speech models for online speech processing. Blocker: None
Left open
Optimize RAG retrieval strategies and data augmentation techniques to enhance fidelity and diversity for sensor-text foundation models. Blocker: Requires the private in-home monitoring sensor dataset and pipeline from the thesis
Developing a foundation model in in-home monitoring data for healthcare applications · Imperial
Left open
Evaluate broader acoustic features and apply noise reduction and data augmentation techniques to improve speech inspiration detection. Blocker: None
Automatic detection of speech inspiration using SVM and decision trees · Cambridge
Left open
Implement data augmentation for acoustic localization by randomly adding or deleting image sources and injecting 30-50 dB SNR noise into RIRs. Blocker: None
Left open
Evaluate text classification augmentation techniques including length-restricted synonym replacement, random word insertion, function word deletion, and single-letter modifications. Blocker: None
Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard
Left open
Evaluate length-based text data augmentation across LSTM and recurrent convolutional neural network architectures on text classification benchmarks. Blocker: None
Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard
Left open
Implement data augmentation on cross-site pretraining datasets in TransEHR to reduce the performance gap between full-model fine-tuning and last-layer MLP tuning. Blocker: Requires multi-site electronic health record datasets (like MIMIC or eICU/private clinical data with credentialing or restrictions).
Robust Representation Learning and Real-Time Serving of Deep Models for Health Time Series · Georgia Tech
Left open
Evaluate the impact of different data augmentation levels and techniques on the informal medical entity recognition (NER) model performance. Blocker: None
Supporting laypeople in learning formal medical terminology · DSpace-CRIS at TU Wien
Left open
Implement audio data augmentations and active learning strategies to improve sample balance and accuracy in BirdNET-based transfer learning. Blocker: None
Left open
Combine length-based data augmentation with standard EDA techniques and evaluate text classification performance on benchmark datasets. Blocker: None
Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.