Chapter Four · failure evidence

What Data Augmentation got wrong, from 94 dissertations

Across diverse machine learning domains, data augmentation strategies frequently degrade model performance or fail to improve upon unaugmented baselines. Major failure modes include corrupting domain specific semantics, generating low quality synthetic samples, inducing overfitting, and introducing harmful noise perturbations. These records come from PhD theses at 28 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Transformations corrupt domain semantics, temporal structures, and ground truth labels

23 theses · 15 institutions

Augmentations such as cropping, swapping, mixing, and temporal shifting frequently alter underlying class semantics and destroy critical structural signals. Across modalities including audio, text, time series, and medical imaging, these aggressive perturbations create invalid ground truth associations that degrade downstream accuracy.

Tried and failed

frame-independent data augmentation applied to video representation learning. Outcome: worse than baseline. Reason: disrupts temporal coherence across video frames, degrading action recognition performance

Action Recognition with Knowledge Transfer · Virginia Tech

Tried and failed

rule-based text data augmentation applied to intent classification. Outcome: worse than baseline. Reason: heuristic word-level perturbations alter sentence semantics and label alignment

Lifelong Machine Learning with Data Efficiency and Knowledge Retention · EPFL

Tried and failed

Naive semantics-preserving data augmentation applied to code vulnerability classification models. Outcome: worse than baseline. Reason: Uncurated uniform transformations introduced noise rather than meaningful invariant training signals.

Refactoring programs to improve the performance of deep learning for vulnerability detection · Iowa State

Considered and rejected

Considered and rejected: Rejected standard computer vision data augmentations (random cropping, horizontal flips, rotations) as they distort frequency and temporal representations in spectrograms.

Applications of Deep Convolutional Neural Networks to Passive Acoustic Monitoring of Baleen Whales · DalSpace

Tried and failed

mixup data augmentation applied to medical image segmentation across population shifts. Outcome: worse than baseline. Reason: linear interpolation created unrealistic images that degraded generalization across pathological groups

Improving the domain generalization and robustness of neural networks for medical imaging · Imperial

Tried and failed

data augmentation in contrastive learning applied to visual similarity representation learning. Outcome: did not generalise. Reason: augmentation corrupted fine-grained similarity signals needed for zero-shot and fine attribute discrimination

Advancements in perceptual quality assessment for interactive media : from mobile cloud gaming to human avatar videos and facial expressions · UT Austin

Considered and rejected

Considered and rejected: Rejected standard data augmentation (rotations, flips, color jitter) for vineyard disease detection because disease symptoms affect radiometric response without altering plant morphology.

Service robotics and machine learning for close-range remote sensing · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Rejected standard pix2pix data augmentations because resizing/rotating narrow bacterial cells produced discontinuities and invalid ground-truth labels.

Bacterial Deepfakes: Generating Synthetic Microscopy Data to Improve Adaptability of Deep Learning-Based Segmentation Models · ResearchWorks

Considered and rejected

Considered and rejected: Discarded very small crops in data augmentation for UNO because cropping occludes critical image information and ruins pseudo-label quality.

Knowledge transfer and retention in deep neural networks · IRIS - UNITN - prod

Considered and rejected

Considered and rejected: Rejected standard image rotations/flips for IMC data augmentation because pixel intensity distributions within segmented cells remain invariant

A Computational Analysis Pipeline for Imaging Mass Cytometry Data for Cancer Research · Queens University Institutional Repository

Tried and failed

perturbation-based contrastive data augmentation applied to relational graph triples. Reason: standard data augmentations alter discrete semantic information in relational triples

Integrating structural and semantic understanding for robust knowledge graph construction: From knowledge graph completion to zero-shot entity linking · Iowa State

Tried and failed

aggressive data augmentations in contrastive learning applied to image and video quality assessment. Reason: augmentations alter or destroy distortion information that defines perceptual quality labels

Learning variable frame rate and unsupervised video quality assessment · UT Austin

Tried and failed

synonym substitution and duplication data augmentation applied to relation classification. Outcome: did not generalise. Reason: None

Building small domain-specific masked language models vs. large generative models for clinical decision support and their effects on users. · MIT

Considered and rejected

Considered and rejected: Rejected Random Swap (RS) and Random Deletion (RD) data augmentation techniques because literature showed they fail to preserve dataset class labels after augmentation.

DATA MINING AND RE-IDENTIFICATION: ANALYSIS OF DATABASE QUERY PATTERNS THAT POSE A THREAT TO ANONYMISED INFORMATION · De Montfort Open Research Archive (DORA)

Considered and rejected

Considered and rejected: Rejected applying synthetic data expansion/augmentation on raw STTF sensor data to preserve actual recorded sensor characteristics.

From ODD Definition to Deployment: Weather-Resilient Perception in Autonomous Driving via Sensor Benchmarking, Adaptation, and Fusion · Carleton University Institutional Repository

Considered and rejected

Considered and rejected: Rejected applying data augmentation to PLC shape-based curve-fitting methods because it restricts degrees of freedom on actual data points

Optimizing Sales Forecasting, Inventory, Pricing and Sourcing Decisions · EPFL

Considered and rejected

Considered and rejected: Rejected shear and strain data augmentation on X-ray datasets because distortive transformations generate unrealistic synthetic fractures

Deep Learning and Augmented Reality for 3D human-machine interaction · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Rejected semantic-relation augmentation (synonyms/antonyms) because lack of learner consensus, cross-association difficulty, formality variance, and need for human guidance negate automation.

Automatic enrichment of word lists with morphological derivatives for computer-assisted learning of vocabulary: Design-based research · Iowa State

Tried and failed

generic time-series data augmentation applied to intertwined multivariate time-series data. Outcome: did not generalise. Reason: altered the underlying semantic meaning of the intertwined multivariate dynamics

Exploring dispersion dynamics in agitated mixers via numerical simulations and machine learning · Imperial

Tried and failed

random time-translation data augmentation applied to time-series neural posterior estimation. Outcome: worse than baseline. Reason: enforcing time-invariance degraded parameter posterior estimation precision compared to fixing the feature alignment

Gravitational Waveform Modelling with Machine Learning and for Eccentric Binary Systems · Cornell

Tried and failed

Random temporal shift data augmentation applied to 1D CNN time-series regression. Outcome: worse than baseline. Reason: Random signal shifts degraded prediction accuracy instead of improving shift invariance.

Virtual metrology applied to milling process · EPFL

Tried and failed

Cutout and intensity shifting data augmentations applied to self-supervised pretraining for 3D object detection. Outcome: worse than baseline. Reason: Augmentations destroyed inherent object information critical for downstream recognition

3D deep learning threat detection for real-time computed tomography baggage screening · Imperial

Considered and rejected

Considered and rejected: Rejected image augmentation using compressing and stretching because altering the height/width ratio distorted air-void circular morphology and confused them with noise.

Three-dimensional Segmentation of Air-void System in Hardened Concrete using Photometric Stereo and Artificial Intelligence Methods · TXST Digital Repository

Complex and adaptive augmentation techniques underperform simpler baselines or unaugmented models

14 theses · 12 institutions

Sophisticated augmentation policies, adaptive expansions, and automated transformation pipelines often achieve worse accuracy and loss than training without any augmentation. Simpler heuristic transformations or clean unaugmented datasets consistently outperform these complex methods across text, vision, and tabular benchmarks.

Lost to a baseline

Duplication and Entity Replacement data augmentation baselines produced higher test RMSE than original unaugmented training across all tensor datasets

Accurate and Trustworthy Recommender Systems: Algorithms and Findings · Georgia Tech

Lost to a baseline

x-vector fine-tuned on ADReSSo2021 with data augmentation achieved 0.6862 accuracy, performing worse than the un-augmented fine-tuned baseline (0.7163).

LEARNING UTTERANCE LEVEL REPRESENTATION FROM SPEECH · JScholarship

Lost to a baseline

On ImageNet test-time augmentation (10 samples), random crop (79.60%), AutoAugment (79.20%), and FastAutoAugment (79.28%) all degraded ResNet-50 performance compared to no augmentation (80.43%)

Incorporating inductive biases into machine learning algorithms · Oxford

Lost to a baseline

Baseline StarGAN-EVC achieved higher Macro-F1 (56.64% with no augmentation vs 55.41% augmented) in SER data augmentation experiments.

Enhancing speech intelligibility through paralinguistic features · Imperial

Lost to a baseline

On RawFooT D45 texture classification, AdaAug (75.27%) and Augerino (78.97%) achieved lower test accuracy than Random Augmentation (79.99%)

Incorporating inductive biases into machine learning algorithms · Oxford

Considered and rejected

Considered and rejected: Rejected class-specific augmentations conditioned only on labels because label-dependent transforms create a mismatch between training and testing that degrades performance below unaugmented or input-agnostic variants

Incorporating inductive biases into machine learning algorithms · Oxford

Considered and rejected

Considered and rejected: Sequence-level paraphrastic data augmentation via beam search or sampled paraphrasing was rejected in favor of greedy-search paraphrasing for the data-augmentation baseline.

Overcoming Data Challenges in Machine Translation · JScholarship

Considered and rejected

Considered and rejected: Rejected naive data augmentation for compositional generalization due to arbitrary heuristics, tuning overhead, and domain specificity.

Compositional Robot Learning for Generalizable Interactions · MIT

Lost to a baseline

Small rotation data augmentation (86.0% accuracy) was beaten by the no-augmentation baseline (86.7% accuracy).

IMAGE QUALITY ASSESSMENT OF ACTIVE SONAR IMAGES THROUGH BAYESIAN DEEP LEARNING · Calhoun

Considered and rejected

Considered and rejected: Decided against complex data augmentations (e.g., MixUp, CutMix, color jitter) for medical segmentation in favor of simple random rotations and flips.

Adaptive and weighted optimization for efficient and robust learning · UT Austin

Considered and rejected

Considered and rejected: Rejected hard augmentations (RandAug) for image consistency regularization in CoPrompt, as it caused severe feature divergence and lower accuracy (79.90%) vs simple augmentations (80.48%).

Representation Learning under Limited Supervision · Queens University Institutional Repository

Lost to a baseline

On Iris UCI dataset, fixed expansion (r=0.1) achieved lower adaptive robust loss (0.0783) than adaptive augmentation (0.0870).

Novel Examination of Interpretable Surrogates and Adversarial Robustness in Machine Learning · YorkSpace

Lost to a baseline

On VLCS, style-transfer data augmentation failed to improve upon the unstylized baseline (72.31% vs 72.49% average accuracy).

Addressing Distributional Shift challenges in Computer Vision for Real-World Applications · IRIS - POLITO - prod

Lost to a baseline

Default Augmentation baseline B (78.8% accuracy) beat MixUp (74.6%), CutMix (76.8%), and RandAugment (76.9%) on Kvasir dataset.

Training Strategy for Limited Labeled Data by Learning from Confusion · Iowa State

Lost to a baseline

On Heart Disease UCI dataset, original unaugmented model achieved lower adaptive robust loss (0.3465) than adaptive augmentation (0.3604).

Novel Examination of Interpretable Surrogates and Adversarial Robustness in Machine Learning · YorkSpace

Lost to a baseline

No Adaptation baseline achieved higher recall on data1 for ResNet (0.969 vs 0.918/0.928) and EfficientNet (0.962 vs 0.859) compared to geometric and DCGAN augmentations.

Learning Effectively from Medical Imaging Datasets · HARVEST

Tried and failed

individual data augmentations during policy training applied to robot imitation learning policies. Outcome: worse than baseline. Reason: individual visual or proprioceptive perturbations severely degraded in-distribution task performance compared to unaugmented baseline

A High Performance Robotics Data Augmentation Framework · Georgia Tech

Lost to a baseline

On Parkinsons UCI dataset, fixed expansion (r=0.1) achieved lower adaptive robust loss (0.1542) than adaptive augmentation (0.1627).

Novel Examination of Interpretable Surrogates and Adversarial Robustness in Machine Learning · YorkSpace

Generative and synthetic data augmentations introduce distributional shift and unrealistic artifacts

15 theses · 12 institutions

Synthetic samples produced by language models, generative adversarial networks, diffusion models, and speech synthesizers often introduce grammatical errors, domain shifts, and unrealistic features. As a result, synthetic samples fail to generalize to unseen test cohorts and fall behind models trained purely on real data.

Tried and failed

synthetic generative data augmentation applied to atypical speech recognition. Outcome: worse than baseline. Reason: synthetic audio failed to outperform simple addition of unperturbed out-of-domain control data

On matching data and model in LF-MMI-based dysarthric speech recognition · EPFL

Tried and failed

synthetic data augmentation using repurposed conditional embeddings applied to automatic speech recognition across accents. Outcome: worse than baseline. Reason: synthetic multi-accent speech generation did not improve downstream model performance over real data training

Voice conversion and text-to-speech for privacy protection applications · Imperial

Tried and failed

generative text data augmentation applied to text quantification training datasets. Outcome: worse than baseline. Reason: generated ungrammatical, nonsensical text with misspellings, degrading training data quality

Quantification learning with deep neural networks · Iowa State

Tried and failed

standard GAN data augmentation applied to photoplethysmography time-series signals. Outcome: worse than baseline. Reason: generated synthetic samples provided limited performance gain over simple baseline augmentation techniques

Toward accurate health monitoring through large-scale Photoplethysmography signal from wearable devices · Georgia Tech

Tried and failed

heavy synthetic data augmentation during pre-training without adaptation applied to semantic segmentation pre-training. Outcome: worse than baseline. Reason: severe domain shift and distortion introduced by augmentations without calibration layers hurt downstream performance

Label-Efficient Visual Understanding with Consistency Constraints · Virginia Tech

Lost to a baseline

ASR models fine-tuned with multi-accent TTS augmentation (FTaug: 7.88% WER, FMaug: 6.79% WER) were beaten by models fine-tuned on real data alone (FT: 7.66% WER, FM: 6.33% WER)

Voice conversion and text-to-speech for privacy protection applications · Imperial

Considered and rejected

Considered and rejected: Rejected Seq2Seq, GAN, and generative LM models (e.g. BART, PPLM) for text augmentation because generated text contained grammatical errors, misspellings, and lacked meaning.

Quantification learning with deep neural networks · Iowa State

Considered and rejected

Considered and rejected: Rejected adding instances via generative data augmentation to balance data diversity, due to introducing confounds from machine-generated text differing from human text.

Experiment Design for Hypotheses About How NLP Models Work · ResearchWorks

Considered and rejected

Considered and rejected: Rejected standard GANs for data augmentation because synthetic data may not accurately represent the true data distribution

A Novel Lightweight Convolutional Neural Network for Medical Image Classification and Segmentation · HARVEST

Tried and failed

generative neural network for data augmentation applied to imbalanced 3D medical image segmentation. Outcome: did not generalise. Reason: generated augmentations caused heavy overfitting to validation data and failed to generalize to unseen test data

Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial

Tried and failed

semi-supervised learning with synthetic blend augmentations applied to medical image outlier and class classification. Outcome: did not generalise. Reason: improves distinct class accuracy but increases confusion between highly similar or fine-grained classes

Machine learning for outlier detection in medical imaging · Imperial

Tried and failed

generative data augmentation for image classification applied to histopathology whole slide image classification. Outcome: did not generalise. Reason: boosted internal validation metrics but failed to improve or degraded performance on external cohort

Advancing Personalized Medicine Through Generative Artificial Intelligence · Georgia Tech

Tried and failed

diffusion-based synthetic data augmentation for classification applied to medical image classification. Outcome: worse than baseline. Reason: accuracy gains saturated and diminished as real training sample sizes increased

Augmenting medical image classifiers with synthetic data across populations · Harvard

Tried and failed

heuristic and back-translation data augmentation for questions applied to machine reading comprehension. Outcome: worse than baseline. Reason: generated synthetic question variations did not provide useful training signal to improve reading comprehension accuracy

Machine Reading Comprehension: Challenges and Approaches · Cornell

Tried and failed

paraphrasing and surface copy for data augmentation applied to style-constrained text dataset augmentation. Reason: generated text failed to match target domain style and stylistic conventions

Contrastive Text Generation · MIT

Lost to a baseline

SemanticGAN-augmented models achieved lower pick-and-place accuracy (0.32 AC, 0.18 AG) than the baseline trained on 100 real samples with no augmentation (0.44 AC, 0.29 AG).

Data-Efficient Learning Frameworks for Adaptive Intelligent Robots in Human-Robot Collaboration Scenarios · IRIS - POLITO - prod

Tried and failed

pretraining with synthetic trajectory data augmentation applied to visuomotor policy learning. Outcome: worse than baseline. Reason: variable quality in synthetic trajectories degraded downstream performance compared to human-only data

Scaling robot learning with heterogeneous data from the real world, simulation, and the web · UT Austin

Data augmentation induces overfitting, feature redundancy, and demographic bias

16 theses · 11 institutions

Multiplying training data volume or repetitively applying fine-grained perturbations leads models to overfit to transformation artifacts and mask geometries. In addition, these expanded datasets can exacerbate subgroup disparities and amplify pre-existing dataset biases.

Tried and failed

asymmetric data augmentation across dataset streams applied to incremental learning with reference data. Outcome: worse than baseline. Reason: model exploited augmentation artifacts present only in the training stream, degrading representation quality

Referencing Unlabelled World Data to Prevent Catastrophic Forgetting in Class-incremental Learning · Virginia Tech

Tried and failed

heavy multi-sample data augmentation applied to certified robustness via randomized smoothing. Outcome: worse than baseline. Reason: increased class-wise accuracy disparity and degraded overall certified accuracy compared to standard single-sample augmentation

Understanding and Improving Representational Robustness of Machine Learning Models · MIT

Considered and rejected

Considered and rejected: Rejected data augmentation for industrial anomaly detection due to risks of overfitting and need for complex augmentation strategies.

Neuro-Symbolic Integration in Artificial Intelligence and its Applications · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Rejected using a neural network parameterization for data augmentation transformations because it easily overfit validation proxies.

Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial

Tried and failed

naive data duplication for dataset augmentation applied to surgical scene segmentation. Outcome: overfit. Reason: overfitting occurred without providing added diversity to the dataset

Image synthesis with class-aware semantic diffusion models for surgical scene segmentation · Imperial

Tried and failed

standard random color jitter data augmentation applied to semantic image segmentation. Outcome: did not generalise. Reason: random perturbations increased subgroup demographic bias and induced excessive false positive predictions

Color Invariant Skin Segmentation · Virginia Tech

Considered and rejected

Considered and rejected: Decided against fixed-edge mask data augmentation during training because it overfit to mask geometry and failed to generalize as well to localized vector orientations.

Dimensionality reduction for validating engine flow simulations · Oxford

Considered and rejected

Considered and rejected: Rejected copy-pasting small object data augmentation due to high likelihood of overfitting from lack of background blending

Tracking and Measuring Objects in Obscure Image Scenarios Through the Lens of Shot Put in Track and Field · Virginia Tech

Considered and rejected

Considered and rejected: Rejected relying solely on post-hoc stain normalization and data augmentation for multi-site deployment, because models still learned residual site-specific signatures

Privacy-Preserving Federated Learning for Secure and Scalable Digital Pathology · DSpace at SUNY Buffalo

Tried and failed

noisy text augmentation for cross-validation applied to transformer language models on small datasets. Outcome: overfit. Reason: augmentations lacked sufficient diversity and heightened model sensitivity to noise, worsening overfitting on small splits

Privacy-preserving dementia processing in speech · Imperial

Tried and failed

data augmentation to mitigate bias applied to tabular recidivism prediction datasets. Reason: the dataset contained fundamental label bias and historical prejudice rather than simple sample representation imbalance

Multi-objective approaches towards trustworthy machine learning · UT Austin

Considered and rejected

Considered and rejected: Rejected synthetic data generation and data augmentation because small dataset size could lead to overfitting, bias amplification, and false validity.

Methods for Classifying Driver Engagement in Autonomous Vehicles Using Physiological Sensors · Carleton University Institutional Repository

Tried and failed

extending training duration with standard data augmentation applied to image classification models. Outcome: overfit. Reason: training baseline models for longer schedules caused overfitting instead of performance gains

Unified confusion-derived learning framework for image classification · Iowa State

Tried and failed

training dataset expansion via data augmentation applied to object detection model training. Outcome: overfit. Reason: tripling the dataset using data augmenters caused model overfitting, worsening error rates

Reducing measurement error in magnetic particle inspection through the optimization of process parameters and artificially intelligent solutions · Iowa State

Tried and failed

adversarial sample augmentation during iterative training applied to reward model training for alignment. Outcome: overfit. Reason: low adversarial data diversity caused the model to overfit when adding too many generated samples

Robust and Flexible Reward Modeling for LLM Alignment · Georgia Tech

Tried and failed

fine-grained rotational data augmentation applied to image classification models. Outcome: overfit. Reason: excessive rotation steps introduced data redundancy that degraded generalization performance

Development of a method to classify and analyse the composition of mixed waste materials in real-time · Cranfield

Improper pipeline placement, multi-stage compounding, and redundant augmentations yield diminishing returns

16 theses · 8 institutions

Sequentially cascading distinct augmentations, applying augmentations continuously throughout entire training schedules, or inserting them into evaluation steps degrades representation learning. Combining multiple transformations or augmenting already well-represented classes provides minimal benefit over individual techniques or single-stage training.

Tried and failed

combining heuristic augmentations with synthetic data applied to accented speech recognition. Outcome: worse than baseline. Reason: None

Voice conversion and text-to-speech for privacy protection applications · Imperial

Tried and failed

sequential cascading of multiple distinct data augmentations applied to video action recognition training. Outcome: worse than baseline. Reason: compounding intra-clip and cross-clip augmentations degraded visual representations compared to stochastic single-augmentation selection

Action Recognition with Knowledge Transfer · Virginia Tech

Tried and failed

data augmentation during clustering optimization applied to deep image clustering network input. Outcome: worse than baseline. Reason: feeding augmented rather than raw samples directly to the clustering head degraded cluster assignment quality

Yet another image clustering framework using deep learning · Iowa State

Considered and rejected

Considered and rejected: Rejected using full signature augmentation (lead-lag + basepoint) universally without testing simpler models, as it degraded performance on the timing group compared to basepoint alone

Modelling childhood risk factors and time-based patterns for respiratory infections with deep learning and life course data · Imperial

Tried and failed

synthetic data augmentation for majority classes applied to semantic image segmentation. Outcome: no signal. Reason: sufficient real features were already present, yielding minimal to diminishing performance returns

Image synthesis with class-aware semantic diffusion models for surgical scene segmentation · Imperial

Tried and failed

combining multiple data augmentation techniques applied to robotic manipulation out-of-distribution generalization. Outcome: did not generalise. Reason: simple pick-and-place tasks limited the utility of augmentations beyond baseline generation

A High Performance Robotics Data Augmentation Framework · Georgia Tech

Tried and failed

unfiltered auxiliary training data augmentation applied to image classification model training. Outcome: did not generalise. Reason: distribution shift and bias in auxiliary data degraded downstream test performance

Probing, Improving, and Verifying Machine Learning Model Robustness · MIT

Tried and failed

data augmentation with poorly conditioned transformations applied to multilayer perceptron training. Outcome: did not generalise. Reason: higher condition numbers in transformed training data increased test error on clean datasets

Computational Tradeoffs and Symmetry in Polynomial Nonnegativity · MIT

Tried and failed

data augmentation during fine-tuning applied to chemical reaction condition prediction. Outcome: worse than baseline. Reason: None

Integrating AI across the Chemistry Discovery Cycle: Advancing Sustainable Chemistry through Digital Methods · EPFL

Tried and failed

data augmentation across multiple sequential training stages applied to few-shot object detection training. Reason: applying augmentation in both stages offered no compounding benefit over single-stage application

Data- and compute-efficient visual recognition and generation · UT Austin

Tried and failed

combining multiple photometric data augmentations applied to low-contrast object detection. Reason: yielded no considerable performance improvement over individual augmentations alone

Object detection for low-contrast complex background applications · Iowa State

Tried and failed

synthetic detection noise for track data augmentation applied to multi-object tracking models. Outcome: no signal. Reason: random object dropout and false positive injection yielded only marginal performance gains

Autonomous Vehicle Perception Quality Assessment · Virginia Tech

Tried and failed

data augmentation on small training subsets applied to keypoint detection in deep pose estimation. Reason: augmentation provided minimal performance gains over non-augmented training when sample size was very small

Video analytics for lameness detection in dairy cattle: Effects of background removal and deep image matting on farm videos · Iowa State

Tried and failed

contrastive data augmentation and textured mesh rendering applied to vision-language navigation. Outcome: worse than baseline. Reason: augmentations and synthetic mesh renders yielded no performance gains on unseen environments

Towards multi-modal AI systems with open-world cognition · Georgia Tech

Tried and failed

continuous mosaic data augmentation throughout training applied to object detection model training. Outcome: worse than baseline. Reason: applying heavy mosaic augmentation for the entire duration degraded final model performance compared to disabling it late in training

Hierarchical transfer learning for small object detection · Iowa State

Considered and rejected

Considered and rejected: Decided against data augmentation during the AIL reward inference evaluation step, as empirically it decreased performance.

On the use of expert data to imitate behavior and accelerate Reinforcement Learning · OpenBU

Noise and blur perturbations corrupt representations and degrade model robustness

10 theses · 6 institutions

Injecting synthetic Gaussian noise, blur, contrast perturbations, and spectral noise fails to improve environmental invariance and degrades task precision across multiple architectures. In addition, training with artificial noise distortions degrades adversarial robustness and increases vulnerability to out-of-distribution artifacts.

Tried and failed

additive white Gaussian noise data augmentation applied to audio speaker diarization. Outcome: worse than baseline. Reason: models showed poor noise robustness and performance degraded across all architectures

Speaker diarization: importance of the modulation spectrum and incorporating uncertainty modelling · Imperial

Tried and failed

data augmentation with noise and reverberation applied to speaker verification x-vector models. Outcome: worse than baseline. Reason: external noise and reverberation augmentation slightly increased equal error rate

Towards Automatic Analysis of Audio Recordings from Children with Autism Spectrum Disorder · Georgia Tech

Tried and failed

synthetic noise data augmentation during pre-training applied to aerial visual object detection models. Outcome: worse than baseline. Reason: degraded detection precision across multiple evaluation categories instead of improving robustness

Translating AI to Impact: Uncertainty and Human-Agent Interactions in Multi-Agent Systems for Public Health and Conservation · Harvard

Tried and failed

Gaussian noise and blur data augmentation applied to generative image and video models. Outcome: worse than baseline. Reason: Degraded generation and reconstruction performance instead of regularizing the models

Probabilistic learning and generation in deep sequence models · Imperial

Tried and failed

Gaussian noise data augmentation applied to adversarial robustness in deep learning. Outcome: worse than baseline. Reason: training on Gaussian-augmented data degraded robustness against adversarial perturbations

Robust Efficient Edge AI: New Principles and Frameworks for Empowering Artificial Intelligence on Edge Devices · Georgia Tech

Tried and failed

adversarial bias field augmentation applied to medical image segmentation domain generalization. Outcome: did not generalise. Reason: increased vulnerability to out-of-distribution spike noise artifacts, degrading segmentation performance

Improving the domain generalization and robustness of neural networks for medical imaging · Imperial

Tried and failed

Gaussian blur data augmentation for defocus robustness applied to microscopy image segmentation. Outcome: did not generalise. Reason: Synthetic Gaussian blur fails to capture real optical defocus characteristics and provides no benefit over unaugmented data

Single-cell Methods and Spatial Analysis for Highly Multiplexed Tissue Images · Harvard

Tried and failed

Random contrast data augmentation applied to low-contrast object detection. Outcome: worse than baseline. Reason: generated uninformative or misleading features that acted as negative training examples

Object detection for low-contrast complex background applications · Iowa State

Tried and failed

spectral noise data augmentation applied to hyperspectral image semantic segmentation. Outcome: worse than baseline. Reason: None

Hyperspectral Remote Sensing for UXO Detection and Damage Assessment on Airfield Pavements · MIT

Considered and rejected

Considered and rejected: Rejected random rotation and random cropping data augmentation because zero-padded image corners produced all-zero artifact patches.

A Unified Framework for Advancing Soil Erosion and Flood Assessment Through Deep Learning and Process-Based Modeling · Publikationssystem UB Tuebingen

Left open by the authors

Problems the authors named and did not get to.

Left open

Develop multimodal data augmentation methods bridging text and audio by pairing text augmentations with text-to-speech models for online speech processing. Blocker: None

Privacy-preserving dementia processing in speech · Imperial

Left open

Optimize RAG retrieval strategies and data augmentation techniques to enhance fidelity and diversity for sensor-text foundation models. Blocker: Requires the private in-home monitoring sensor dataset and pipeline from the thesis

Developing a foundation model in in-home monitoring data for healthcare applications · Imperial

Left open

Evaluate broader acoustic features and apply noise reduction and data augmentation techniques to improve speech inspiration detection. Blocker: None

Automatic detection of speech inspiration using SVM and decision trees · Cambridge

Left open

Implement data augmentation for acoustic localization by randomly adding or deleting image sources and injecting 30-50 dB SNR noise into RIRs. Blocker: None

Deep Learning Approach to Simultaneously Localize Acoustic Source and Receiver with a Single Room Impulse Response · UT Austin

Left open

Evaluate text classification augmentation techniques including length-restricted synonym replacement, random word insertion, function word deletion, and single-letter modifications. Blocker: None

Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard

Left open

Evaluate length-based text data augmentation across LSTM and recurrent convolutional neural network architectures on text classification benchmarks. Blocker: None

Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard

Left open

Implement data augmentation on cross-site pretraining datasets in TransEHR to reduce the performance gap between full-model fine-tuning and last-layer MLP tuning. Blocker: Requires multi-site electronic health record datasets (like MIMIC or eICU/private clinical data with credentialing or restrictions).

Robust Representation Learning and Real-Time Serving of Deep Models for Health Time Series · Georgia Tech

Left open

Evaluate the impact of different data augmentation levels and techniques on the informal medical entity recognition (NER) model performance. Blocker: None

Supporting laypeople in learning formal medical terminology · DSpace-CRIS at TU Wien

Left open

Implement audio data augmentations and active learning strategies to improve sample balance and accuracy in BirdNET-based transfer learning. Blocker: None

Avian Conservation Bioacoustics in Human-modified Landscapes: Community Responses to Forest Management and Riparian Urbanization in the Pacific Northwest · ResearchWorks

Left open

Combine length-based data augmentation with standard EDA techniques and evaluate text classification performance on benchmark datasets. Blocker: None

Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.