Chapter Four · failure evidence
What Transfer Learning & Domain Adaptation got wrong, from 56 dissertations
The records document various failures of transfer learning and domain adaptation across diverse imaging, sequential, biological, and physical tasks. In many settings, transferred representations and adaptation algorithms introduced negative transfer, failed to bridge domain shifts, or were beaten by simpler target-specific baselines. These records come from PhD theses at 20 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Pretrained natural image representations fail to transfer to specialized non-natural image modalities
Models pretrained on natural images or external datasets struggled to transfer effectively to specialized visual domains like thermal images, microscopy, RF spectrograms, and malware graphs. In many instances, the pretraining representations introduced inductive biases that caused overfitting or underperformed relative to training from scratch.
Tried and failed
transfer learning for image segmentation applied to noisy infrared thermal images. Outcome: worse than baseline. Reason: struggled to segment low signal-to-noise regions compared to hybrid fuzzy clustering
Autonomous Experimentation to Accelerate Boiling Heat Transfer Research · MIT
Tried and failed
transfer learning with natural image pretrained weights applied to satellite land cover image segmentation. Reason: None
Tried and failed
ImageNet pre-trained transfer learning applied to calcium imaging neural decoding. Outcome: worse than baseline. Reason: natural image visual representations transfer poorly to fluorescence microscopy signals compared to training from scratch
Optimizing sensorimotor behaviors through information integration and mental simulation · MIT
Tried and failed
transfer learning with natural image pretraining applied to malware graph visual representation classification. Outcome: worse than baseline. Reason: None
Developing Robust Models, Algorithms, Databases and Tools With Applications to Cybersecurity and Healthcare · Georgia Tech
Tried and failed
transfer learning with natural image pre-trained models applied to simulated microscopy image classification. Outcome: worse than baseline. Reason: Pre-training bias hindered adaptation compared to training from random initialization
Tried and failed
transfer learning with natural image pre-trained weights applied to RF spectral eigengram object detection. Outcome: worse than baseline. Reason: natural image visual features did not transfer well compared to random initialization for non-visual spectrogram representations
Intelligently Leveraging Multi-Channel Image Processing Neural Networks for Multi-View Co-Channel Signal Detection · Virginia Tech
Tried and failed
ImageNet pre-trained transfer learning applied to grayscale image classification. Outcome: did not generalise. Reason: Natural image features do not transfer effectively to grayscale, manufacturing, or medical domains
Engineering-Driven Learning Approaches for Bio-Manufacturing and Personalized Medicine · Georgia Tech
Tried and failed
transfer learning with pre-trained convolutional neural networks applied to droplet deposition density map classification. Outcome: overfit. Reason: models underperformed severely on test data and suffered from poor generalization and overfitting
Automated Exploration of High-Mix, Low Volume Direct Write Design Spaces Through Artificial Intelligence · Georgia Tech
Considered and rejected
Considered and rejected: Rejected using transfer learning from external datasets for the 1000-species classification model to avoid introducing domain training biases.
Deep Learning and Continual Learning Techniques for Plant Image Analysis Tiefes Lernen und kontinuierliche Lernmethoden für die Analyse von Pflanzenbildern · open_UMR Marburg DSpace 10.0
Transfer learning and domain adaptation underperform training from scratch or local target baselines
Across multiple applications, transfer learning models failed to improve upon simple baselines trained solely on the target dataset or simpler non-transfer representations. Complex adaptation schemes and pretraining were routinely matched or beaten by standard empirical risk minimization, random weight initialization, and simple dataset merging.
Tried and failed
transfer learning via large-scale pretraining applied to industrial semantic segmentation. Outcome: worse than baseline. Reason: pretraining on related domain datasets underperformed training from scratch on the target domain
Lost to a baseline
Transfer learning on Goodfellow CNN (10s) on ICU test set achieved F1=0.85, matching zero-shot ECG-FM v1 (F1=0.85) and beaten by non-transfer Inception v3 recurrence plots (F1=0.88).
Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients: A Comparison of Artificial Intelligence Approaches · Queens University Institutional Repository
Lost to a baseline
Transfer learning models for zinc(II) salphen emission energy (Fingerprint + Ridge R2 = 0.831 ± 0.095, MAE = 0.0769 ± 0.0216 eV) showed no significant improvement over the baseline model trained purely on the 49 zinc(II) salphen dataset (Fingerprint + Ridge R2 = 0.825 ± 0.216, MAE = 0.0706 ± 0.0429 eV)
Considered and rejected
Considered and rejected: Rejected pure transfer learning due to decreased classification performance compared to retraining dataset-specific models.
Automated Design and Optimization of Metallic Alloys · unevada
Tried and failed
domain adaptation and generalization algorithms applied to tabular neuroimaging classification across sites. Outcome: worse than baseline. Reason: Specialized domain adaptation methods failed to significantly outperform standard empirical risk minimization.
Fair and Generalizable Machine Learning for Neuroimaging · Penn
Lost to a baseline
Multi-source domain adaptation methods were outperformed by simple merging on the Heintz-Buschart real-data target domain.
Tried and failed
transfer learning from lower-fidelity pre-trained neural networks applied to neural network force field training. Outcome: worse than baseline. Reason: pre-trained weights provided no faster convergence or accuracy gain over random initialization
Machine-learning models for analysis of biomass reactions and prediction of reaction energies · Georgia Tech
Lost to a baseline
Smaller fully-connected neural network models (e.g., 5-2 and 5-5) performed worse when trained with transfer learning (train RMSE 2.80 mm and 2.49 mm) than when trained only with self-supervised learning (2.20 mm and 1.72 mm) or supervised learning (0.97 mm and 1.48 mm).
Enabling Shape-Based Approaches for Autonomous Percutaneous Interventions with Sensorized Needles · JScholarship
Lost to a baseline
Transfer learning models for platinum(II) NNC emission energy (Mordred + RF R2 = 0.760 ± 0.123, test R2 = 0.18, test MAE = 0.14 eV) showed no significant improvement over the baseline model trained purely on the platinum(II) NNC dataset (Mordred + RF R2 = 0.723 ± 0.241, MAE = 0.0927 ± 0.0279 eV)
Distribution alignment and unsupervised domain adaptation fail to bridge domain gaps or cause negative transfer
Adversarial and distribution alignment techniques frequently failed to bridge domain differences in tasks like depth completion and histopathology, sometimes even degrading accuracy compared to unadapted baselines. When domain shift was minimal or target data distributions were already balanced, domain adaptation mechanisms provided no meaningful signal and produced negative transfer.
Tried and failed
unsupervised domain adaptation with distribution alignment applied to large-scale image classification. Outcome: worse than baseline. Reason: negative transfer occurred due to poor target label estimation when source data was plentiful
TOWARDS ROBUST VISUAL PERCEPTION SYSTEMS IN REAL-WORLD ENVIRONMENTS · Cornell
Tried and failed
unsupervised domain adaptation with adversarial training applied to intra-position gesture classification. Reason: source and target distributions within identical positions lacked sufficient domain shift for adaptation benefits
Adaptive gesture recognition for human-robot interface using mechanomyography (MMG) · Imperial
Tried and failed
adversarial feature and output alignment applied to LiDAR depth completion domain adaptation. Outcome: no signal. Reason: provided minimal performance impact and failed to bridge the domain gap
Domain adaptation for semantic and 3D tasks · Imperial
Tried and failed
increasing backbone neural network depth applied to unsupervised domain adaptation. Outcome: did not generalise. Reason: deeper networks did not resolve domain discrepancy and slightly reduced feature transferability across domains
Domain Adaptation with a Classifier Trained by Robust Pseudo-Labels · Virginia Tech
Tried and failed
unsupervised domain adaptation applied to tasks with small domain discrepancy. Outcome: worse than baseline. Reason: source and target domains already had minimal discrepancy, rendering adaptation alignment ineffective
Tried and failed
unsupervised domain adaptation for multiple instance learning applied to out-of-distribution histopathology slide classification. Outcome: did not generalise. Reason: feature and pixel adaptation failed to provide statistically significant improvements on distinct external target cohort
Tried and failed
learning prior class distribution in optimal transport applied to unsupervised domain adaptation. Outcome: no signal. Reason: the target domain had a balanced class distribution matching a uniform prior
A prototype-oriented framework for deep transfer learning applications · UT Austin
Lost to a baseline
Unsupervised domain adaptation alone using CycleGAN without BWE (11.50% EER / 0.532 minDCF on SRE16-YUE-eval40) degraded performance relative to the unadapted/no-BWE baseline (7.46% EER / 0.382 minDCF).
Robust Speaker Recognition using Perceptual and Adversarial Speech Enhancement · JScholarship
Considered and rejected
Considered and rejected: Adversarial domain adaptation was rejected in favor of pseudo-shot learning because adversarial training aligns only domain-level distributions without ensuring correct class-level feature alignment
Coping with the distribution change in soil classification with LIBS · oURspace
Negative transfer occurs across severely mismatched tasks, environments, and biological systems
Severe discrepancies between source and target settings, such as mismatched signal-to-noise ratios, distinct geographic environments, and disparate genomic architectures, led to complete transfer failure. Direct cross-domain transfers from preclinical models or synthetic datasets degraded predictive accuracy because the underlying representations lacked shared inductive structures.
Tried and failed
transfer learning with synthetic pretraining and finetuning applied to subsurface seismic fault segmentation. Outcome: did not generalise. Reason: negative transfer and inductive bias mismatch between distinct synthetic seismic geological models
Multiscale Integration of Cross-Modal Subsurface Data for Reservoir Characterization under Label-Constrained Environments · Georgia Tech
Lost to a baseline
At very low task similarity rho in transfer learning, standard learning with no transfer (delta = 0) beats both hard and soft transfer.
Estimation and Learning via Convex Optimization: Asymptotics, Phase Transitions, and New Algorithms · Harvard
Tried and failed
zero-shot transfer learning with pretrained audio embeddings applied to idiosyncratic atypical vocalization classification. Outcome: no signal. Reason: generic pretrained representations failed to capture idiosyncratic acoustic patterns of atypical vocalizations
Foundations of Cognitive, Affective, and Communicative Systems for Neurodiverse Individuals · MIT
Tried and failed
cross-region transfer learning for acoustic classification applied to audio event detection. Outcome: did not generalise. Reason: models trained on external or online datasets failed to transfer to a new geographic acoustic environment
Listening in on the forest: use of bioacoustics to preserve soundscapes and rare species · Imperial
Lost to a baseline
Transferring from easier (high SNR) to harder (low SNR) synthetic domains was beaten by the limited-data baseline when SNR ranges did not overlap
Foundations of Radio Frequency Transfer Learning · Virginia Tech
Lost to a baseline
Generic zero-shot AudioSet transfer learning achieved only 51.1% accuracy on self-talk, performing barely above chance
Tried and failed
direct cross-dataset model transfer without domain adaptation applied to cross-study drug sensitivity prediction. Outcome: did not generalise. Reason: distribution mismatch between source and target datasets degraded predictive performance without domain mapping
Application of advanced machine learning based approaches in cancer precision medicine · Texas Tech
Tried and failed
direct transfer learning from preclinical cell lines applied to clinical patient drug response prediction. Outcome: did not generalise. Reason: distribution shift and biological discrepancies between in vitro cell models and in vivo human tumors
Predicting cancer patient response to chemotherapy using machine learning from small data · Imperial
Tried and failed
transfer learning across related domains applied to polygenic risk score prediction. Outcome: did not generalise. Reason: insufficient shared genetic architecture between auxiliary and target traits
TRANSFER LEARNING IN CLASSIFICATION AND REGRESSION WITH SUMMARY STATISTICS · Penn
Transfer learning is outperformed by simpler heuristics and standard control models in sequential and dynamical systems
In time series, physical tracking, and degradation forecasting, transfer learning architectures were beaten by basic persistent baselines, model-based controllers, and null models. These models failed to extrapolate dynamic trends such as late-life battery aging and performed worse than basic single-cluster or semi-supervised baselines.
Lost to a baseline
Zero-information transfer learning (testing source classifiers directly on target memory data with simple score averaging) performed significantly worse than typical unidimensional memory classification (t(42) = 1.90, p = 0.032)
More than sum of its parts : investigating episodic memory as a multidimensional cognitive process · UT Austin
Tried and failed
transfer learning for long-term degradation prediction applied to battery state of health estimation. Outcome: did not generalise. Reason: model failed to extrapolate capacity loss during late-life aging phases beyond trained cycle horizons
Modeling and Simulation of Power System with High Penetration of Inverter-based Resources · Georgia Tech
Lost to a baseline
Model-free tracking MSE on target SISO system after transfer learning (2.135e-5) was slightly worse than model-based control with an accurate predictor.
Control of Agentic Systems Using the Newton-Raphson Controller · Georgia Tech
Lost to a baseline
Transfer learning on 2 clusters (MAE 0.127) performed worse than the K=1 base model (MAE 0.123) for LSTM on the 4-week cell-A sample.
A Framework for Generalizing Uncertainty in Mobile Network Traffic Prediction · Virginia Tech
Lost to a baseline
Multimodal transfer learning flood prediction achieved 0.783 accuracy (1-year), lost to naive persistent baseline (0.895 accuracy).
Lost to a baseline
Cold-started and warm-started transfer learning CNN models on M. buryatense copper response performed no better than null baseline models trained on shuffled sequences (F1 ~0.34 vs shuffled F1 ~0.33; Pearson r ~0.26 vs shuffled r ~0.19).
Lost to a baseline
Real-time unsupervised domain adaptation performed statistically worse than the semi-supervised baseline on level ground walking R2 (0.64 vs 0.77).
Enabling Scalable, Versatile, and Robust Control for Robotic Exoskeletons · Georgia Tech
Lost to a baseline
Transfer learning of eqt-pnw was outperformed by semblance ensembling in cross-domain picking accuracy on noisy data
Observing Seismic Variations by Earth and Lab Fluids and Fractures · Harvard
Constrained transfer strategies and frozen feature representations degrade downstream performance
Restricting transfer learning to frozen early layers or fixed feature representations resulted in lower prediction quality than full model fine-tuning. Furthermore, transferring linear projections to nonlinear kernels and extracting representations without target label awareness discarded predictive information and failed to mitigate representation bias.
Considered and rejected
Considered and rejected: Rejected standard transfer learning that updates only pre-trained model weights without retaining source dataset access, due to suboptimality under the data processing inequality.
Fair and Generalizable Machine Learning for Neuroimaging · Penn
Tried and failed
transfer learning with frozen feature layers applied to time series classification. Outcome: worse than baseline. Reason: None
Statistical Machine Learning on Time Series with Applications to Manufacturing and Healthcare · Texas Tech
Considered and rejected
Considered and rejected: Transfer learning by freezing/constraining early neural network layers was rejected because it degraded prediction performance relative to full-network fine-tuning.
Building Blocks of Neural Network Intermolecular Interaction Potentials · Georgia Tech
Tried and failed
transferring linear representation projections to nonlinear kernels applied to molecular property prediction. Outcome: did not generalise. Reason: linear optimization does not account for higher-order feature interactions introduced by polynomial kernel powers
A general and efficient framework for atomistic machine learning · EPFL
Tried and failed
weight decay regularization applied to fixed-feature transfer learning. Outcome: did not generalise. Reason: does not mitigate transfer of representation bias
Considered and rejected
Considered and rejected: Rejected transfer component analysis (TCA) for domain adaptation because it extracts representations without utilizing source domain labels, which can discard features with moderate domain mismatch but high predictive power
Towards Robust Machine Learning for Health Applications · Publikationssystem UB Tuebingen
Left open by the authors
Problems the authors named and did not get to.
Left open
Apply formal transfer learning algorithms to measure and mitigate negative transfer between simulated genomic data and empirical target datasets. Blocker: None
The Inference of Selective Sweep Parameters from their Genomic Footprint · Cornell
Left open
Scale the multi-step Vision Transformer transfer learning framework to broader medical image classification and segmentation tasks across modalities. Blocker: None
Detecting Covid-19 Effectively With Transformers And CNN-based Deep Learning Mechanisms · Texas Tech
Left open
Implement transfer learning using pre-trained computer vision models on low-light plant image datasets for leaf segmentation. Blocker: Requires the private low-light and EMCCD luminescence plant imaging dataset from the thesis lab
Automation of Luminescence Quantitation for High-Throughput Plant Phenotyping Using Image Processing and Deep Learning · TXST Digital Repository
Left open
Develop transfer learning and domain adaptation methods to relax transportability assumptions between clinical trial surrogate marker datasets. Blocker: No specific statistical framework or target algorithm is defined beyond general concepts
Heterogeneous surrogate markers in clinical trials and real-world settings · UT Austin
Left open
Develop domain adaptation algorithms combined with spectral data calibration for LIBS soil classification across distribution shifts. Blocker: Access to the thesis's specific calibrated LIBS soil spectral datasets.
Coping with the distribution change in soil classification with LIBS · oURspace
Left open
Develop a theoretical framework analyzing how contrastive learning affects feature distribution alignment in domain adaptation. Blocker: Lacks specific mathematical framework, formulation, or concrete hypotheses to test
Exploring Deep Representation Learning on Vision and Language Intelligence · DukeSpace
Left open
Analyze theoretical and empirical conditions under which simple multi-source data merging outperforms domain adaptation in regression settings. Blocker: None
Left open
Investigate domain adaptation methods for short-sequence LSTM ensembles applied to significantly different source and target dynamic systems. Blocker: The objective is a broad research direction without specific target datasets, adaptation approaches, or evaluation metrics.
Parameter optimization for enhanced modeling of dynamic systems · Iowa State
Left open
Evaluate pretrained transformer models with subword tokenization and transfer learning on scraped job postings to predict bankruptcy and corporate growth. Blocker: Scraped job listings dataset may not be publicly archived with the thesis
Assessing Corporate Growth and Bankruptcy Risk Using Public Data Proxies · Harvard
Left open
Develop unified time-series foundation models across diverse datasets using transfer learning and domain adaptation techniques. Blocker: None
Graph-based Time-series Forecasting in Deep Learning · Virginia Tech
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.