Chapter Four · failure evidence
What Multi-Task Learning got wrong, from 46 dissertations
Multi-task learning records show frequent failures when joint training across multiple objectives leads to negative transfer, optimization instability, or worse accuracy than single-task baselines. Effective multi-task performance depends heavily on appropriate loss weighting, balanced parameter sharing, proper task scheduling, and explicit task conditioning. These records come from PhD theses at 17 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Negative transfer causes multi-task models to underperform single-task baselines
Joint multi-task training across diverse or conflicting objectives frequently degrades performance compared to specialized single-task models. Across vision, language, biology, and reinforcement learning domains, task competition and negative transfer result in lower accuracy, higher error, or inferior policies.
Tried and failed
multi-task learning with shared encoder applied to multi-modal object detection and segmentation. Outcome: worse than baseline. Reason: negative transfer between distinct task objectives degraded individual task performance
Co-Designing Efficient Systems and Algorithms for Sparse and Quantized Deep Learning Computing · MIT
Tried and failed
multi-task transfer without task similarity weighting applied to text classification across diverse tasks. Outcome: worse than baseline. Reason: incorporating all source tasks equally without similarity grouping caused negative transfer
Human Learning-Augmented Machine Learning Frameworks for Text Analytics · Virginia Tech
Tried and failed
single multi-task CNN for distinct tissue structures applied to histological structure identification in confocal microscopy. Outcome: worse than baseline. Reason: joint learning of distinct morphological targets degraded performance relative to specialized single-task models
Tried and failed
multi-task reinforcement learning query policies applied to heterogeneous entity penetration testing. Outcome: worse than baseline. Reason: negative transfer across distinct entity types degraded performance on other tasks
Securing access control using machine learning and formal methods · UT Austin
Tried and failed
joint multi-task learning with auxiliary task applied to few-shot image segmentation. Outcome: worse than baseline. Reason: None
Learning to manipulate images with image segmentation, search and synthesis · UT Austin
Lost to a baseline
Task-conditioned multi-task BC lost to the single-agent BC baseline on the Gravity 1.5 HalfCheetah task due to negative transfer.
Data efficiency in imitation learning with a focus on object manipulation · Imperial
Lost to a baseline
Single-task variants (STVs) beat multi-task models in HH income prediction when data subsampling reached lowest fractions due to negative transfer.
Behaviorally Informed Machine Learning for Human Mobility · ResearchWorks
Lost to a baseline
On breast cancer dataset GSE4922, single-task analysis achieved 0.83 AUC whereas multi-task learning (MTL) achieved 0.81 AUC.
High-Dimensional Data Integration with Multiple Heterogeneous and Outlier Contaminated Tasks · YorkSpace
Lost to a baseline
3D multi-task U-Net (Dice 0.831) lost to 3D single-task U-Net (Dice 0.867) on white matter segmentation
Improving deep-learning segmentation performance in 3D neuroimaging with minimal manual annotations · Oxford
Tried and failed
multi-task reinforcement learning on diverse dynamics applied to adaptive numerical step-size controllers. Outcome: worse than baseline. Reason: joint training produced suboptimal hybrid policies that performed worse than single-task policies
Reinforcement Learning for Self-adapting Time Discretizations of Complex Systems · Virginia Tech
Lost to a baseline
Single-task learning outperformed multi-task learning with UDA on facial Action Unit recognition (0.78 vs 0.75 mean F1-accuracy with Bottleneck).
Low-Resource Neural Adaptation: A Unified Data Adaptation Framework for Neural Networks · ResearchWorks
Lost to a baseline
On predicate grounding, PIPELINE achieved 74.0 F1 vs MULTI-TASK achieving 72.3 F1.
Information Extraction on Scientific Literature under Limited Supervision · Georgia Tech
Lost to a baseline
PointPillars single-task baseline achieved 66.3% AP@0.7 vs PillarFlowNet multi-task at 54.7% AP@0.7 and single-task at 65.4% AP@0.7
Deep Sensor Data Fusion for Environmental Perception of Automated Systems · Publikationssystem UB Tuebingen
Lost to a baseline
On single-source speaker verification trials, fine-tuned sequential single-channel X-vector baseline beat the proposed multi-task network (3.80% vs 4.87% EER on 10s segments).
Deep Learning Approaches for Auditory Perception in Robotics · EPFL
Lost to a baseline
Single-task damage detection achieved 97.5 ± 1.5 Real AUC, slightly outperforming the multi-task TransReI3D (97.3 ± 2.2).
Tackling Data Challenges in Computer Vision · IRIS - POLITO - prod
Lost to a baseline
Single-task fine-tuning beat joint multi-task fine-tuning on UAV-Captions (BLEU-4 53.27 vs 49.02) and RSVQA-LR (88.56% vs 88.13% avg accuracy)
Advanced Methods for Remote Sensing Image Captioning · IRIS - UNITN - prod
Lost to a baseline
MFBERT multi-task learning on Tox21 (mean ROC-AUC 0.80) lost to MoleculeNet single-task models (best ROC-AUC 0.85)
Contextual representations of the chemical space for task agnostic machine learning methods · Imperial
Tried and failed
joint multi-task training with shared representations applied to heterogeneous multi-task learning. Outcome: did not generalise. Reason: task competition causing lower per-task generalization than isolated models
Tried and failed
joint multi-task surrogate regression modeling applied to multi-objective molecular property prediction. Outcome: worse than baseline. Reason: joint multi-property training underperformed separate single-property models
AI for materials design: Generative AI with multi-fidelity strategies · Iowa State
Tried and failed
multi-task learning with pretrained molecular transformers applied to multi-target bioassay toxicity prediction. Outcome: worse than baseline. Reason: None
Contextual representations of the chemical space for task agnostic machine learning methods · Imperial
Gradient imbalance and inappropriate loss weighting destabilize multi-task optimization
Assigning extreme, static, or uncalibrated loss weights across multiple tasks leads to gradient domination by auxiliary tasks or severe optimization instability. This imbalance prevents shared networks from converging or causes the primary task performance to collapse entirely.
Tried and failed
extreme weighting in multi-task contrastive learning applied to medical image classification. Outcome: worse than baseline. Reason: setting the contrastive loss weight either too low or too high degraded target task performance
Integrating domain knowledge and deep learning for enhanced chest X-ray diagnosis and localization · UT Austin
Tried and failed
multi-task neural network with shared weights applied to coupled multiphysics field prediction. Outcome: did not converge. Reason: gradient imbalance between the different coupled physical output loss terms
Considered and rejected
Considered and rejected: Multi-task adversarial learning was rejected early due to training instability, lack of guarantees in debiasing latent representations, and inability to handle unmeasured confounders.
Unsound foundations: refining AI’s role in audio-based COVID-19 detection · Imperial
Tried and failed
high auxiliary loss weighting in multi-task reinforcement learning applied to multi-agent policy learning. Outcome: unstable. Reason: auxiliary concept gradients dominate the policy optimization, causing task performance collapse
Considered and rejected
Considered and rejected: Rejected equal weighting across multi-task reconstruction errors in CVAE in favor of adaptively learned task-dependent uncertainty weighting
Semantic neural representation for SLAM and scene understanding · Imperial
Unconstrained or rigid parameter sharing creates representational interference across tasks
Forcing complete parameter sharing or using unconstrained shared layers across heterogeneous tasks prevents necessary task-specific specialization. These rigid architectures lead to inter-task conflicts, poor sample efficiency, or complete failure to learn complex multi-task behaviors.
Tried and failed
dense attention with unconstrained parameter sharing applied to multi-task multi-agent reinforcement learning. Outcome: did not generalise. Reason: leads to negative transfer or poor sample efficiency across heterogeneous tasks
Robust and Scalable Multiagent Reinforcement Learning in Adversarial Scenarios · MIT
Tried and failed
full parameter sharing across multi-task reinforcement learning applied to multi-policy multi-MDP optimization. Outcome: worse than baseline. Reason: sharing all layers prevented task-specific specialization, degrading performance to single-task baseline levels
Learning for optimized provisioning of network resources · Imperial
Considered and rejected
Considered and rejected: Rejected one-hot coding multi-task input vectors for contingency DSA because it increased data dimensions, degraded training efficiency, and introduced inter-task conflicts in a single network.
Lost to a baseline
In Multi-Task RL (MTRL) on MinAtar, TopKRouter (0.51 ± 0.02) underperformed the dense baseline (0.63 ± 0.03).
Intelligent interaction at scale · Oxford
Considered and rejected
Considered and rejected: Rejected non-modular MLPs with concatenated goal-state inputs for multi-task robot manipulation after achieving 0% success on training tasks
Interactive and Explainable Methods in Machine Learning with Humans · Georgia Tech
Flawed task scheduling and sequential updates cause negative transfer and catastrophic forgetting
Updating multi-task models sequentially or selecting tasks via uniform or random schedules results in negative transfer and catastrophic forgetting of earlier tasks. Effective multi-task convergence requires proportional sampling or simultaneous joint optimization rather than decoupled sequential updates.
Tried and failed
random task selection in multi-task learning applied to dialogue act classification. Outcome: worse than baseline. Reason: causes negative transfer across tasks during training
Enhancing speech intelligibility through paralinguistic features · Imperial
Tried and failed
sequential multi-task training across schemas applied to cross-ontology information extraction. Outcome: worse than baseline. Reason: sequential training underperformed compared to simultaneous joint training across ontologies
Towards Generalizable Information Extraction with Limited Supervision · Virginia Tech
Tried and failed
multi-task fine-tuning of a language model applied to inter-parameter dependency and value generation. Outcome: worse than baseline. Reason: catastrophic forgetting degraded performance on dependency detection task
Effective Automation of Black-Box Testing for REST APIs with Machine Learning and Language Models · Georgia Tech
Lost to a baseline
Reverse Annealing and Uniform Sampling multi-task schemes lost to Proportional Sampling in unseen in-context learning test loss during instruction-tuning.
Towards Understanding Conflict & Transfer in Multi-Task Deep Learning · JScholarship
Considered and rejected
Considered and rejected: Avoided sequential task parameter updates in incompatible multi-task reinforcement learning due to vulnerability to catastrophic forgetting.
Missing task conditioning, corrupted labels, and auxiliary redundancy degrade shared representations
Multi-task networks fail when inputs lack task specification prompts, when untargeted tasks introduce default negative labels that contaminate representations, or when modality features are naively averaged. Furthermore, auxiliary multi-task objectives provide no performance benefit when pretrained representations already capture the necessary target signals.
Tried and failed
multi-task learning with negative default labeling applied to sparse group-specific text classification. Outcome: worse than baseline. Reason: default negative labels on untargeted tasks contaminated shared representations, collapsing minority recall
Toxic language and target detection under sparsity by modeling group-specific representations · UT Austin
Considered and rejected
Considered and rejected: Rejected null prompts in text-to-text multi-task learning because models fail to infer task specification from input text distributions alone (leading to severe negative transfer).
Towards Understanding Conflict & Transfer in Multi-Task Deep Learning · JScholarship
Considered and rejected
Considered and rejected: Rejected naive feature averaging in multi-task ControlNet/T2I adapters because it dilutes modality information in contradictory spatial regions.
Enabling Conditional Generation Without Training: Plug & Play Generative Modelling using Diffusion Models · JScholarship
Tried and failed
multi-task learning with pretrained transformer representations applied to multi-word expression parsing and tagging. Outcome: no signal. Reason: pretrained language model representations already capture auxiliary task information, rendering joint training benefits redundant
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop training schemes to mitigate negative transfer during joint 3D detection and map segmentation multi-task learning in BEVFusion. Blocker: None
Efficient Deep Learning with Sparsity: Algorithms, Systems, and Applications · MIT
Left open
Develop a multi-task joint training scheme for BEVFusion to eliminate negative transfer between 3D object detection and bird's-eye view segmentation. Blocker: None
Co-Designing Efficient Systems and Algorithms for Sparse and Quantized Deep Learning Computing · MIT
Left open
Extend AxialNet to multi-task learning combining CT abnormality prediction with lesion detection using the public DeepLesion dataset. Blocker: None
Towards Fully Automated Interpretation of Volumetric Medical Images with Deep Learning · DukeSpace
Left open
Develop a theoretical framework explaining why learned multi-task representations outperform single-agent baselines in non-identical drone navigation environments. Blocker: None
Designing policy optimization algorithms for multi-agent reinforcement learning · Georgia Tech
Left open
Develop multi-task debiasing architectures to evaluate and debias auxiliary tasks alongside primary NLP tasks across diverse benchmarks. Blocker: None
Robustness in natural language processing · Imperial
Left open
Evaluate multi-task regression runtime prediction performance under alternative offloading, pipeline selection, and data querying scenarios. Blocker: None
Predicting Performance Run-time Metrics in Fog Manufacturing using Multi-task Learning · Virginia Tech
Left open
Implement and evaluate task-aware distillation with multi-filter routing in a multi-task learning setting. Blocker: None
On Parameter Efficiency of Neural Language Models · Georgia Tech
Left open
Benchmark the IL-SOAR algorithm against multi-task baselines like IMPALA across standard non-UAV reinforcement learning environments. Blocker: None
Multi-Task Reinforcement Learning: From Single-Agent to Multi-Agent Systems · Virginia Tech
Left open
Develop and train a multi-task neural network architecture with shared representations that jointly predicts tight constraints and optimal final times. Blocker: None
Left open
Evaluate Context Label Learning (CoLab) under domain shifts and multi-task settings using a shared backbone model to identify subclasses across tasks. Blocker: None
Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.