Chapter Four · failure evidence
What Image Segmentation & Computer Vision got wrong, from 95 dissertations
The records document practical breakdowns and engineering trade-offs encountered across computer vision and image segmentation systems. Across these trials, deep neural models and heuristic tools frequently fail due to domain shifts, heavy computational demands, boundary confusion, and underperformance relative to simpler baselines. These records come from PhD theses at 32 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Complex neural architectures often underperform simpler models or established baselines
Attempts to deploy diffusion models, promptable foundation models, adversarial training, and single-stage instance detectors often yielded higher error rates and worse segmentation accuracy than standard discriminative baselines. In addition, simpler classifiers, basic U-Nets, and traditional clustering algorithms regularly matched or exceeded the performance of these larger designs.
Tried and failed
diffusion models applied to deterministic semantic segmentation. Outcome: worse than baseline. Reason: did not outperform standard discriminative models on deterministic prediction benchmarks
Fast and Future: Towards Efficient Forecasting in Video Semantic Segmentation · EPFL
Tried and failed
promptable foundation model post-processing for segmentation applied to color-sensitive semantic segmentation. Outcome: worse than baseline. Reason: expanded false positives by segmenting non-target objects sharing color features
A Machine Learning Approach to Recognize Environmental Features Associated with Social Factors · Virginia Tech
Lost to a baseline
A simple 1% labeled-frame threshold baseline beat the complex combined labeling-and-segmentation dual-threshold method in prognostic stratification (85.96% vs 83.33%).
Development of Lung Ultrasound Quantitative Approaches and Automatic Semi-Quantitative Strategies: In Silico, In Vitro, and Clinical Studies · IRIS - UNITN - prod
Lost to a baseline
Original SAM achieved higher recall (0.53 vs 0.41) than the fine-tuned cropping approach on bare cropland, though driven by over-segmentation/false positives.
A Unified Framework for Advancing Soil Erosion and Flood Assessment Through Deep Learning and Process-Based Modeling · Publikationssystem UB Tuebingen
Lost to a baseline
Direct 2D segmentation without priors (Deep-Single) had higher surface error outliers on noisy B-scans compared to the graph-based Aura baseline.
RETINAL OCT IMAGE ANALYSIS USING DEEP LEARNING · JScholarship
Lost to a baseline
ResNet with ASPP module on radargram segmentation achieved overall accuracy slightly lower than literature SVM baselines
Advanced methods for simulation-based performance assessment and analysis of radar sounder data · IRIS - UNITN - prod
Lost to a baseline
Epistemic uncertainty data selection (Hausdorff 4.15 mm) lost to baseline fully supervised (Hausdorff 3.97 mm) in MRI segmentation
Improving deep-learning segmentation performance in 3D neuroimaging with minimal manual annotations · Oxford
Tried and failed
Double deep image prior unsupervised segmentation applied to pathological tissue segmentation. Outcome: worse than baseline. Reason: yielded very low overlap and aggregated Jaccard index unsuitable for complex tissue structures
Recognition, retrieval, and harmonisation for multicentre clinical data analysis · Imperial
Lost to a baseline
Adv-trained ResDSN Coarse 3D medical image segmentation performance on clean data (79.09% accuracy/Dice) lost to baseline ResDSN Coarse (87.84%).
TOWARDS DEEP LEARNING ROBUSTNESS FOR COMPUTER VISION IN THE REAL WORLD · JScholarship
Lost to a baseline
Custom-trained ilastik pixel classifier underperformed generic pre-trained deep learning models (Cellpose and StarDist) on nuclear segmentation of crowded/touching cells
Exploring tumour heterogeneity and responses to therapy using single-cell resolved microscopy · Imperial
Lost to a baseline
DeepLabv3 achieved lower accuracy metrics (F1 score 0.751 vs 0.809) on visible image segmentation compared to FPN despite having over 2x the parameter count.
Condition Assessment of Civil Infrastructure and Materials Using Deep Learning · Virginia Tech
Lost to a baseline
SegResNet demonstrated slightly higher sensitivity in overall liver tumor segmentation than SmoothSegNet despite having lower accuracy and Dice score due to over-segmentation.
Knowledge-Informed Weakly-Supervised Deep Learning Models for Cancer Applications · Georgia Tech
Lost to a baseline
SDXL-Turbo achieved lower semantic segmentation performance (36.99% overall Dice-AUC) compared to full SDXL (38.27% overall Dice-AUC).
Considered and rejected
Considered and rejected: Rejected using a dual 3D segmentation head ensemble for domain generalization in LiDOG as it performed worse than the auxiliary 2D BEV projection decoder (30.07 vs 44.18 mIoU).
Handling Domain Shift in 3D Point Cloud Perception · IRIS - UNITN - prod
Tried and failed
real-time one-stage instance segmentation applied to sidewalk semantic segmentation. Reason: produced inaccurate segmentations and exhibited high sensitivity to input image resolution without specific backbone tuning
Tried and failed
single-stage real-time instance segmentation applied to onboard robotic vision. Outcome: worse than baseline. Reason: produced lower-quality segmentation masks compared to an optimised two-stage instance segmentation model
Autonomous exploration and object reconstruction with an MAV · Imperial
Tried and failed
fine-tuned real-time instance segmentation models applied to anatomical structures in surgical video. Outcome: did not generalise. Reason: model suffered extreme false positive rates and low specificity on unseen test video frames
Lost to a baseline
U-Net beat U-KAN on Sentinel-1 crop field segmentation Recall (85.56% vs 77.50%)
Spatio-Temporal Machine Learning for Ecology and Crisis Management · IRIS - POLITO - prod
Tried and failed
hardcoding domain-specific anatomical constraints into neural network applied to image segmentation models. Outcome: worse than baseline. Reason: rigid boundary rules reduced segmentation accuracy and introduced software instability
A study of the relationship between monocotyledonous plant anatomy and water · University of Nottingham Repository
Lost to a baseline
CoBEVT achieved 60.4% mIoU on vehicle BEV segmentation and 63.0% on drivable area, beating CoBEVFusion (59.5% and 61.7% mIoU)
Enhancing Perception for Autonomous Vehicles · Queens University Institutional Repository
Lost to a baseline
K-means achieved higher Precision in traffic frame segmentation (0.89 vs 0.87 for HHGATSD)
Novel optimization methods and model for improving sustainability and efficiency in last-mile logistics · DeustoTeka
Cross-dataset distribution shifts and synthetic data fail to generalize to real clinical and field imagery
Segmentation networks trained on synthetic objects, specific scanner resolutions, or narrow geographic datasets experienced severe performance drops and merge errors when transferred to real target domains. Generative augmentations and heuristic synthetic adjustments failed to bridge these protocol gaps and occasionally created artifacts that further degraded test accuracy.
Tried and failed
U-Net segmentation cross-dataset transfer applied to histology image segmentation. Outcome: did not generalise. Reason: Domain shift from differing staining and imaging protocols reduced accuracy to simple thresholding baseline
Computational modelling of diffusion magnetic resonance imaging based on cardiac histology · Imperial
Tried and failed
direct cross-resolution neural network deployment applied to medical image segmentation. Outcome: did not generalise. Reason: resolution and signal-to-noise mismatch between high-resolution multi-average training data and lower-resolution single acquisitions
Deep learning-based analysis of multiple sclerosis lesions with high and ultra-high field MRI · EPFL
Tried and failed
generative neural network for data augmentation applied to imbalanced 3D medical image segmentation. Outcome: did not generalise. Reason: generated augmentations caused heavy overfitting to validation data and failed to generalize to unseen test data
Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial
Tried and failed
off-the-shelf object segmentation without domain-specific fine-tuning applied to surgical video anatomy detection. Outcome: did not generalise. Reason: models pre-trained on common natural images fail to recognize specialized surgical anatomy
XMARCUS: A Pathway Towards Remote Robotic Surgery Coaching · Virginia Tech
Tried and failed
direct transfer of synthetic-trained 3D CNNs applied to real 3D volumetric image segmentation. Outcome: did not generalise. Reason: domain gap between synthetic training data and noisy real data caused discontinuous predictions and low recall
Multiscale Integration of Cross-Modal Subsurface Data for Reservoir Characterization under Label-Constrained Environments · Georgia Tech
Tried and failed
direct cross-dataset model transfer without domain adaptation applied to electron microscopy 3D segmentation. Outcome: did not generalise. Reason: acquisition variations across datasets caused severe undersegmentation and high merge errors
Grounded: Inference via Local Signals & Learned Representations by Organic & Artificial Systems · Harvard
Tried and failed
single-modality deep learning auto-segmentation for error screening applied to radiotherapy target volume contour validation. Outcome: did not generalise. Reason: CT-only segmentation produced an unacceptably high false positive rate for anomaly detection
Tried and failed
mixup data augmentation applied to medical image segmentation across population shifts. Outcome: worse than baseline. Reason: linear interpolation created unrealistic images that degraded generalization across pathological groups
Improving the domain generalization and robustness of neural networks for medical imaging · Imperial
Tried and failed
adversarial bias field augmentation applied to medical image segmentation domain generalization. Outcome: did not generalise. Reason: increased vulnerability to out-of-distribution spike noise artifacts, degrading segmentation performance
Improving the domain generalization and robustness of neural networks for medical imaging · Imperial
Lost to a baseline
Heuristic TEA degraded performance for DeepMedic on cross-site prostate segmentation (Site B) compared to no TEA (67.4% vs 71.7% DSC with learned class-specific TRA).
Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial
Lost to a baseline
Models trained with synthetic long cells showed decreased segmentation performance on in-distribution short cells compared to models trained exclusively on short cells.
Bacterial Deepfakes: Generating Synthetic Microscopy Data to Improve Adaptability of Deep Learning-Based Segmentation Models · ResearchWorks
Considered and rejected
Considered and rejected: Rejected fully simulated datasets as sole ground truth for defect segmentation because synthetic defects fail to capture real shape complexity, boundary definition, and noise interactions.
End-to-End Artificial Intelligence-Based Pipeline for Quantitative Reporting of Lung V/Q Scintigraphy—VQ-SPRINT: Segmentation, Pseudo-planar Generation, and Registration INtegration Tool · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Rejected supervised machine-learning organ segmentation trained on external datasets (CT-ORG, AAPM Thoracic) due to domain shifts and severe disease exclusion in prior benchmarks, choosing unsupervised morphological processing instead.
Towards Fully Automated Interpretation of Volumetric Medical Images with Deep Learning · DukeSpace
Tried and failed
semantic segmentation model trained on urban driving applied to aerial building facade and ground segmentation. Outcome: did not generalise. Reason: domain shift between urban street-level perspective and aerial or campus imagery degraded feature recognition
Aerial Cadastral and Flood Assessment for Disaster Risk Management in Appalachia · Virginia Tech
Tried and failed
synthetic object injection for anomaly segmentation applied to road scene obstacle detection. Outcome: did not generalise. Reason: biases model toward large nearby objects, missing small distant obstacles and causing false positives
High computational complexity and inference latency prevent real-time deployment
Volumetric 3D networks, dense vision transformers, multi-modal fusion heads, and connected component post-processing created excessive runtime delays and hardware strain. Because of these processing bottlenecks, researchers rejected pixel-wise architectures or selected lighter models to maintain acceptable frame rates in robotics and clinical imaging.
Tried and failed
pointwise Gaussian process regression applied to point cloud segmentation. Outcome: too slow. Reason: computational cost was too high for real-time processing and struggled with occlusion
Online vehicle trajectory extraction based on LiDAR data · Texas Tech
Tried and failed
star-convex object detection for tracking and segmentation applied to large-scale high-throughput image sequences. Outcome: too slow. Reason: per-frame inference caused excessive overall processing times on long high-throughput video sequences
Towards systems biophotonics in microscopy, medicine, and robotics · Georgia Tech
Tried and failed
connected-component analysis post-processing applied to neural network image segmentation masks. Outcome: too slow. Reason: Increased inference time tenfold without providing any meaningful improvement in segmentation accuracy
Lost to a baseline
Single-image U-Net and DeepLabV3 segmentation models significantly outperformed all fusion architectures (CMNeXt, MMSFormer, StitchFusion) in inference latency (5.0–5.2 ms vs 18.2–83.9 ms).
Autonomous System for Identifying and Capturing Floating Waste · Georgia Tech
Lost to a baseline
The proposed hybrid segmentation algorithm required double the processing time of LSA and OHRH baseline methods due to dynamic per-loop metric updates
An improved segmentation and classification method for building extraction from RGB images using GEOBIA framework · Queens University Institutional Repository
Lost to a baseline
VGG-16 semantic segmentation achieved only 2 fps with frequent misidentifications and lost to manual contact determination
Characterisation of a Novel Bioadhesive Used by the Ctenophore Pleurobrachia pileus · Research Repository UCD
Considered and rejected
Considered and rejected: Rejected CNN-based face segmentation models for the real-time wearable implementation due to high computational complexity.
Considered and rejected
Considered and rejected: Direct 3D deep learning segmentation rejected due to excessive computational expense and hardware demands for real-time ultrasound imaging, opting for 2D slice-by-slice U-Net.
Forward-Viewing Ultrasound Guidance of a Robotically-Steered Guidewire for Peripheral Interventions · Georgia Tech
Considered and rejected
Considered and rejected: Rejected instance object segmentation (e.g., Mask R-CNN/UNet) in favor of bounding-box object detection (YOLOv4) due to inference latency and higher data annotation burden.
Vision-Based Force Planning and Voice-Based Human-Machine Interface of an Assistive Robotic Exoskeleton Glove for Brachial Plexus Injuries · Virginia Tech
Considered and rejected
Considered and rejected: Rejected pixel-wise semantic/instance segmentation and GrabCut for object extraction due to high computational overhead in robotics.
Scene understanding via scene graph Szenenverständnis mittels Szenengraph · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Rejected 3D CNNs and R-CNN architectures for dynamic frame segmentation because driving video frames lack strict temporal dependence and U-Net was more computationally efficient.
Methods for Classifying Driver Engagement in Autonomous Vehicles Using Physiological Sensors · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Decided against transformer architectures for semantic segmentation due to extreme computational expense and large data requirements compared to CNNs.
Integration of machine learning for enhanced digital rock physics workflows · UT Austin
Considered and rejected
Considered and rejected: Decided against semantic segmentation architectures (pixel-wise segmentation) due to high computational cost and because area composition estimation does not require precise item boundary segmentation.
Development of a method to classify and analyse the composition of mixed waste materials in real-time. · Cranfield
Considered and rejected
Considered and rejected: Rejected standard 3D CNNs from scratch for organ segmentation due to training instability, lack of 3D pre-trained weights, and high computational cost
Towards Robust Deep Learning for Medical Image Analysis · JScholarship
Semantic segmentation struggles with overlapping boundaries and topological continuity
Pixel-wise semantic classification frequently merged adjacent or intersecting instances into single undifferentiated masks and produced jagged or fragmented lines. Standard loss functions and uncropped single-stage networks failed to maintain instance boundaries, fine curvilinear structures, and topological continuity.
Tried and failed
centerline and distance-based segmentation loss functions applied to thin curvilinear structure segmentation. Outcome: worse than baseline. Reason: degraded segmentation performance compared to standard Dice loss
Pavement Crack Segmentation with Dense Local Geometry Features and Boundary Enhancement Loss · Georgia Tech
Lost to a baseline
Single-class semantic segmentation achieved 85% Mean IoU, whereas two-class training achieved lower Mean IoU (80%) due to irregular, jagged predicted boundaries on metric bars.
Integration and classification of spatial data for 3D modelling and monitoring of built heritage · IRIS - POLITO - prod
Lost to a baseline
Mesmer whole-cell segmentations produced more rounded boundaries, under-capturing irregular membrane/cytoplasmic extensions compared to traditional watershed segmentation.
Computational Methods to Process and Analyze High-dimensional Imaging Mass Cytometry Datasets in Pathological Tissue Samples · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected semantic segmentation (pixel-wise shape classification) because precise item borders are unneeded for mass/area composition estimation and fails on cluttered waste
Development of a method to classify and analyse the composition of mixed waste materials in real-time · Cranfield
Considered and rejected
Considered and rejected: Rejected relying purely on monocular semantic segmentations without instance labels because overlapping obstacles merge into single impassable blocks.
Perception Enabled Planning for Autonomous Systems · Cornell
Considered and rejected
Considered and rejected: Semantic segmentation via standard U-Net rejected due to inability to differentiate multiple or intersecting instruments without extensive post-processing
Detection and 3D Localization of Surgical Instruments for Image-Guided Surgery · JScholarship
Considered and rejected
Considered and rejected: Rejected text annotations from the WiSe dataset because its semantic segmentation masks lacked instance granularity (words/lines) needed for tracking.
Lecture Video Summarization by Detection and Representation of Content · DSpace at SUNY Buffalo
Considered and rejected
Considered and rejected: Directly using semantic segmentation 'otherprop' class alone for object extraction without plane detection (rejected due to ScanNet under-segmentation causing poor object boundary precision)
Object change detection for autonomous indoor robots in open-world settings · DSpace-CRIS at TU Wien
Considered and rejected
Considered and rejected: Single-stage semantic segmentation network on uncropped images was rejected due to unacceptable false positive rates on internal organelles and out-of-focus boundaries.
Quantifying RTK Signal Transduction Processes Across The Plasma Membrane · JScholarship
Tried and failed
semantic segmentation of spatially adjacent overlapping features applied to medical image segmentation of multiple findings. Outcome: did not generalise. Reason: models confuse boundary distinctions when multiple distinct localized patterns co-occur in close spatial proximity
AI Systems for Understanding and Grounding Radiology Reports · Harvard
Tried and failed
pixel-wise loss functions for continuous line segmentation applied to crack detection in structural images. Reason: standard pixel losses fail to preserve topological continuity, producing fragmented segmentations
Considered and rejected
Considered and rejected: Rejected stateful panoptic fusion for streaming sectors because it causes under-segmentation for objects spanning into upcoming unseen sectors.
On the Use of Vision and Range Data for Scene Understanding · JScholarship
Considered and rejected
Considered and rejected: Segmentation-only turning lane validation was rejected because it generated noisy, disconnected segmentation masks for valid lanes
Enriching Digital Maps with Aerial Imagery and GPS Data · MIT
Intensity and heuristic thresholding techniques fail in noisy and variable environments
Global, adaptive, and batch thresholding strategies proved ineffective when target objects were obscured by confounding hyper-intensities, uneven illumination, or low contrast. These rigid heuristic cutoffs caused extensive false negatives on subtle structures and misclassified background noise as foreground targets.
Tried and failed
intensity thresholding segmentation applied to anatomical structure segmentation in medical imaging. Outcome: worse than baseline. Reason: confounding hyper-intensities prevent accurate separation of target structures from background tissue
Deep Learning for Localizing and Segmenting Anatomies in Medical Imaging · Cornell
Tried and failed
Gaussian mixture model automatic thresholding applied to time-lapse microscopy image segmentation. Outcome: worse than baseline. Reason: Failed to accurately segment features across time-lapse frames compared to a static threshold
Precipitation Dynamics at the Solution-Solution Interface in Confined Geometries, and the Effects of Organics on Precipitate Evolution · Georgia Tech
Tried and failed
color thresholding and skeletonization applied to overlapping root system segmentation. Reason: heuristic segmentation failed to separate complex overlapping structures from background
Tried and failed
simple intensity thresholding for segmentation applied to subcellular molecular cluster detection. Reason: missed low-intensity functional clusters and failed to resolve closely spaced assemblies
MECHANISMS OF TRANSCRIPTION FACTOR HUB FORMATION AND FUNCTION DURING EMBRYONIC DEVELOPMENT · Penn
Tried and failed
Automated intensity thresholding segmentation applied to subcellular focal adhesion quantification. Outcome: worse than baseline. Reason: Batch thresholding masks failed to detect features accurately, producing high false-negative rates versus manual quantification
Mechanotaxis and mechanoresistance in cancer · Imperial
Tried and failed
fixed-threshold normalized difference index segmentation applied to surface water mapping across regions. Outcome: did not generalise. Reason: a single threshold caused overestimation at some sites while completely omitting target features at others
Tried and failed
heuristic contrast adjustment and thresholding preprocessing applied to grayscale electron microscopy segmentation. Outcome: worse than baseline. Reason: degraded image quality and disrupted deep learning feature extraction
Deep Learning Approach for Cell Nuclear Pore Detection and Quantification over High Resolution 3D Data · Virginia Tech
Considered and rejected
Considered and rejected: Rejected global and adaptive threshold-based segmentation baselines (e.g., Otsu's method) because histogram-based thresholding inherently fails to distinguish root from non-root foreground noise.
ITErRoot: High Throughput Segmentation of 2-Dimensional Root System Architecture · HARVEST
Considered and rejected
Considered and rejected: Thresholding pixel values alone for laser line extraction (lacked repeatable precision required for structured light sensing, necessitating regional segmentation / binary image conversion).
Volumetric flow monitoring of biomass through an industrial grinder · Iowa State
Considered and rejected
Considered and rejected: Rejected using global thresholding segmentation due to poor multiclass classification compared to watershed and Bayesian/machine-learning approaches
Wettability characterisation of sandstone and carbonate rocks using X-ray micro-CT imaging · Imperial
Considered and rejected
Considered and rejected: Rejected purely automated thresholding (Otsu/RenyiEntropy) for final segmentation because overestimation and rigidity included background and cell bodies instead of true synaptic puncta.
Lack of spatial, multimodal, or temporal context degrades segmentation accuracy
Omitting elevation models, historical viewpoints, high-contrast modalities, or bidirectional volumetric context deprived models of critical structural priors. These localized or unimodal approaches suffered from perspective ambiguities, high false positive rates, and truncated peripheral context.
Tried and failed
omitting digital elevation data in visual segmentation applied to aerial imagery semantic segmentation. Outcome: worse than baseline. Reason: lacks crucial topographic priors, causing false positive segmentations above natural elevation thresholds
Monitoring and understanding treeline dynamics in the Swiss Alps from 80 years of aerial imagery · EPFL
Tried and failed
omitting specialized high-contrast input modalities in segmentation applied to cortical lesion detection. Outcome: worse than baseline. Reason: standard input sequences lacked sufficient contrast to resolve subtle cortical boundaries without specialized imaging sequences
Deep learning-based analysis of multiple sclerosis lesions with high and ultra-high field MRI · EPFL
Tried and failed
fixed margin bounding box cropping and resizing applied to semantic image segmentation. Outcome: worse than baseline. Reason: fixed margin crops either truncate peripheral features or introduce irrelevant distracting background context
Deep face tracking and parsing in the wild · Imperial
Considered and rejected
Considered and rejected: Single-scan batch inference for online segmentation was rejected due to lower accuracy from lacking historical viewpoint context.
3D Segmentation and Damage Analysis from Robotic Scans of Disaster Sites · Georgia Tech
Tried and failed
fine-tuning 2D models for volumetric data applied to 3D medical image segmentation. Outcome: worse than baseline. Reason: lack of bidirectional volumetric context across slices
AI Systems for Understanding and Grounding Radiology Reports · Harvard
Considered and rejected
Considered and rejected: Rejected standard CNN local feature extraction alone for overlap segmentation because it loses global context, resulting in subpar segmentation
Cervical Cell Separation using Deep Learning Techniques · Carleton University Institutional Repository
Tried and failed
pure vision transformer with smaller patch size applied to 3D spatiotemporal microstructural degradation prediction. Outcome: worse than baseline. Reason: Lacks inductive spatial biases of CNNs for complex 3D volumetric sequences
TransVNet: Predicting bone degradation using ViT and virtual dataset of cellular microstructures · Iowa State
Tried and failed
deep feature classification on unmasked raw images applied to pairwise visual disambiguation. Outcome: worse than baseline. Reason: pretrained visual features failed to distinguish ambiguous relationships without explicit foreground segmentation masks
Pushing the Boundaries of 3D Spatial Understanding · Cornell
Tried and failed
unimodal visual feature segmentation applied to spatial accident risk prediction. Outcome: worse than baseline. Reason: spatial dispersion and high false positive rates without auxiliary structural and dynamic context
Enhancing autonomous vehicle decision-making through scenario-based traffic rule integration · Imperial
Automated segmentation pipelines fall short of manual or semi-automated expert workflows
Fully automated deep learning systems exhibited inconsistent contouring failures, lesion spiculation biases, and error propagation compared to expert manual masking. These models struggled to encode subjective clinical expertise and institutional protocol differences, forcing researchers to retain manual intervention.
Lost to a baseline
Active contour segmentation was more biased than commercial semi-automatic segmentation on low-spiculation lesions at 2.5 mm slice thickness
Truth-based Radiomics for Prediction of Lung Cancer Prognosis · DukeSpace
Tried and failed
deep learning auto-segmentation for automated quality assurance applied to clinical target volume delineation. Outcome: did not generalise. Reason: Protocol variability, error propagation from upstream inputs, and inability to encode subjective clinical judgment.
Lost to a baseline
Def-RgDL SMG segmentation (DSC 53.8%) was inferior to DL alone without guidance (DSC 63.5%) in a case where registration propagated a false vessel contour.
Efficient and Intelligent Radiotherapy Planning and Adaptation · DSpace at UTSWMED
Lost to a baseline
Manual clover dry matter fraction prediction from Hansen et al. [101] achieved 7.8% standard deviation, surpassing the thesis's ERFNet grass/legumes segmentation in complex outdoor lighting
Mobile vision system for estimation of soil and plant properties Mobiles Bildverarbeitungssystem für die Schätzung von Boden- und Pflanzeneigenschaften · DSpace-CRIS at TU Wien
Lost to a baseline
Manual full-organ segmentation outperformed the simpler, faster ROI-based baseline method in test-retest repeatability across nearly all sequences.
Multiparametrische Magnetresonanztomographie der Nieren in einer prospektiven Probanden- und Patientenstudie: Retest-Reliabilität funktioneller Gewebeparameter und Evaluation einer Deep Learning-basierten Organsegmentierung · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Excluded automated computer-vision machine learning classification in favor of manual human coding due to current ML limitations in recognizing complex semantic scene context.
Lost to a baseline
Automated nnU-NET SAT segmentation suffered complete or partial failure in 13% of datasets, requiring manual correction compared to reliable manual masking.
Magnetic Resonance Imaging and Spectroscopy Methods for Studying Obesity: Applications for Bariatric Surgery · University of Nottingham Repository
Lost to a baseline
CardioINSIGHT automated CT segmentation inconsistently failed compared to manual multi-slice 2D contour interpolation.
Feasibility of improving risk stratification in the inherited cardiac conditions · Imperial
Left open by the authors
Problems the authors named and did not get to.
Left open
Benchmark TEDM against foundation models like SAM for semi-supervised medical image segmentation under domain shift and out-of-distribution conditions. Blocker: None
Left open
Evaluate and improve unsupervised instance segmentation failure cases caused by background clutter and edge ambiguity. Blocker: None
Enhancing Unsupervised Instance Segmentation with Exemplars · Carleton University Institutional Repository
Left open
Develop advanced post-processing techniques specifically tailored for gradient-based weakly-supervised semantic segmentation methods. Blocker: The thesis provides no specific design, algorithm, or mathematical formulation for the intended post-processing techniques.
Learning without Expert Labels for Multimodal Data · Virginia Tech
Left open
Implement and evaluate alternative classification algorithms beyond K-NN for historical document character classification in the press variant identification pipeline. Blocker: Lack of the specific historical print dataset and upstream segmentation pipeline output used in the thesis.
Automated identification of press variants in old documents · De Montfort Open Research Archive (DORA)
Left open
Develop physiology-aware machine learning models to reduce segmentation errors in multiplexed tissue images. Blocker: Lacks specific definition, architecture, or formulation for what constitutes 'physiology-aware' models
Single-cell Methods and Spatial Analysis for Highly Multiplexed Tissue Images · Harvard
Left open
Develop generalist deep learning segmentation architectures beyond U-Net to segment multiple cellular compartments across imaging modalities. Blocker: Lacks specific architectural design requirements, target benchmark datasets, and evaluation metrics
Grounded: Inference via Local Signals & Learned Representations by Organic & Artificial Systems · Harvard
Left open
Extend weakly supervised satellite semantic segmentation to self-supervised frameworks with automatic error detection for refining low-resolution labels. Blocker: None
Weak-Supervised Deep Learning Methods for the Analysis of Multi-Source Satellite Remote Sensing Images · IRIS - UNITN - prod
Left open
Combine pixel-wise, region-wise, and fuzzy boundary-wise loss functions and evaluate segmentation performance across standard metrics. Blocker: None
Incorporating fuzzy-based methods to deep learning models for semantic segmentation · University of Nottingham Repository
Left open
Develop a transfer-learning module to generalize brain metastasis segmentation across varied institutional MRI acquisition protocols. Blocker: Access to diverse multi-institutional brain metastasis MRI datasets with expert segmentations
Advancing Radiotherapy Treatment Through Artificial Intelligence-Driven Approaches · DSpace at UTSWMED
Left open
Adapt the PrinCut-Auto segmentation framework to work across multiple imaging modalities such as MRI, two-photon, and confocal microscopy. Blocker: Lacks specific algorithmic techniques or concrete benchmarks for achieving generalizability across modalities
Animal Internal Motion Analysis with Unsupervised Machine Learning Methods · Virginia Tech
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.