Chapter Four · failure evidence
What Natural Language Processing got wrong, from 44 dissertations
The evaluated doctoral records examine a range of natural language processing techniques, from classical lexical matching to large language models across diverse operational domains. Many advanced language models struggled with domain shifts, task-specific constraints, and competitive baselines, while simpler models frequently encountered fundamental representational limitations. These records come from PhD theses at 20 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Complex language models and dense embeddings underperformed simpler baselines on classification and extraction tasks
Fine-tuned or zero-shot large language models and Sentence-BERT embeddings trailed smaller encoder architectures, simple tokenizers, or classical machine learning models on token extraction and text classification. Several prompting and embedding approaches failed to exceed baseline performance or produced elevated false positive rates across clinical, legal, and academic datasets.
Tried and failed
direct zero-shot prompting of large language models applied to clinical note information extraction. Outcome: did not generalise. Reason: failed to adapt and performed poorly on advanced reasoning models without structured reasoning or fine-tuning
Machine Learning Approaches for Drug Combination Discovery · Cornell
Tried and failed
Sentence-BERT embeddings for textual features applied to tabular text feature representation. Outcome: worse than baseline. Reason: Trailed simple tokenization accuracy in predictive modeling
Tried and failed
fine-tuned large language models for sequence extraction applied to named entity recognition. Outcome: worse than baseline. Reason: Generative LLMs underperformed specialized smaller encoder models like BERT on token-level span classification tasks.
Consistency-aware and LLM-assisted methods for named entity recognition · Iowa State
Lost to a baseline
General language representation from LLAMA-65B embedding (AUROC 0.53–0.68) performed worse or comparable to simple utterance-level RoBERTa sentiment analysis (AUROC 0.63–0.69)
Multimodal assessment of neuropsychiatric disorders using audiovisual recordings · Georgia Tech
Tried and failed
prompt engineering and fine-tuning large language models applied to legal statutory ambiguity identification. Outcome: no signal. Reason: models failed to surpass random chance performance across all prompting strategies and fine-tuning variants
Disambiguating Large Language Model Performance on the Ambiguities of Law, Reasoning, and the Future · Harvard
Lost to a baseline
Transformer language model reaction embeddings failed to outperform Random Forest, Extreme Gradient Boosting, or Feedforward Neural Networks using standard fingerprint or DFT features.
Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT
Tried and failed
prompt-based zero-shot large language model classification applied to student text behavior classification. Outcome: worse than baseline. Reason: produced substantially lower AUC and more than double the false positives compared to embedding-based classifiers
Pretrained models suffered from domain mismatch and representation gaps across specialized formats
Language models pretrained on general or formal corpora failed when transferred to informal patient text, unseen programming languages, or specialized financial news. Direct embedding of raw code or mathematical formulas also faced substantial representation gaps with natural language specifications, requiring auxiliary descriptive text.
Tried and failed
domain-specific pretrained language model applied to patient-authored text classification. Outcome: worse than baseline. Reason: linguistic mismatch between formal clinical training corpora and informal patient-authored text
Developing Novel NLP-Enabled Timely and Accurate Decision Making for Precision Medicine · Georgia Tech
Considered and rejected
Considered and rejected: Rejected training BERT from scratch on raw security logs without natural language or code pre-training
Language Models and Cybersecurity - Applications and Current Limits · IRIS - POLITO - prod
Tried and failed
distilled general language model sentiment analysis applied to domain-specific financial news sentiment classification. Outcome: worse than baseline. Reason: produced poor directional accuracy and negatively correlated with domain-specific models during divergence periods
An Analysis of the Impact of News-Media Sentiment on Cryptocurrency Pricing · UT Austin
Tried and failed
Domain-adapted language model for text sentiment features applied to downstream financial market direction prediction. Outcome: worse than baseline. Reason: Domain-specific fine-tuning added computational cost without improving downstream task accuracy over general-purpose language models
Tried and failed
pretrained code language model embeddings applied to unseen programming language code retrieval. Outcome: did not generalise. Reason: the model was not pretrained on the target programming language
Structure-aware graph representation learning using graph neural networks · Iowa State
Considered and rejected
Considered and rejected: Rejected using raw source code embeddings directly to match natural language misuse rules due to representation gap; replaced code with LLM-generated natural language summaries.
API Utility Enhancement: From Traditional Software to Deep Learning Frameworks · YorkSpace
Considered and rejected
Considered and rejected: Rejected using raw LaTeX formulas directly for dense retrieval, finding that LLM-generated natural-language slogans significantly improve retrieval quality
Neural Methods for Plasma Simulation and Mathematical Formalization · ResearchWorks
Surface lexical features and rule-based methods failed to capture semantic nuance and technical terminology
Rule-based systems and general readability metrics struggled with complex clinical negation and domain-specific medical vocabulary. Surface word features and lexical overlap metrics showed weak correlation with classification accuracy and failed to provide reliable error detection or population risk stratification.
Tried and failed
rule-based natural language processing applied to clinical report diagnosis extraction. Reason: failed to correctly handle complex clinical negation, causing poor specificity and low F1 score
Tried and failed
General-domain readability metrics and span complexity models applied to medical text simplification evaluation. Outcome: did not generalise. Reason: Lexical metrics and general-corpus models fail to capture specialized technical and clinical vocabulary needs
Studying Text Revision in Scientific Writing · Georgia Tech
Considered and rejected
Considered and rejected: Rejected using EHR free text/unstructured clinical notes via natural language processing due to unrecognized false negatives, potential false positives, and impracticality for large population risk stratification.
Considered and rejected
Considered and rejected: Automated natural language systems for initial error categorization, rejected because they cannot detect subtle text nuances or apply clinical judgment.
Improving the Error Review Process for Incident Reports at UT Southwestern Through the Use of a Standardized Taxonomy Tool · DSpace at UTSWMED
Tried and failed
direct lexical word features in discrete choice applied to social media engagement prediction. Outcome: worse than baseline. Reason: yielded lower prediction accuracy and poorer interpretability than aggregate topic features
Tried and failed
TF-IDF cosine distance and lexical overlap metrics applied to predicting neural classifier performance. Outcome: no signal. Reason: lexical similarity metrics showed weak and inconsistent correlation with classification success across diverse texts
The computational bridge: Interfacing theory and data in cognitive science · Cornell
Architectural context limits and shallow networks failed to handle long sequences and compositional horizons
Fixed context windows and sequence length constraints hindered transformer fine-tuning on long regulatory documents and prevented short-sequence models from generalizing to longer texts. Recurrent and feed-forward architectures were similarly rejected or failed when tasked with unconstrained summarization or long compositional navigation environments.
Tried and failed
fine-tuning transformer language models applied to long-document regulatory text classification. Reason: token context window constraints relative to long document lengths and lack of content-specific sentence summarization
Managerial oversight : how Congress solves problems in the age of regulatory governance · UT Austin
Tried and failed
open-source small language model summarization applied to unconstrained open-ended user data. Outcome: worse than baseline. Reason: struggled with unconstrained summarization tasks compared to larger proprietary baselines
Language Modeling with Few-shot Language Feedback · Georgia Tech
Tried and failed
neural models trained on short sequences applied to regular language classification. Outcome: did not generalise. Reason: training distribution on shorter sequences fails to generalize to longer evaluation sequences
COMPOSITIONAL GENERALIZATION IN INSTRUCTION FOLLOWING TASKS · Penn
Considered and rejected
Considered and rejected: Rejected using feed-forward networks directly for language modelling and sentence classification due to poor performance compared to sequential/recurrent architectures.
Considered and rejected
Considered and rejected: Rejected standard end-to-end Sequence-to-Sequence CNN-LSTM models for vision-language navigation because they memorize trajectories and suffer poor generalization (~2-4% success) in unseen environments with long compositional horizons.
Information fusion for decision making · Iowa State
Generative language models produced generic, ungrounded, or non-actionable outputs in evaluative and reporting tasks
Prompted and fine-tuned models omitted core numerical metrics in clinical reports, lost citation grounding, or outputted generic solutions instead of identifying specific student errors. In educational and clinical settings, generated critiques were repetitive and failed to provide actionable or beneficial guidance.
Tried and failed
prompting large language models for domain summarization applied to clinical report generation. Reason: inconsistent incorporation of core numerical metrics led to clinically inappropriate recommendations
Precision Medicine in Diabetes Using Continuous Glucose Monitoring · MIT
Tried and failed
supervised fine-tuning for structured reasoning applied to large language models for clinical grounding. Outcome: worse than baseline. Reason: optimizing for diagnostic correctness collapsed citation grounding and degraded structured output format compliance
Causal Inference and Evidence-Grounded Language Models for Trustworthy Personalized Clinical Decision Support · Georgia Tech
Tried and failed
prompting large language models for critique generation applied to evaluating complex student writing. Outcome: worse than baseline. Reason: generated critiques were vague, generic, and repetitive compared to human instructor assessments
Addressing teachers’ needs for integrating generative AI into a pedagogical feedback model · Iowa State
Tried and failed
large language models for automated grading applied to student engineering derivations and solutions. Reason: Outputted full solutions and generic suggestions rather than pinpointing specific error locations in student work
Developing an Automated Practice Environment and Feedback Engine for Guided Instruction in the Deformable Bodies Course · Virginia Tech
Tried and failed
automated natural language feedback and grading applied to student homework free responses. Reason: students did not perceive automated neural network grading access as sufficiently beneficial
Critical thinking in automated homework · Texas Tech
Compression and federated optimization techniques degraded internal model representations and training dynamics
Layerwise hidden state matching and tensor-train decomposition degraded lower-level feature representations and overly constrained student model learning during compression. In federated language modeling, momentum-based optimizers underperformed non-momentum baselines by destroying mini-batch gradient token sparsity.
Tried and failed
layerwise hidden representation matching in knowledge distillation applied to cross-architecture language model compression. Outcome: worse than baseline. Reason: unfiltered layerwise intermediate state matching overly constrained student learning compared to output probability distillation
On Parameter Efficiency of Neural Language Models · Georgia Tech
Tried and failed
tensor-train decomposition for neural network compression applied to dense language model layers. Outcome: worse than baseline. Reason: degrades broad, weak lower-level feature representation required for commonsense reasoning tasks
Generative language model compression with algebraic approaches · Imperial
Lost to a baseline
In language modeling (Shakespeare charRNN and StackOverflow LSTM), momentum-based federated optimizers (FedAvgMom, MimeMom, MimeAdam) underperformed non-momentum baselines (FedAvgSGD, FedAvgAdagrad) because dense momentum updates destroy mini-batch gradient token sparsity.
Uncurated or noisy training corpora degraded generative model outputs
Fine-tuning language models on scientific papers with poorly formatted equations resulted in noisy delimiter generation. Similarly, training models on raw social media data produced rude and non-factual text that lowered generation quality across standard evaluation metrics.
Tried and failed
instruction tuning large language models applied to scientific document summarization. Reason: training data contained poorly formatted equations and notations causing noisy generated delimiters
Improving Access to ETD Elements Through Chapter Categorization and Summarization · Virginia Tech
Tried and failed
training language models on raw social media data applied to counter-misinformation response generation. Outcome: worse than baseline. Reason: raw uncurated responses were frequently rude and lacked factual evidence, degrading output quality across metrics
Leveraging AI to Combat Misinformation by Empowering Crowds and Evaluating Detectors · Georgia Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Extend the clinical graph attention network to incorporate unstructured clinical notes using large language model embeddings alongside structured patient features. Blocker: Access to credentialed clinical datasets containing linked tabular features and free-text notes (e.g., PhysioNet/MIMIC)
Essays on developing artificial intelligence solutions for patient-centered healthcare delivery · UT Austin
Left open
Use natural language processing on inpatient clinical notes to extract consult questions and identify resulting treatment changes. Blocker: Requires access to private electronic health record (EHR) clinical notes data
Essays on Physician Consults · Penn
Left open
Apply textual sentiment analysis to alternative macroeconomic texts to predict other macroeconomic indicators. Blocker: None
The Usefulness of Textual Sentiment Analysis for Macroeconomics · UT Austin
Left open
Compare narrative readability and tone between US GAAP and IFRS annual reports across regulatory changes using existing textual analysis metrics. Blocker: None
Narratives in corporate annual reports: the drivers and impacts of narrative's readability and tone. · Cranfield
Left open
Analyze how stylistic and lexical properties of natural language feedback, such as emotional tone, affect LLM performance across different models. Blocker: None
Language Modeling with Few-Shot Language Feedback · Georgia Tech
Left open
Learn linear transformations across embedding spaces to enable domain transfer of self-regulated learning classification between chemistry and formal logic text datasets. Blocker: Requires the private student textual learning interaction datasets from the specific chemistry and formal logic domains used in the thesis.
Left open
Apply advanced large language models and hybrid quantum-classical algorithms to source code representations for software defect prediction. Blocker: The unfinished work lacks concrete specifications, specific model architectures, or target benchmarks
Left open
Apply neuron redundancy insights to prune, distill, and optimize inference costs of code-trained language models. Blocker: None
Analyzing redundancy in code-trained language models · Iowa State
Left open
Develop training algorithms beyond standard maximum likelihood estimation and non-autoregressive methods for neural language generation models. Blocker: The goal is a broad research direction without concrete specifications, objectives, or defined algorithmic formulations
Towards a Deeper Understanding of Neural Language Generation · MIT
Left open
Apply lifelong learning fine-tuning techniques to large language models for generating automated code review comments across evolving codebases. Blocker: None
Towards Sustainable AI for Continuous Integration Quality Gates · Queens University Institutional Repository
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.