Chapter Four · failure evidence

What Natural Language Processing got wrong, from 44 dissertations

The evaluated doctoral records examine a range of natural language processing techniques, from classical lexical matching to large language models across diverse operational domains. Many advanced language models struggled with domain shifts, task-specific constraints, and competitive baselines, while simpler models frequently encountered fundamental representational limitations. These records come from PhD theses at 20 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Complex language models and dense embeddings underperformed simpler baselines on classification and extraction tasks

7 theses · 7 institutions

Fine-tuned or zero-shot large language models and Sentence-BERT embeddings trailed smaller encoder architectures, simple tokenizers, or classical machine learning models on token extraction and text classification. Several prompting and embedding approaches failed to exceed baseline performance or produced elevated false positive rates across clinical, legal, and academic datasets.

Tried and failed

direct zero-shot prompting of large language models applied to clinical note information extraction. Outcome: did not generalise. Reason: failed to adapt and performed poorly on advanced reasoning models without structured reasoning or fine-tuning

Machine Learning Approaches for Drug Combination Discovery · Cornell

Tried and failed

Sentence-BERT embeddings for textual features applied to tabular text feature representation. Outcome: worse than baseline. Reason: Trailed simple tokenization accuracy in predictive modeling

Machine Learning-Based Predictive Modeling for OTC and exercise recommendation for knee joint pain with long-tail classification · Texas Tech

Tried and failed

fine-tuned large language models for sequence extraction applied to named entity recognition. Outcome: worse than baseline. Reason: Generative LLMs underperformed specialized smaller encoder models like BERT on token-level span classification tasks.

Consistency-aware and LLM-assisted methods for named entity recognition · Iowa State

Lost to a baseline

General language representation from LLAMA-65B embedding (AUROC 0.53–0.68) performed worse or comparable to simple utterance-level RoBERTa sentiment analysis (AUROC 0.63–0.69)

Multimodal assessment of neuropsychiatric disorders using audiovisual recordings · Georgia Tech

Tried and failed

prompt engineering and fine-tuning large language models applied to legal statutory ambiguity identification. Outcome: no signal. Reason: models failed to surpass random chance performance across all prompting strategies and fine-tuning variants

Disambiguating Large Language Model Performance on the Ambiguities of Law, Reasoning, and the Future · Harvard

Lost to a baseline

Transformer language model reaction embeddings failed to outperform Random Forest, Extreme Gradient Boosting, or Feedforward Neural Networks using standard fingerprint or DFT features.

Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT

Tried and failed

prompt-based zero-shot large language model classification applied to student text behavior classification. Outcome: worse than baseline. Reason: produced substantially lower AUC and more than double the false positives compared to embedding-based classifiers

MEASURING AND UNDERSTANDING STUDENTS’ SELF-REGULATED LEARNING IN TEXTUAL DATA IN COMPUTER-BASED LEARNING ENVIRONMENTS · Penn

Pretrained models suffered from domain mismatch and representation gaps across specialized formats

7 theses · 7 institutions

Language models pretrained on general or formal corpora failed when transferred to informal patient text, unseen programming languages, or specialized financial news. Direct embedding of raw code or mathematical formulas also faced substantial representation gaps with natural language specifications, requiring auxiliary descriptive text.

Tried and failed

domain-specific pretrained language model applied to patient-authored text classification. Outcome: worse than baseline. Reason: linguistic mismatch between formal clinical training corpora and informal patient-authored text

Developing Novel NLP-Enabled Timely and Accurate Decision Making for Precision Medicine · Georgia Tech

Considered and rejected

Considered and rejected: Rejected training BERT from scratch on raw security logs without natural language or code pre-training

Language Models and Cybersecurity - Applications and Current Limits · IRIS - POLITO - prod

Tried and failed

distilled general language model sentiment analysis applied to domain-specific financial news sentiment classification. Outcome: worse than baseline. Reason: produced poor directional accuracy and negatively correlated with domain-specific models during divergence periods

An Analysis of the Impact of News-Media Sentiment on Cryptocurrency Pricing · UT Austin

Tried and failed

Domain-adapted language model for text sentiment features applied to downstream financial market direction prediction. Outcome: worse than baseline. Reason: Domain-specific fine-tuning added computational cost without improving downstream task accuracy over general-purpose language models

Machine Learning for Financial Market Forecasting · Harvard

Tried and failed

pretrained code language model embeddings applied to unseen programming language code retrieval. Outcome: did not generalise. Reason: the model was not pretrained on the target programming language

Structure-aware graph representation learning using graph neural networks · Iowa State

Considered and rejected

Considered and rejected: Rejected using raw source code embeddings directly to match natural language misuse rules due to representation gap; replaced code with LLM-generated natural language summaries.

API Utility Enhancement: From Traditional Software to Deep Learning Frameworks · YorkSpace

Considered and rejected

Considered and rejected: Rejected using raw LaTeX formulas directly for dense retrieval, finding that LLM-generated natural-language slogans significantly improve retrieval quality

Neural Methods for Plasma Simulation and Mathematical Formalization · ResearchWorks

Surface lexical features and rule-based methods failed to capture semantic nuance and technical terminology

6 theses · 6 institutions

Rule-based systems and general readability metrics struggled with complex clinical negation and domain-specific medical vocabulary. Surface word features and lexical overlap metrics showed weak correlation with classification accuracy and failed to provide reliable error detection or population risk stratification.

Tried and failed

rule-based natural language processing applied to clinical report diagnosis extraction. Reason: failed to correctly handle complex clinical negation, causing poor specificity and low F1 score

Artificial intelligence applications for the acquisition, analysis and reporting of cardiovascular magnetic resonance imaging · Imperial

Tried and failed

General-domain readability metrics and span complexity models applied to medical text simplification evaluation. Outcome: did not generalise. Reason: Lexical metrics and general-corpus models fail to capture specialized technical and clinical vocabulary needs

Studying Text Revision in Scientific Writing · Georgia Tech

Considered and rejected

Considered and rejected: Rejected using EHR free text/unstructured clinical notes via natural language processing due to unrecognized false negatives, potential false positives, and impracticality for large population risk stratification.

ASSESSING THE EFFECT OF COMPUTABLE PHENOTYPES IN PREDICTING HEALTHCARE UTILIZATION AND POTENTIAL RACIAL DISPARITIES IN IDENTIFYING TYPE 2 DIABETES POPULATION · JScholarship

Considered and rejected

Considered and rejected: Automated natural language systems for initial error categorization, rejected because they cannot detect subtle text nuances or apply clinical judgment.

Improving the Error Review Process for Incident Reports at UT Southwestern Through the Use of a Standardized Taxonomy Tool · DSpace at UTSWMED

Tried and failed

direct lexical word features in discrete choice applied to social media engagement prediction. Outcome: worse than baseline. Reason: yielded lower prediction accuracy and poorer interpretability than aggregate topic features

How Words Move Hearts: Interpretable Machine Learning Models of Bias, Engagement, and Influence in Socio-Political Systems · EPFL

Tried and failed

TF-IDF cosine distance and lexical overlap metrics applied to predicting neural classifier performance. Outcome: no signal. Reason: lexical similarity metrics showed weak and inconsistent correlation with classification success across diverse texts

The computational bridge: Interfacing theory and data in cognitive science · Cornell

Architectural context limits and shallow networks failed to handle long sequences and compositional horizons

5 theses · 5 institutions

Fixed context windows and sequence length constraints hindered transformer fine-tuning on long regulatory documents and prevented short-sequence models from generalizing to longer texts. Recurrent and feed-forward architectures were similarly rejected or failed when tasked with unconstrained summarization or long compositional navigation environments.

Tried and failed

fine-tuning transformer language models applied to long-document regulatory text classification. Reason: token context window constraints relative to long document lengths and lack of content-specific sentence summarization

Managerial oversight : how Congress solves problems in the age of regulatory governance · UT Austin

Tried and failed

open-source small language model summarization applied to unconstrained open-ended user data. Outcome: worse than baseline. Reason: struggled with unconstrained summarization tasks compared to larger proprietary baselines

Language Modeling with Few-shot Language Feedback · Georgia Tech

Tried and failed

neural models trained on short sequences applied to regular language classification. Outcome: did not generalise. Reason: training distribution on shorter sequences fails to generalize to longer evaluation sequences

COMPOSITIONAL GENERALIZATION IN INSTRUCTION FOLLOWING TASKS · Penn

Considered and rejected

Considered and rejected: Rejected using feed-forward networks directly for language modelling and sentence classification due to poor performance compared to sequential/recurrent architectures.

Categorical tools for natural language processing · Oxford

Considered and rejected

Considered and rejected: Rejected standard end-to-end Sequence-to-Sequence CNN-LSTM models for vision-language navigation because they memorize trajectories and suffer poor generalization (~2-4% success) in unseen environments with long compositional horizons.

Information fusion for decision making · Iowa State

Generative language models produced generic, ungrounded, or non-actionable outputs in evaluative and reporting tasks

5 theses · 5 institutions

Prompted and fine-tuned models omitted core numerical metrics in clinical reports, lost citation grounding, or outputted generic solutions instead of identifying specific student errors. In educational and clinical settings, generated critiques were repetitive and failed to provide actionable or beneficial guidance.

Tried and failed

prompting large language models for domain summarization applied to clinical report generation. Reason: inconsistent incorporation of core numerical metrics led to clinically inappropriate recommendations

Precision Medicine in Diabetes Using Continuous Glucose Monitoring · MIT

Tried and failed

supervised fine-tuning for structured reasoning applied to large language models for clinical grounding. Outcome: worse than baseline. Reason: optimizing for diagnostic correctness collapsed citation grounding and degraded structured output format compliance

Causal Inference and Evidence-Grounded Language Models for Trustworthy Personalized Clinical Decision Support · Georgia Tech

Tried and failed

prompting large language models for critique generation applied to evaluating complex student writing. Outcome: worse than baseline. Reason: generated critiques were vague, generic, and repetitive compared to human instructor assessments

Addressing teachers’ needs for integrating generative AI into a pedagogical feedback model · Iowa State

Tried and failed

large language models for automated grading applied to student engineering derivations and solutions. Reason: Outputted full solutions and generic suggestions rather than pinpointing specific error locations in student work

Developing an Automated Practice Environment and Feedback Engine for Guided Instruction in the Deformable Bodies Course · Virginia Tech

Tried and failed

automated natural language feedback and grading applied to student homework free responses. Reason: students did not perceive automated neural network grading access as sufficiently beneficial

Critical thinking in automated homework · Texas Tech

Compression and federated optimization techniques degraded internal model representations and training dynamics

3 theses · 3 institutions

Layerwise hidden state matching and tensor-train decomposition degraded lower-level feature representations and overly constrained student model learning during compression. In federated language modeling, momentum-based optimizers underperformed non-momentum baselines by destroying mini-batch gradient token sparsity.

Tried and failed

layerwise hidden representation matching in knowledge distillation applied to cross-architecture language model compression. Outcome: worse than baseline. Reason: unfiltered layerwise intermediate state matching overly constrained student learning compared to output probability distillation

On Parameter Efficiency of Neural Language Models · Georgia Tech

Tried and failed

tensor-train decomposition for neural network compression applied to dense language model layers. Outcome: worse than baseline. Reason: degrades broad, weak lower-level feature representation required for commonsense reasoning tasks

Generative language model compression with algebraic approaches · Imperial

Lost to a baseline

In language modeling (Shakespeare charRNN and StackOverflow LSTM), momentum-based federated optimizers (FedAvgMom, MimeMom, MimeAdam) underperformed non-momentum baselines (FedAvgSGD, FedAvgAdagrad) because dense momentum updates destroy mini-batch gradient token sparsity.

Optimization methods for collaborative learning · EPFL

Uncurated or noisy training corpora degraded generative model outputs

2 theses · 2 institutions

Fine-tuning language models on scientific papers with poorly formatted equations resulted in noisy delimiter generation. Similarly, training models on raw social media data produced rude and non-factual text that lowered generation quality across standard evaluation metrics.

Tried and failed

instruction tuning large language models applied to scientific document summarization. Reason: training data contained poorly formatted equations and notations causing noisy generated delimiters

Improving Access to ETD Elements Through Chapter Categorization and Summarization · Virginia Tech

Tried and failed

training language models on raw social media data applied to counter-misinformation response generation. Outcome: worse than baseline. Reason: raw uncurated responses were frequently rude and lacked factual evidence, degrading output quality across metrics

Leveraging AI to Combat Misinformation by Empowering Crowds and Evaluating Detectors · Georgia Tech

Left open by the authors

Problems the authors named and did not get to.

Left open

Extend the clinical graph attention network to incorporate unstructured clinical notes using large language model embeddings alongside structured patient features. Blocker: Access to credentialed clinical datasets containing linked tabular features and free-text notes (e.g., PhysioNet/MIMIC)

Essays on developing artificial intelligence solutions for patient-centered healthcare delivery · UT Austin

Left open

Use natural language processing on inpatient clinical notes to extract consult questions and identify resulting treatment changes. Blocker: Requires access to private electronic health record (EHR) clinical notes data

Essays on Physician Consults · Penn

Left open

Apply textual sentiment analysis to alternative macroeconomic texts to predict other macroeconomic indicators. Blocker: None

The Usefulness of Textual Sentiment Analysis for Macroeconomics · UT Austin

Left open

Compare narrative readability and tone between US GAAP and IFRS annual reports across regulatory changes using existing textual analysis metrics. Blocker: None

Narratives in corporate annual reports: the drivers and impacts of narrative's readability and tone. · Cranfield

Left open

Analyze how stylistic and lexical properties of natural language feedback, such as emotional tone, affect LLM performance across different models. Blocker: None

Language Modeling with Few-Shot Language Feedback · Georgia Tech

Left open

Learn linear transformations across embedding spaces to enable domain transfer of self-regulated learning classification between chemistry and formal logic text datasets. Blocker: Requires the private student textual learning interaction datasets from the specific chemistry and formal logic domains used in the thesis.

MEASURING AND UNDERSTANDING STUDENTS’ SELF-REGULATED LEARNING IN TEXTUAL DATA IN COMPUTER-BASED LEARNING ENVIRONMENTS · Penn

Left open

Apply advanced large language models and hybrid quantum-classical algorithms to source code representations for software defect prediction. Blocker: The unfinished work lacks concrete specifications, specific model architectures, or target benchmarks

Enhancing Software Defect Prediction: Investigating Diverse Representations of Source Code as Feature Values in Classical and Quantum Machine Learning Approaches · HARVEST

Left open

Apply neuron redundancy insights to prune, distill, and optimize inference costs of code-trained language models. Blocker: None

Analyzing redundancy in code-trained language models · Iowa State

Left open

Develop training algorithms beyond standard maximum likelihood estimation and non-autoregressive methods for neural language generation models. Blocker: The goal is a broad research direction without concrete specifications, objectives, or defined algorithmic formulations

Towards a Deeper Understanding of Neural Language Generation · MIT

Left open

Apply lifelong learning fine-tuning techniques to large language models for generating automated code review comments across evolving codebases. Blocker: None

Towards Sustainable AI for Continuous Integration Quality Gates · Queens University Institutional Repository

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.