Chapter Four · failure evidence

What Rule-Based & Expert Systems got wrong, from 35 dissertations

The evaluated records demonstrate that rule-based systems and manual heuristics frequently falter when confronting linguistic variation, stochastic operating environments, and rigid threshold boundaries. In multiple domains, deterministic rules also suffer from maintenance bottlenecks, fail under distribution shifts, and lag behind statistical or machine learning baselines. These records come from PhD theses at 19 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Rule-based natural language processing fails to handle linguistic variation and complex syntax

9 theses · 8 institutions

Rigid syntactic patterns and regular expression heuristics fail to capture document layout variations, complex negation, and idiosyncratic semantic shifts. Handcrafted text rules and parsers achieve poor precision or recall compared to learned models and struggle with varied linguistic structures.

Tried and failed

rule-based natural language processing applied to clinical report diagnosis extraction. Reason: failed to correctly handle complex clinical negation, causing poor specificity and low F1 score

Artificial intelligence applications for the acquisition, analysis and reporting of cardiovascular magnetic resonance imaging · Imperial

Tried and failed

heuristic regular-expression rules applied to document metadata extraction. Outcome: did not generalise. Reason: rule-based heuristics failed to capture diverse document layouts and formatting variations across a larger corpus

Automatic Metadata Extraction Incorporating Visual Features from Scanned Electronic Theses and Dissertations 2021 ACM/IEEE JOINT CONFERENCE ON DIGITAL LIBRARIES (JCDL 2021) · Virginia Tech

Tried and failed

rule-based pattern matching NLP applied to inter-parameter dependency extraction from documentation. Outcome: no signal. Reason: rigid syntactic patterns failed to match diverse real-world natural language documentation structures

Effective Automation of Black-Box Testing for REST APIs with Machine Learning and Language Models · Georgia Tech

Tried and failed

rule-based text data augmentation applied to intent classification. Outcome: worse than baseline. Reason: heuristic word-level perturbations alter sentence semantics and label alignment

Lifelong Machine Learning with Data Efficiency and Knowledge Retention · EPFL

Tried and failed

rule-based compositional semantics and derivational morphology applied to compound participle meaning prediction. Outcome: did not generalise. Reason: rules failed to capture idiosyncratic semantic shifts and semi-productivity patterns

The semantics of past participles · UT Austin

Tried and failed

rule-based dependency parser with candidate scoring applied to natural language to formal specification translation. Reason: candidate completions lacking matching input categories were overly penalized by the negative scoring mechanism

Framework for Automatic Translation of Hardware Specifications Written in English to a Formal Language · Virginia Tech

Lost to a baseline

Rule-based baseline NLP (EasyCIE_GUI) achieved an F1 of only 0.43 and precision of 0.34 (at 0.59 recall), losing substantially to deep learning models (BiLSTM F1 0.64–0.66, precision 0.62–0.67).

Surgical Site Infection (SSI) Identification Across Multiple Facilities and Surgery Types Using Multimodal Data and Deep Learning · ResearchWorks

Considered and rejected

Considered and rejected: Rejected using only plain rule-based parsing in ChemDataExtractor due to brittle failure on minor grammatical variations and low inorganic recall (56%).

Magnetic and Superconducting Materials Discovery: Employing Data Science, Natural Language Processing and Machine Learning · Cambridge

Considered and rejected

Considered and rejected: Rejected pure rule-based comparative adjective/adverb parsing due to the complexity of building grammar rules for varied linguistic structures.

Compa: A Comparative Retrieval and Analytical Engine for Consumer Products · TXST Digital Repository

Deterministic rules and static heuristics struggle in stochastic and dynamic environments

6 theses · 4 institutions

Deterministic logical models and static heuristic thresholds break down when deployed in dynamic, volatile, or noisy environments. These systems fail to adapt across long time horizons and ignore baseline human or operational stochasticity.

Tried and failed

rule-based behavior inference applied to human response behavior under disruption. Reason: produced systematic overestimation by ignoring baseline stochasticity in human behavior

Toward a Resilient Public Transportation System: Effective Monitoring and Control under Service Disruptions · MIT

Tried and failed

deterministic rule-based logical models applied to human activity estimation. Outcome: worse than baseline. Reason: strict logical rules cannot handle noise and uncertainty compared to probabilistic modeling

Cognitive Human Activity and Plan Recognition for Human-Robot Collaboration · MIT

Considered and rejected

Considered and rejected: Rejected rule-based heuristics and MPC for long-term operational optimization due to lack of adaptability in dynamic environments and poor scaling over long-horizon stochastic tasks.

Edge Computing in Space: Design Optimization and Reinforcement Learning Scheduling of Onboard Computing Satellites · MIT

Lost to a baseline

simpler rule-based or gradient-based dispatch methods excelled over heuristic algorithms under stable conditions, only faltering in highly volatile environments

Optimization of Design Parameters for Last-Mile Delivery Drones · UT Austin

Tried and failed

static heuristic-based threshold switching rules applied to dynamic architectural protocol selection. Outcome: worse than baseline. Reason: simple thresholds fail to capture complex workload dynamics, yielding incorrect decisions in half the cases

AI-DRIVEN ADAPTIVE DISTRIBUTED SYSTEMS IN UNTRUSTED ENVIRONMENTS · Penn

Tried and failed

deterministic rule-based modeling of dynamic systems applied to agent survival simulation. Outcome: did not generalise. Reason: Failed to survive and became erratic outside a narrow payoff range

An adaptive agent-based multicriteria simulation system · Cranfield

Rigid thresholds and hardcoded boundaries cause false positives and poor boundary behavior

5 theses · 5 institutions

Handcrafted numerical thresholds and crisp conditional rules fail to handle boundary cases fairly and generate high rates of false positives. These rigid cutoff approaches lead to oversimplified models or unmaintainable rule sets that do not generalize across diverse styles.

Tried and failed

heuristic rule-based constraint extraction applied to circuit symmetry constraint detection. Outcome: worse than baseline. Reason: produced higher false positive rates and generated unnecessary constraints compared to learned representations

Layout automation for custom integrated circuits · UT Austin

Considered and rejected

Considered and rejected: Default high dependency thresholds in the Flexible Heuristics Miner algorithm were rejected because they led to oversimplified process models in low-structured domains.

Rezeptions- und Interpretationsprozesse von Lehrpersonen bei datengestützten Entscheidungen. Exploration und Förderung · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Crisp rule-based expert systems using rigid mathematical thresholds were rejected because they fail to handle boundary cases fairly and create overly complex, unmaintainable rule bases.

Strategie działania inteligentnych systemów wspierających kształcenie operujące na danych nieprecyzyjnych Strategies of operation for intelligent tutoring systems operating on imprecise data · AMUR - Repozytorium Uniwersytetu im. Adama Mickiewicza w Poz

Considered and rejected

Considered and rejected: Fixing static μ thresholds rule-based was rejected in favor of dynamically low-pass filtering synaptic weight efficacy

Building new memories using the past: Interplay between semantic and episodic memory in brains and machines · Imperial

Considered and rejected

Considered and rejected: Rejected classical rule-based conditional programming (hardcoded angle thresholds) because it produced false positives while moving/pretending to shoot and could not generalize across diverse shooting styles.

Live Perception and Real Time Motion Prediction with Deep Neural Networks and Machine Learning · Harvard

Manual rule systems suffer from poor maintainability and inability to learn or scale

5 theses · 5 institutions

Expert deduction engines and handcrafted anti-pattern heuristics cannot incorporate new observations or generalize to new problem domains. Adding criteria requires extensive manual reprogramming, while crowdsourced rules often encode shallow heuristics that quickly plateau.

Tried and failed

extracting decision rules from human-voted behavior applied to sequential decision-making in disrupted environments. Outcome: did not generalise. Reason: crowdsourced rules encoded shallow, suboptimal heuristics with plateauing performance gains

Managing The Gig Economy Via Behavioral And Operational Lenses · Penn

Considered and rejected

Considered and rejected: Rejected rule-based classifiers due to poor handling of continuous numerical attributes like GPA and GRE scores.

Towards Better Interpretability of Machine Learning-Based Decision Support Systems · DSpace at SUNY Buffalo

Tried and failed

rule-based anti-pattern heuristics for manual diagnosis applied to identifying algorithmic complexity vulnerabilities. Reason: developers could not accurately diagnose complex edge cases using manual pattern rules

Theory and Patterns for Avoiding Regex Denial of Service · Virginia Tech

Considered and rejected

Considered and rejected: Rejected fuzzy logic/AI rule-based expert systems due to inflexibility requiring full reprogramming for added criteria

A practical decision support tool for the design of automated manufacturing systems: incorporating human factors alongside other considerations in the design · Cranfield

Considered and rejected

Considered and rejected: Rejected rule-based deduction engines (e.g. from IoIF) because semantic rules cannot learn from new observations or scale across new problem domains.

MM-ADM: A model-based approach to multidisciplinary design to support automated decision-making · Georgia Tech

Heuristic decision rules and synthetic models fail under distribution shifts

5 theses · 5 institutions

Simple heuristics and synthetic rule-based error models fail to generalize to real error distributions or unrepresentative baselines. Heuristic decision support systems can increase user response latency without improving accuracy, while stopping heuristics miss human-understandable failure patterns.

Considered and rejected

Considered and rejected: Rejected automatically distinguishing between genuinely unknown biochemical gaps and reconstruction errors in dead-end tests due to poor accuracy of previous automated heuristics.

The fundamentals of genome-scale metabolic models and their application to the study of evolution and cancer · OpenBU

Tried and failed

heuristic training for counterfactual forecasting applied to decision making under unrepresentative baselines. Outcome: did not generalise. Reason: training failed when faced with unrepresentative baseline distributions or prospective conditional framing

FAST AND FRUGAL STATISTICAL HEURISTICS: TRANSFER OF COUNTERFACTUAL FORECASTING TRAINING WITHIN AND ACROSS DOMAINS · Penn

Considered and rejected

Considered and rejected: Rejected synthetic rule-based error injection models because their error distributions failed to generalize to real generation errors

Fine-grained evaluation for text summarization · UT Austin

Tried and failed

rule-based heuristic decision support system applied to human decision-making under degraded information. Outcome: worse than baseline. Reason: increased user response times without improving accuracy, especially for lower-performing users

Understanding and Supporting Decision Making in Denied and Degraded Environments · Georgia Tech

Considered and rejected

Considered and rejected: Rejected simple confidence-based early-stopping heuristics and generic upweighting methods for error mitigation because they fail to capture consistent, human-understandable failure modes.

A Data-Based Perspective on Model Reliability · MIT

Rule-based classifiers and filters underperform probabilistic and machine learning baselines

4 theses · 4 institutions

Rule-based classifiers and heuristic filters achieve lower accuracy than probabilistic models and decision trees on classification tasks. Simple heuristic mechanisms produce repetitive predictions and fail to capture complex causal topological interactions.

Lost to a baseline

List Recall at K=20 was only 0.046 for rule-based method vs 0.008 for the baseline, suffering from repetitive predictions across adjacent time windows.

Artificial Intelligence Methods and Evaluation Strategies for Detecting Future Customer Needs from User Generated Content · Research Repository UCD

Lost to a baseline

Heuristic rules-based algorithm achieved lower overall gait event identification accuracy across locomotion modes (94.87%) compared to the unsupervised BP-AR-HMM (99.6%).

Machine Learning and Wearable Sensors for the Estimation of Biomechanical Variables Outside the Laboratory · Scholars' Bank

Lost to a baseline

Rule-based PART classifier (0.5005 accuracy on investment cost) lost to C4.5 Trees J48 (0.571 average accuracy) and Naive Bayes (0.5091 accuracy on investment cost)

Renewable Energy Communities: A Preference Learning Approach to Evaluate Differential Participation in Cities · IRIS - POLITO - prod

Tried and failed

heuristic filter-based feature selection applied to transient stability classification. Outcome: worse than baseline. Reason: standard filters failed to capture complex causal topological interactions compared to Markov blanket selection

Topological changes in data-driven dynamic security assessment for power system control · Imperial

Left open by the authors

Problems the authors named and did not get to.

Left open

Extend the tree-based rule extraction heuristic to estimate continuous variables and generalize across different problem sizes. Blocker: Lack of specific target continuous variables, validation metrics, or concrete generalization methodology

A Machine Learning-Based Heuristic to Explain Game-Theoretic Models · Virginia Tech

Left open

Test whether fast-and-frugal counterfactual forecasting heuristics transfer to real-world scenarios lacking accessible ground truth. Blocker: No specific real-world domain, evaluation methodology, or benchmark dataset is defined for scenarios lacking ground truth

FAST AND FRUGAL STATISTICAL HEURISTICS: TRANSFER OF COUNTERFACTUAL FORECASTING TRAINING WITHIN AND ACROSS DOMAINS · Penn

Left open

Adapt the rule-extraction and hierarchical decision tree heuristic to explain policy functions in reinforcement learning models. Blocker: The proposal is an exploratory direction without a specified target RL environment, policy model, or evaluation benchmark.

A Machine Learning-Based Heuristic to Explain Game-Theoretic Models · Virginia Tech

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.