Chapter Four · failure evidence
What Rule-Based & Expert Systems got wrong, from 35 dissertations
The evaluated records demonstrate that rule-based systems and manual heuristics frequently falter when confronting linguistic variation, stochastic operating environments, and rigid threshold boundaries. In multiple domains, deterministic rules also suffer from maintenance bottlenecks, fail under distribution shifts, and lag behind statistical or machine learning baselines. These records come from PhD theses at 19 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Rule-based natural language processing fails to handle linguistic variation and complex syntax
Rigid syntactic patterns and regular expression heuristics fail to capture document layout variations, complex negation, and idiosyncratic semantic shifts. Handcrafted text rules and parsers achieve poor precision or recall compared to learned models and struggle with varied linguistic structures.
Tried and failed
rule-based natural language processing applied to clinical report diagnosis extraction. Reason: failed to correctly handle complex clinical negation, causing poor specificity and low F1 score
Tried and failed
heuristic regular-expression rules applied to document metadata extraction. Outcome: did not generalise. Reason: rule-based heuristics failed to capture diverse document layouts and formatting variations across a larger corpus
Tried and failed
rule-based pattern matching NLP applied to inter-parameter dependency extraction from documentation. Outcome: no signal. Reason: rigid syntactic patterns failed to match diverse real-world natural language documentation structures
Effective Automation of Black-Box Testing for REST APIs with Machine Learning and Language Models · Georgia Tech
Tried and failed
rule-based text data augmentation applied to intent classification. Outcome: worse than baseline. Reason: heuristic word-level perturbations alter sentence semantics and label alignment
Lifelong Machine Learning with Data Efficiency and Knowledge Retention · EPFL
Tried and failed
rule-based compositional semantics and derivational morphology applied to compound participle meaning prediction. Outcome: did not generalise. Reason: rules failed to capture idiosyncratic semantic shifts and semi-productivity patterns
The semantics of past participles · UT Austin
Tried and failed
rule-based dependency parser with candidate scoring applied to natural language to formal specification translation. Reason: candidate completions lacking matching input categories were overly penalized by the negative scoring mechanism
Framework for Automatic Translation of Hardware Specifications Written in English to a Formal Language · Virginia Tech
Lost to a baseline
Rule-based baseline NLP (EasyCIE_GUI) achieved an F1 of only 0.43 and precision of 0.34 (at 0.59 recall), losing substantially to deep learning models (BiLSTM F1 0.64–0.66, precision 0.62–0.67).
Surgical Site Infection (SSI) Identification Across Multiple Facilities and Surgery Types Using Multimodal Data and Deep Learning · ResearchWorks
Considered and rejected
Considered and rejected: Rejected using only plain rule-based parsing in ChemDataExtractor due to brittle failure on minor grammatical variations and low inorganic recall (56%).
Considered and rejected
Considered and rejected: Rejected pure rule-based comparative adjective/adverb parsing due to the complexity of building grammar rules for varied linguistic structures.
Compa: A Comparative Retrieval and Analytical Engine for Consumer Products · TXST Digital Repository
Deterministic rules and static heuristics struggle in stochastic and dynamic environments
Deterministic logical models and static heuristic thresholds break down when deployed in dynamic, volatile, or noisy environments. These systems fail to adapt across long time horizons and ignore baseline human or operational stochasticity.
Tried and failed
rule-based behavior inference applied to human response behavior under disruption. Reason: produced systematic overestimation by ignoring baseline stochasticity in human behavior
Tried and failed
deterministic rule-based logical models applied to human activity estimation. Outcome: worse than baseline. Reason: strict logical rules cannot handle noise and uncertainty compared to probabilistic modeling
Cognitive Human Activity and Plan Recognition for Human-Robot Collaboration · MIT
Considered and rejected
Considered and rejected: Rejected rule-based heuristics and MPC for long-term operational optimization due to lack of adaptability in dynamic environments and poor scaling over long-horizon stochastic tasks.
Lost to a baseline
simpler rule-based or gradient-based dispatch methods excelled over heuristic algorithms under stable conditions, only faltering in highly volatile environments
Optimization of Design Parameters for Last-Mile Delivery Drones · UT Austin
Tried and failed
static heuristic-based threshold switching rules applied to dynamic architectural protocol selection. Outcome: worse than baseline. Reason: simple thresholds fail to capture complex workload dynamics, yielding incorrect decisions in half the cases
AI-DRIVEN ADAPTIVE DISTRIBUTED SYSTEMS IN UNTRUSTED ENVIRONMENTS · Penn
Tried and failed
deterministic rule-based modeling of dynamic systems applied to agent survival simulation. Outcome: did not generalise. Reason: Failed to survive and became erratic outside a narrow payoff range
An adaptive agent-based multicriteria simulation system · Cranfield
Rigid thresholds and hardcoded boundaries cause false positives and poor boundary behavior
Handcrafted numerical thresholds and crisp conditional rules fail to handle boundary cases fairly and generate high rates of false positives. These rigid cutoff approaches lead to oversimplified models or unmaintainable rule sets that do not generalize across diverse styles.
Tried and failed
heuristic rule-based constraint extraction applied to circuit symmetry constraint detection. Outcome: worse than baseline. Reason: produced higher false positive rates and generated unnecessary constraints compared to learned representations
Layout automation for custom integrated circuits · UT Austin
Considered and rejected
Considered and rejected: Default high dependency thresholds in the Flexible Heuristics Miner algorithm were rejected because they led to oversimplified process models in low-structured domains.
Rezeptions- und Interpretationsprozesse von Lehrpersonen bei datengestützten Entscheidungen. Exploration und Förderung · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Crisp rule-based expert systems using rigid mathematical thresholds were rejected because they fail to handle boundary cases fairly and create overly complex, unmaintainable rule bases.
Strategie działania inteligentnych systemów wspierających kształcenie operujące na danych nieprecyzyjnych Strategies of operation for intelligent tutoring systems operating on imprecise data · AMUR - Repozytorium Uniwersytetu im. Adama Mickiewicza w Poz
Considered and rejected
Considered and rejected: Fixing static μ thresholds rule-based was rejected in favor of dynamically low-pass filtering synaptic weight efficacy
Considered and rejected
Considered and rejected: Rejected classical rule-based conditional programming (hardcoded angle thresholds) because it produced false positives while moving/pretending to shoot and could not generalize across diverse shooting styles.
Live Perception and Real Time Motion Prediction with Deep Neural Networks and Machine Learning · Harvard
Manual rule systems suffer from poor maintainability and inability to learn or scale
Expert deduction engines and handcrafted anti-pattern heuristics cannot incorporate new observations or generalize to new problem domains. Adding criteria requires extensive manual reprogramming, while crowdsourced rules often encode shallow heuristics that quickly plateau.
Tried and failed
extracting decision rules from human-voted behavior applied to sequential decision-making in disrupted environments. Outcome: did not generalise. Reason: crowdsourced rules encoded shallow, suboptimal heuristics with plateauing performance gains
Managing The Gig Economy Via Behavioral And Operational Lenses · Penn
Considered and rejected
Considered and rejected: Rejected rule-based classifiers due to poor handling of continuous numerical attributes like GPA and GRE scores.
Towards Better Interpretability of Machine Learning-Based Decision Support Systems · DSpace at SUNY Buffalo
Tried and failed
rule-based anti-pattern heuristics for manual diagnosis applied to identifying algorithmic complexity vulnerabilities. Reason: developers could not accurately diagnose complex edge cases using manual pattern rules
Theory and Patterns for Avoiding Regex Denial of Service · Virginia Tech
Considered and rejected
Considered and rejected: Rejected fuzzy logic/AI rule-based expert systems due to inflexibility requiring full reprogramming for added criteria
Considered and rejected
Considered and rejected: Rejected rule-based deduction engines (e.g. from IoIF) because semantic rules cannot learn from new observations or scale across new problem domains.
MM-ADM: A model-based approach to multidisciplinary design to support automated decision-making · Georgia Tech
Heuristic decision rules and synthetic models fail under distribution shifts
Simple heuristics and synthetic rule-based error models fail to generalize to real error distributions or unrepresentative baselines. Heuristic decision support systems can increase user response latency without improving accuracy, while stopping heuristics miss human-understandable failure patterns.
Considered and rejected
Considered and rejected: Rejected automatically distinguishing between genuinely unknown biochemical gaps and reconstruction errors in dead-end tests due to poor accuracy of previous automated heuristics.
Tried and failed
heuristic training for counterfactual forecasting applied to decision making under unrepresentative baselines. Outcome: did not generalise. Reason: training failed when faced with unrepresentative baseline distributions or prospective conditional framing
Considered and rejected
Considered and rejected: Rejected synthetic rule-based error injection models because their error distributions failed to generalize to real generation errors
Fine-grained evaluation for text summarization · UT Austin
Tried and failed
rule-based heuristic decision support system applied to human decision-making under degraded information. Outcome: worse than baseline. Reason: increased user response times without improving accuracy, especially for lower-performing users
Understanding and Supporting Decision Making in Denied and Degraded Environments · Georgia Tech
Considered and rejected
Considered and rejected: Rejected simple confidence-based early-stopping heuristics and generic upweighting methods for error mitigation because they fail to capture consistent, human-understandable failure modes.
Rule-based classifiers and filters underperform probabilistic and machine learning baselines
Rule-based classifiers and heuristic filters achieve lower accuracy than probabilistic models and decision trees on classification tasks. Simple heuristic mechanisms produce repetitive predictions and fail to capture complex causal topological interactions.
Lost to a baseline
List Recall at K=20 was only 0.046 for rule-based method vs 0.008 for the baseline, suffering from repetitive predictions across adjacent time windows.
Artificial Intelligence Methods and Evaluation Strategies for Detecting Future Customer Needs from User Generated Content · Research Repository UCD
Lost to a baseline
Heuristic rules-based algorithm achieved lower overall gait event identification accuracy across locomotion modes (94.87%) compared to the unsupervised BP-AR-HMM (99.6%).
Machine Learning and Wearable Sensors for the Estimation of Biomechanical Variables Outside the Laboratory · Scholars' Bank
Lost to a baseline
Rule-based PART classifier (0.5005 accuracy on investment cost) lost to C4.5 Trees J48 (0.571 average accuracy) and Naive Bayes (0.5091 accuracy on investment cost)
Renewable Energy Communities: A Preference Learning Approach to Evaluate Differential Participation in Cities · IRIS - POLITO - prod
Tried and failed
heuristic filter-based feature selection applied to transient stability classification. Outcome: worse than baseline. Reason: standard filters failed to capture complex causal topological interactions compared to Markov blanket selection
Topological changes in data-driven dynamic security assessment for power system control · Imperial
Left open by the authors
Problems the authors named and did not get to.
Left open
Extend the tree-based rule extraction heuristic to estimate continuous variables and generalize across different problem sizes. Blocker: Lack of specific target continuous variables, validation metrics, or concrete generalization methodology
A Machine Learning-Based Heuristic to Explain Game-Theoretic Models · Virginia Tech
Left open
Test whether fast-and-frugal counterfactual forecasting heuristics transfer to real-world scenarios lacking accessible ground truth. Blocker: No specific real-world domain, evaluation methodology, or benchmark dataset is defined for scenarios lacking ground truth
Left open
Adapt the rule-extraction and hierarchical decision tree heuristic to explain policy functions in reinforcement learning models. Blocker: The proposal is an exploratory direction without a specified target RL environment, policy model, or evaluation benchmark.
A Machine Learning-Based Heuristic to Explain Game-Theoretic Models · Virginia Tech
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.