Chapter Four · failure evidence
What Graph Machine Learning got wrong, from 69 dissertations
Across these doctoral theses, graph machine learning approaches often struggled against simpler non-graph baselines, suffered from architectural oversmoothing, and failed to generalize across divergent network topologies. Practitioners also encountered substantial difficulties with noisy graph construction, computational bottlenecks on large networks, and representation collapse in knowledge graphs and heterophilic settings. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Graph neural networks underperformed simpler tabular and heuristic baselines
Deep graph neural networks and learned graph embeddings repeatedly failed to surpass standard tabular models, random forests, and multilayer perceptrons across tasks such as molecular property prediction and node classification. Researchers frequently rejected or abandoned complex GNN architectures because simpler baselines achieved comparable or superior accuracy with greater interpretability and lower complexity.
Tried and failed
graph convolutional network without pretraining applied to molecular property prediction. Outcome: worse than baseline. Reason: failed to exceed sequence-based models and barely matched simple baselines
Photoionization Detection of Volatile Organic Compounds · Harvard
Tried and failed
graph neural network on atomic coordinates applied to predicting global scalar material properties. Outcome: worse than baseline. Reason: provided only marginal accuracy improvement over random forest baselines despite much higher complexity
Using Deep Learning to Understand and Design Heterogeneous Materials · MIT
Tried and failed
graph neural networks with language model embeddings applied to chemical reaction yield prediction. Outcome: worse than baseline. Reason: failed to outperform random forests trained on molecular fingerprints and physical descriptors
Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT
Tried and failed
increasing graph neural network architecture complexity applied to predicting electronic structure properties of crystals. Outcome: worse than baseline. Reason: advanced architectures provided only marginal accuracy gains while sacrificing physical interpretability
Computational and Data-Driven Design of Perturbed Metal Sites for Catalytic Transformations · Virginia Tech
Tried and failed
augmenting graph neural networks with handcrafted molecular descriptors applied to molecular property prediction from semi-empirical calculations. Outcome: worse than baseline. Reason: hyperparameter tuning and manual descriptor concatenation yielded minimal performance gains or degraded predictive accuracy
High-throughput virtual screening of molecules for photon conversion · Imperial
Tried and failed
GNNs on static population graphs applied to tabular and imaging phenotype regression. Outcome: worse than baseline. Reason: Static heuristic graph construction introduces suboptimal connectivity compared to standard MLP baselines
Deep learning for interpretable brain age estimation · Imperial
Tried and failed
learned molecular graph embeddings applied to small molecule representation for drug-target interaction. Outcome: worse than baseline. Reason: learned embeddings failed to outperform simple fingerprint baselines
Lost to a baseline
On CoauthorCS, simple Dense feed-forward baseline without graph structure (94.88±0.21%) outperformed GCN (93.85±0.23%), GAT (93.80±0.38%), and L-CAT (93.65±0.23%)
Mitigating Practical Limitations of GNNs through Explicit Inductive Biases · Publikationssystem UB Tuebingen
Lost to a baseline
Baseline graph-agnostic MLP (MAE 3.73 years) beat static population graph GCNs (MAE 3.89 to 5.19 years) and GATs (MAE 4.07 to 5.38 years).
Deep learning for interpretable brain age estimation · Imperial
Considered and rejected
Considered and rejected: Rejected Graph Neural Networks for node-level epidemic status inference in Chapter 1 in favor of standard MLPs for simplicity and generalizability to non-network models.
Approximate Bayesian Inference for Network Processes · Harvard
Considered and rejected
Considered and rejected: Rejected using complex non-linear deep learning or graph neural networks (GNNs) requiring structural input, choosing single-parameter-per-element compositional linear/ReLU models to retain human-interpretable chemical heuristics.
Learning Simple Chemical Heuristics to Model and Discover Materials · MIT
Lost to a baseline
GraphSAGE GNN (PR-AUC 0.3709) was substantially beaten across all matched feature tiers by tabular XGBoost (PR-AUC 0.5185)
THE AI EDGE IN CONFLICT PREDICTION THROUGH ANALYSIS OF THE INFORMATION ENVIRONMENT · Calhoun
Considered and rejected
Considered and rejected: Rejected standalone deep graph neural networks (GNNs/CNNs) as primary operational classifiers due to black-box opacity, lack of auditability for financial regulators, and high computational inference costs.
AI-Driven Financial Market Surveillance: Detecting and Preventing Misconduct in Decentralized Systems · IRIS - POLITO - prod
Tried and failed
Graph neural networks with Mol2vec embeddings applied to drug synergy prediction. Outcome: did not generalise. Reason: Poor test performance and inability to generalize compared to standard fingerprint-based feed-forward networks
Combination Antibiotic-Focused Machine Learning Models for Integration into Experimental Workflows · Harvard
Considered and rejected
Considered and rejected: Decided against deep graph neural network similarity learning due to lack of interpretability.
Graph matching with applications to network analysis · OpenBU
Lost to a baseline
Feature-based graph embedding methods MUSAE (AUC 0.5261 w/o POI, 0.4940 w/ POI), TENE (AUC 0.4928 w/o POI, 0.5316 w/ POI), and ASNE (AUC 0.4904 w/o POI, 0.5203 w/ POI) were outperformed by the raw feature baseline (AUC 0.5603 w/o POI, 0.7984 w/ POI).
Individual Human Mobility Models for sustainable cities applications · IRIS - SNS - prod
Tried and failed
graph neural networks as surrogate models applied to optimal power flow proxies. Outcome: worse than baseline. Reason: numerical instability and lower prediction accuracy compared to standard deep neural networks
Synergizing Machine Learning and Optimization: Scalable Real-Time Risk Assessment in Power Systems · Georgia Tech
Flawed graph construction and indiscriminate neighborhood aggregation introduced harmful noise
Constructing graphs from heuristic thresholds, synthetic features, or unpruned paths injected noise that degraded representation learning across benchmarks. In addition, uniform neighborhood aggregation and multi-hop expansions frequently diluted critical signals by incorporating irrelevant neighbors and disconnected triplets.
Tried and failed
combining language models with graph structure embeddings applied to link prediction on sparse graphs. Outcome: worse than baseline. Reason: graph structure information introduced noise on sparse graphs, degrading prediction performance
Temporal link prediction in the wild · Imperial
Tried and failed
Graph convolutional networks with synthetic node features applied to non-attributed graph community detection. Outcome: worse than baseline. Reason: synthetic features like landing probabilities or random vectors fail to provide meaningful signals for GCN aggregation
On Statistical Learning for Structural Data: Data Fusion and Semi-supervised Learning · Harvard
Tried and failed
graph neural network with mean neighborhood aggregation applied to k-nearest neighbor similarity graph regression. Outcome: worse than baseline. Reason: uniform aggregation failed to adaptively weight varying neighbor relevance
Graph Neural Network Architectures for interpretable I/O bottleneck analysis in high-performance computing systems · Iowa State
Tried and failed
graph convolutional networks with neighborhood feature aggregation applied to influencer detection in social networks. Outcome: worse than baseline. Reason: aggregating neighborhood node features introduced noise and degraded predictive performance
Analyzing Networks with Hypergraphs: Detection, Classification, and Prediction · Virginia Tech
Tried and failed
random landmark node selection for embedding applied to graph representation learning. Outcome: worse than baseline. Reason: random selection captured less structural importance than degree-based landmark selection across datasets
Considered and rejected
Considered and rejected: Using full dependency graphs without pruning to entity-mention paths hurt representation learning in GraphTransfer.
Low-Resource Event Extraction · Scholars' Bank
Tried and failed
expanding knowledge graph embeddings to multi-hop neighborhoods applied to multimodal cross-attention feature enrichment. Outcome: worse than baseline. Reason: expanded neighborhood graphs introduced irrelevant concept noise that degraded cross-modal alignment
Commonsense for Zero-Shot Natural Language Video Localization · Virginia Tech
Tried and failed
random walk graph embedding on full hierarchy applied to unfiltered large-scale medical ontology. Outcome: worse than baseline. Reason: generating embeddings without domain-presence filtering introduced noise from irrelevant or overly granular concepts
Artificial Intelligence-Driven Clinical Decision Support for Antibiotic Optimisation · Imperial
Tried and failed
inserting multiple knowledge graph triplets into text applied to entity linking language models. Outcome: worse than baseline. Reason: inserting more candidate triplets introduced excessive noise compared to using a single triplet
Leveraging Structured Knowledge to Enable Efficient AI Assistants · Georgia Tech
Tried and failed
random graph connectivity in graph neural networks applied to named entity recognition. Outcome: worse than baseline. Reason: destroying structural layout priors removes informative relational context
Tried and failed
high correlation threshold for graph edge construction applied to node embedding correlation graph reconstruction. Outcome: worse than baseline. Reason: overly restrictive threshold caused excessive false negatives by omitting true dependencies
Learning and Reconstructing Conflicts in O-RAN · Virginia Tech
Tried and failed
k-nearest neighbor graph construction from features applied to adversarial graph purification and learning. Outcome: worse than baseline. Reason: Feature-only kNN connections discard spectral and topological structural information needed for accurate graph recovery
Accurate and Efficient Representation Learning on Large-Scale Graphs · Cornell
Graph models failed to generalize across varying topologies and out of distribution structures
Models trained on specific graph structures struggled to transfer to networks of different sizes, connectivity layouts, or longer evaluation horizons. Transductive attention mechanisms and structural assumptions broke down when tested on multi-component compositions or unseen expansion graphs.
Tried and failed
broadly trained graph neural network property prediction applied to ternary crystal magnetic property screening. Outcome: did not generalise. Reason: generalized crystal graph model exhibited systematic errors on multi-component compositions without domain-specific fine-tuning
The discovery and design of rare-earth-free magnets using machine learning · UT Austin
Tried and failed
discrete generative diffusion on graph adjacency matrices applied to knowledge graph link correction. Outcome: overfit. Reason: severe test set overfitting and failure to generalize beyond training graph structures
Scalable Methods for Knowledge Graph Reasoning and Generation · EPFL
Tried and failed
Magnetic Laplacian on continuous directed graphs applied to fuzzy directed graph representation learning. Outcome: did not generalise. Reason: Fails to separate in- and out-neighbor messages when generalized to continuous edge directions, reducing expressiveness
Developing Differentiable Toolkits for Computational Biology · Harvard
Tried and failed
predicting structural network paradoxes using linear correlation applied to node attributes in complex graphs. Outcome: did not generalise. Reason: linear correlation cannot overcome the influence of specific adverse graph topologies
Well-defined graph-theoretic paradoxes · Cornell
Tried and failed
extending friendship paradox inequality to directed networks applied to directed network topology and node degrees. Outcome: did not generalise. Reason: degree gap inequality fails to hold even on simple three-node directed graph structures
Well-defined graph-theoretic paradoxes · Cornell
Considered and rejected
Considered and rejected: Rejected container-level node modeling in GNN graphs because generalization suffered compared to host-level node modeling with isolated application signatures.
Self-adaptive serverless edge computing for edge intelligence · DSpace-CRIS at TU Wien
Considered and rejected
Considered and rejected: Rejected Graph Attention Networks (GAT) because they operate in a transductive setting requiring the entire graph structure during training, hindering generalization to unseen expansion stations.
Data-Driven Bike-Share Ridership Prediction and Network Optimization · YorkSpace
Considered and rejected
Considered and rejected: Rejected node-level proximity preserving embedding methods (e.g., DeepWalk, Node2Vec) because aligning embeddings across different graphs is problematic.
Individual Human Mobility Models for sustainable cities applications · IRIS - SNS - prod
Tried and failed
learned optimizers parameterized by graph recurrent networks applied to distributed optimization. Outcome: did not generalise. Reason: performance degraded or saturated when evaluated on time horizons longer than seen during training
Tried and failed
zero-shot policy transfer across network topologies applied to optimization algorithm parameter tuning. Outcome: did not generalise. Reason: policies trained on one graph structure fail on differently sized or structured graphs without retraining
Designing policy optimization algorithms for multi-agent reinforcement learning · Georgia Tech
Stacking deeper graph layers caused oversmoothing and feature collapse
Increasing the depth of graph convolutions and message passing layers led to severe oversmoothing where node representations collapsed into uniform vectors. This architectural deepening degraded discriminability across tasks including traffic forecasting, tracking, routing, and physical simulation.
Tried and failed
deep message passing in graph neural networks applied to molecular property prediction. Outcome: did not generalise. Reason: Excessive message passing steps caused feature collapse and over-smoothing of node embeddings.
Advances in Artificial Intelligence for Accelerated Discovery of Energy Storage Polymers · Georgia Tech
Tried and failed
graph attention networks for relational modeling applied to inter-agent interaction modeling. Outcome: worse than baseline. Reason: Suffered from oversmoothing and failed to outperform a simple multilayer perceptron.
Unified and Multimodal Learning for Gaze Prediction in Naturalistic Settings · EPFL
Tried and failed
deepening spatio-temporal graph convolutional networks applied to motion and action prediction. Outcome: worse than baseline. Reason: increasing graph convolution depth caused over-smoothing of node representations
On the motion and action prediction using deep graph models · UT Austin
Tried and failed
stacking deeper continuous graph ODE blocks applied to spatio-temporal traffic forecasting. Outcome: overfit. Reason: increasing neural ODE-GNN depth from 2 to 6 layers caused overfitting relative to a single block
Graph-based Multi-ODE Neural Networks for Spatio-Temporal Traffic Forecasting · Virginia Tech
Tried and failed
increasing graph convolutional neural network layer depth applied to multi-object tracking feature association. Outcome: worse than baseline. Reason: deeper graph convolutions cause feature over-smoothing and degrade node discriminability
A Graph Convolutional Neural Network Based Approach for Object Tracking Using Augmented Detections With Optical Flow · Virginia Tech
Tried and failed
increasing graph convolutional network depth applied to small graph path prediction. Reason: over-smoothing and diminishing returns on small 5-node graphs beyond four layers
Space Layout Optimization For Natural Ventilation Using Machine Learning Techniques · Harvard
Tried and failed
standard graph convolutional networks applied to fully connected graphs in routing problems. Outcome: no signal. Reason: Low-pass filtering caused oversmoothing on complete graphs, producing uniform edge predictions that failed to guide search.
DEEP UNSUPERVISED MODELS LEVERAGING LEARNING AND REASONING · Cornell
Tried and failed
deep single-stage graph neural network applied to flow simulation on network graphs. Outcome: worse than baseline. Reason: increasing network depth caused severe oversmoothing of node representations
Knowledge graph representations suffered from entity confusion and uninformative aggregation
Learning representations without negative sampling or structured relation constraints left models unable to distinguish similar yet distinct entities. Furthermore, min pooling and uniform sampling failed to retain discriminative features and produced disjoint subgraphs lacking semantic coherence.
Tried and failed
variational approximation with generic proposal distributions applied to graph neural network explanation. Outcome: did not converge. Reason: generic variational distributions fail to effectively approximate the complex posterior distribution of subgraph rationales
Learning-based search algorithm design · Georgia Tech
Tried and failed
non-negative representation learning without contrastive structured loss applied to knowledge graph completion. Outcome: worse than baseline. Reason: cannot sufficiently distinguish similar yet distinct entities without negative samples or structured relation constraints
Tried and failed
graph neural network embeddings for clustering applied to pairwise difference latent type identification. Outcome: worse than baseline. Reason: Node embeddings produced false positives and false negatives, underperforming baseline hierarchical clustering.
Tried and failed
knowledge graph augmented retrieval generation applied to document question answering. Outcome: worse than baseline. Reason: vector search retrieved incorrect documents, compounding retrieval errors downstream
Loud Yet Invisible: A Humanist-Designed Pipeline for Unlocking the Early Modern Archive · HARVEST
Tried and failed
min pooling across edge-embedded graph dimensions applied to text graph neural networks. Outcome: worse than baseline. Reason: failed to retain discriminative features compared to max or average pooling
Improving Text Classification Using Graph-based Methods · Virginia Tech
Tried and failed
uniform edge sampling from knowledge graphs applied to generating coherent multi-relational subgraphs. Reason: sampled disconnected or disjoint triplets that lacked semantic coherence for natural text expression
Collaborative AI Agents in the Era of Large Language Models · EPFL
Tried and failed
LLM reranking with web search retrieval applied to knowledge graph completion. Outcome: worse than baseline. Reason: Noisy external search results and mismatch with closed-world benchmark evaluation criteria.
From graphs to truth: towards efficient knowledge graph fusion for factual verification · Imperial
Lost to a baseline
DistMult and ComplEx embedding models were outperformed by TransE across all evaluated triple types in the threat knowledge graph.
Facilitating decision-making in large distributed systems with selfish and adversarial actors · OpenBU
High computational overhead and memory bottlenecks limited scalability on large graphs
Scaling graph architectures to full-scale networks encountered out-of-memory errors and communication bottlenecks during distributed execution. High-dimensional tensor variants and complex attention mechanisms were rejected or beaten due to excessive compute costs without consistent performance gains.
Tried and failed
attentive fingerprint graph neural network applied to molecular property and activity prediction. Outcome: too slow. Reason: increased computational time without consistent performance gains over graph attention networks
Exploring Graph Neural Networks for Molecular Activity Prediction · Harvard
Considered and rejected
Considered and rejected: Rejected RDF-based triple stores and SPARQL querying for large-scale graph learning due to computational inefficiencies and poor storage performance relative to labeled property graphs.
MM-ADM: A model-based approach to multidisciplinary design to support automated decision-making · Georgia Tech
Considered and rejected
Considered and rejected: Rejected high-dimensional tensor GNN variants for MLN inference due to computational intractability on large knowledge graphs, choosing GNN with tunable embeddings.
Knowledge Reasoning with Graph Neural Networks · Georgia Tech
Tried and failed
monolithic centralized neural network for large graphs applied to large-scale grid power flow optimization. Outcome: infeasible cost. Reason: exceeded GPU memory and suffered large prediction errors at full system scale
Advances in Large-Scale Power System Operations: Reconstruction, Reliability, Learning · Georgia Tech
Lost to a baseline
Multi-GPU GraphSage scaling on ogbn-papers100M with small TT ranks (rank 8) achieves less than 2x speedup on 8 GPUs due to non-embedding compute dominance
Fast and compact neural network via Tensor-Train reparameterization · Georgia Tech
Lost to a baseline
For small-world network graphs (such as Barabasi-Albert, cond-mat-2005, loc-Brightkite), deterministic probing scaled poorly and was outperformed by adaptive Hutch++ due to the number of required colors scaling up with graph size.
Approximation of Matrix Functions Arising in Physics and Network Science: Theoretical and Computational Aspects · IRIS - SNS - prod
Lost to a baseline
DGL-GPU training was slower than PyG-GPU on the smallest graph dataset PPI for full-batch GraphSAGE.
Efficient Large-Scale Graph Neural Network Training · TXST Digital Repository
Feature smoothing and message passing broke down on heterophilic graphs
Standard feature and label propagation algorithms failed when applied to heterophilic graphs or networks with low homophily. Regularization techniques and neighborhood smoothing assumptions collapsed because connected nodes did not share similar labels or attributes.
Tried and failed
weight decay and dropout regularization applied to GNNs on heterophilic directed graphs. Outcome: did not generalise. Reason: regularization techniques failed to improve performance on heterophilic graph data
Deep learning on real-world graphs · Imperial
Tried and failed
label and residual propagation on graphs applied to heterophilic and negatively correlated networks. Outcome: worse than baseline. Reason: standard smoothing assumes positive homophily, failing when connected nodes have dissimilar labels
Modeling and Inferring Attributed Graphs · Cornell
Tried and failed
feature propagation for missing node attributes applied to heterophilous or low-homophily graphs. Outcome: worse than baseline. Reason: smoothing features across connected nodes fails when neighbors do not share similar attributes or classes
Deep learning on real-world graphs · Imperial
Left open by the authors
Problems the authors named and did not get to.
Left open
Evaluate 3D graph neural network architectures incorporating geometric features like distances, angles, and torsions for molecular activity prediction. Blocker: None
Exploring Graph Neural Networks for Molecular Activity Prediction · Harvard
Left open
Implement graph neural networks to encode drug compound chemical structures as molecular graphs instead of one-hot encodings for learning-to-rank models. Blocker: None
Robust learning to rank models and their biomedical applications · OpenBU
Left open
Develop graph neural network embeddings or structure-agnostic representations for chemical structures in Bayesian optimization. Blocker: Lacks specific target problem, chemical dataset, and concrete baseline metrics
Bayesian optimisation in chemical problems · Imperial
Left open
Develop graph neural network architectures to predict edges and their directionality over higher-order representations such as graph-products and tensorial adjacencies. Blocker: None
Structure-aware graph representation learning using graph neural networks · Iowa State
Left open
Implement and evaluate graph neural network architectures and graph normalizing flows for biomolecular conformation generation on protein structures. Blocker: None
Physically Interpretable Biomolecular Conformation Generation with A Deep Probabilistic Framework · Harvard
Left open
Implement multitask learning neural networks with parallel regression and classification heads to predict physicochemical properties from GC chromatographic data. Blocker: None
Machine Learning for Structure-Agnostic Chemical Analysis from Chromatographic Data · Virginia Tech
Left open
Pre-train or augment graph neural network molecular activity predictors using large synthetic chemical datasets to improve performance on low-data empirical benchmarks. Blocker: None
Exploring Graph Neural Networks for Molecular Activity Prediction · Harvard
Left open
Develop graph neural network and transformer surrogate models trained on expanded MOF adsorption datasets to improve prediction and extrapolation. Blocker: None
Efficient and Accurate Incorporation of Flexibility and Defects Into the Modeling of Adsorption in Metal-Organic Frameworks · Georgia Tech
Left open
Implement a Graph Neural Network to replace the embedding prediction network in KD-EMD for spatial and relational knowledge distillation. Blocker: None
Left open
Develop graph neural networks and graph variational autoencoders to optimize subgraph sampling and motif search on connectome graphs. Blocker: Lacks specific target architectures, quantitative efficiency benchmarks, or concrete graph sampling algorithms specified in the thesis
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.