Chapter Four · failure evidence

What Graph Machine Learning got wrong, from 69 dissertations

Across these doctoral theses, graph machine learning approaches often struggled against simpler non-graph baselines, suffered from architectural oversmoothing, and failed to generalize across divergent network topologies. Practitioners also encountered substantial difficulties with noisy graph construction, computational bottlenecks on large networks, and representation collapse in knowledge graphs and heterophilic settings. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Graph neural networks underperformed simpler tabular and heuristic baselines

16 theses · 10 institutions

Deep graph neural networks and learned graph embeddings repeatedly failed to surpass standard tabular models, random forests, and multilayer perceptrons across tasks such as molecular property prediction and node classification. Researchers frequently rejected or abandoned complex GNN architectures because simpler baselines achieved comparable or superior accuracy with greater interpretability and lower complexity.

Tried and failed

graph convolutional network without pretraining applied to molecular property prediction. Outcome: worse than baseline. Reason: failed to exceed sequence-based models and barely matched simple baselines

Photoionization Detection of Volatile Organic Compounds · Harvard

Tried and failed

graph neural network on atomic coordinates applied to predicting global scalar material properties. Outcome: worse than baseline. Reason: provided only marginal accuracy improvement over random forest baselines despite much higher complexity

Using Deep Learning to Understand and Design Heterogeneous Materials · MIT

Tried and failed

graph neural networks with language model embeddings applied to chemical reaction yield prediction. Outcome: worse than baseline. Reason: failed to outperform random forests trained on molecular fingerprints and physical descriptors

Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT

Tried and failed

increasing graph neural network architecture complexity applied to predicting electronic structure properties of crystals. Outcome: worse than baseline. Reason: advanced architectures provided only marginal accuracy gains while sacrificing physical interpretability

Computational and Data-Driven Design of Perturbed Metal Sites for Catalytic Transformations · Virginia Tech

Tried and failed

augmenting graph neural networks with handcrafted molecular descriptors applied to molecular property prediction from semi-empirical calculations. Outcome: worse than baseline. Reason: hyperparameter tuning and manual descriptor concatenation yielded minimal performance gains or degraded predictive accuracy

High-throughput virtual screening of molecules for photon conversion · Imperial

Tried and failed

GNNs on static population graphs applied to tabular and imaging phenotype regression. Outcome: worse than baseline. Reason: Static heuristic graph construction introduces suboptimal connectivity compared to standard MLP baselines

Deep learning for interpretable brain age estimation · Imperial

Tried and failed

learned molecular graph embeddings applied to small molecule representation for drug-target interaction. Outcome: worse than baseline. Reason: learned embeddings failed to outperform simple fingerprint baselines

Learning the language of biomolecular interactions · MIT

Lost to a baseline

On CoauthorCS, simple Dense feed-forward baseline without graph structure (94.88±0.21%) outperformed GCN (93.85±0.23%), GAT (93.80±0.38%), and L-CAT (93.65±0.23%)

Mitigating Practical Limitations of GNNs through Explicit Inductive Biases · Publikationssystem UB Tuebingen

Lost to a baseline

Baseline graph-agnostic MLP (MAE 3.73 years) beat static population graph GCNs (MAE 3.89 to 5.19 years) and GATs (MAE 4.07 to 5.38 years).

Deep learning for interpretable brain age estimation · Imperial

Considered and rejected

Considered and rejected: Rejected Graph Neural Networks for node-level epidemic status inference in Chapter 1 in favor of standard MLPs for simplicity and generalizability to non-network models.

Approximate Bayesian Inference for Network Processes · Harvard

Considered and rejected

Considered and rejected: Rejected using complex non-linear deep learning or graph neural networks (GNNs) requiring structural input, choosing single-parameter-per-element compositional linear/ReLU models to retain human-interpretable chemical heuristics.

Learning Simple Chemical Heuristics to Model and Discover Materials · MIT

Lost to a baseline

GraphSAGE GNN (PR-AUC 0.3709) was substantially beaten across all matched feature tiers by tabular XGBoost (PR-AUC 0.5185)

THE AI EDGE IN CONFLICT PREDICTION THROUGH ANALYSIS OF THE INFORMATION ENVIRONMENT · Calhoun

Considered and rejected

Considered and rejected: Rejected standalone deep graph neural networks (GNNs/CNNs) as primary operational classifiers due to black-box opacity, lack of auditability for financial regulators, and high computational inference costs.

AI-Driven Financial Market Surveillance: Detecting and Preventing Misconduct in Decentralized Systems · IRIS - POLITO - prod

Tried and failed

Graph neural networks with Mol2vec embeddings applied to drug synergy prediction. Outcome: did not generalise. Reason: Poor test performance and inability to generalize compared to standard fingerprint-based feed-forward networks

Combination Antibiotic-Focused Machine Learning Models for Integration into Experimental Workflows · Harvard

Considered and rejected

Considered and rejected: Decided against deep graph neural network similarity learning due to lack of interpretability.

Graph matching with applications to network analysis · OpenBU

Lost to a baseline

Feature-based graph embedding methods MUSAE (AUC 0.5261 w/o POI, 0.4940 w/ POI), TENE (AUC 0.4928 w/o POI, 0.5316 w/ POI), and ASNE (AUC 0.4904 w/o POI, 0.5203 w/ POI) were outperformed by the raw feature baseline (AUC 0.5603 w/o POI, 0.7984 w/ POI).

Individual Human Mobility Models for sustainable cities applications · IRIS - SNS - prod

Tried and failed

graph neural networks as surrogate models applied to optimal power flow proxies. Outcome: worse than baseline. Reason: numerical instability and lower prediction accuracy compared to standard deep neural networks

Synergizing Machine Learning and Optimization: Scalable Real-Time Risk Assessment in Power Systems · Georgia Tech

Flawed graph construction and indiscriminate neighborhood aggregation introduced harmful noise

12 theses · 9 institutions

Constructing graphs from heuristic thresholds, synthetic features, or unpruned paths injected noise that degraded representation learning across benchmarks. In addition, uniform neighborhood aggregation and multi-hop expansions frequently diluted critical signals by incorporating irrelevant neighbors and disconnected triplets.

Tried and failed

combining language models with graph structure embeddings applied to link prediction on sparse graphs. Outcome: worse than baseline. Reason: graph structure information introduced noise on sparse graphs, degrading prediction performance

Temporal link prediction in the wild · Imperial

Tried and failed

Graph convolutional networks with synthetic node features applied to non-attributed graph community detection. Outcome: worse than baseline. Reason: synthetic features like landing probabilities or random vectors fail to provide meaningful signals for GCN aggregation

On Statistical Learning for Structural Data: Data Fusion and Semi-supervised Learning · Harvard

Tried and failed

graph neural network with mean neighborhood aggregation applied to k-nearest neighbor similarity graph regression. Outcome: worse than baseline. Reason: uniform aggregation failed to adaptively weight varying neighbor relevance

Graph Neural Network Architectures for interpretable I/O bottleneck analysis in high-performance computing systems · Iowa State

Tried and failed

graph convolutional networks with neighborhood feature aggregation applied to influencer detection in social networks. Outcome: worse than baseline. Reason: aggregating neighborhood node features introduced noise and degraded predictive performance

Analyzing Networks with Hypergraphs: Detection, Classification, and Prediction · Virginia Tech

Tried and failed

random landmark node selection for embedding applied to graph representation learning. Outcome: worse than baseline. Reason: random selection captured less structural importance than degree-based landmark selection across datasets

Graph Embedding for Retrieval · EPFL

Considered and rejected

Considered and rejected: Using full dependency graphs without pruning to entity-mention paths hurt representation learning in GraphTransfer.

Low-Resource Event Extraction · Scholars' Bank

Tried and failed

expanding knowledge graph embeddings to multi-hop neighborhoods applied to multimodal cross-attention feature enrichment. Outcome: worse than baseline. Reason: expanded neighborhood graphs introduced irrelevant concept noise that degraded cross-modal alignment

Commonsense for Zero-Shot Natural Language Video Localization · Virginia Tech

Tried and failed

random walk graph embedding on full hierarchy applied to unfiltered large-scale medical ontology. Outcome: worse than baseline. Reason: generating embeddings without domain-presence filtering introduced noise from irrelevant or overly granular concepts

Artificial Intelligence-Driven Clinical Decision Support for Antibiotic Optimisation · Imperial

Tried and failed

inserting multiple knowledge graph triplets into text applied to entity linking language models. Outcome: worse than baseline. Reason: inserting more candidate triplets introduced excessive noise compared to using a single triplet

Leveraging Structured Knowledge to Enable Efficient AI Assistants · Georgia Tech

Tried and failed

random graph connectivity in graph neural networks applied to named entity recognition. Outcome: worse than baseline. Reason: destroying structural layout priors removes informative relational context

From Structured Document To Structured Knowledge · MIT

Tried and failed

high correlation threshold for graph edge construction applied to node embedding correlation graph reconstruction. Outcome: worse than baseline. Reason: overly restrictive threshold caused excessive false negatives by omitting true dependencies

Learning and Reconstructing Conflicts in O-RAN · Virginia Tech

Tried and failed

k-nearest neighbor graph construction from features applied to adversarial graph purification and learning. Outcome: worse than baseline. Reason: Feature-only kNN connections discard spectral and topological structural information needed for accurate graph recovery

Accurate and Efficient Representation Learning on Large-Scale Graphs · Cornell

Graph models failed to generalize across varying topologies and out of distribution structures

9 theses · 9 institutions

Models trained on specific graph structures struggled to transfer to networks of different sizes, connectivity layouts, or longer evaluation horizons. Transductive attention mechanisms and structural assumptions broke down when tested on multi-component compositions or unseen expansion graphs.

Tried and failed

broadly trained graph neural network property prediction applied to ternary crystal magnetic property screening. Outcome: did not generalise. Reason: generalized crystal graph model exhibited systematic errors on multi-component compositions without domain-specific fine-tuning

The discovery and design of rare-earth-free magnets using machine learning · UT Austin

Tried and failed

discrete generative diffusion on graph adjacency matrices applied to knowledge graph link correction. Outcome: overfit. Reason: severe test set overfitting and failure to generalize beyond training graph structures

Scalable Methods for Knowledge Graph Reasoning and Generation · EPFL

Tried and failed

Magnetic Laplacian on continuous directed graphs applied to fuzzy directed graph representation learning. Outcome: did not generalise. Reason: Fails to separate in- and out-neighbor messages when generalized to continuous edge directions, reducing expressiveness

Developing Differentiable Toolkits for Computational Biology · Harvard

Tried and failed

predicting structural network paradoxes using linear correlation applied to node attributes in complex graphs. Outcome: did not generalise. Reason: linear correlation cannot overcome the influence of specific adverse graph topologies

Well-defined graph-theoretic paradoxes · Cornell

Tried and failed

extending friendship paradox inequality to directed networks applied to directed network topology and node degrees. Outcome: did not generalise. Reason: degree gap inequality fails to hold even on simple three-node directed graph structures

Well-defined graph-theoretic paradoxes · Cornell

Considered and rejected

Considered and rejected: Rejected container-level node modeling in GNN graphs because generalization suffered compared to host-level node modeling with isolated application signatures.

Self-adaptive serverless edge computing for edge intelligence · DSpace-CRIS at TU Wien

Considered and rejected

Considered and rejected: Rejected Graph Attention Networks (GAT) because they operate in a transductive setting requiring the entire graph structure during training, hindering generalization to unseen expansion stations.

Data-Driven Bike-Share Ridership Prediction and Network Optimization · YorkSpace

Considered and rejected

Considered and rejected: Rejected node-level proximity preserving embedding methods (e.g., DeepWalk, Node2Vec) because aligning embeddings across different graphs is problematic.

Individual Human Mobility Models for sustainable cities applications · IRIS - SNS - prod

Tried and failed

learned optimizers parameterized by graph recurrent networks applied to distributed optimization. Outcome: did not generalise. Reason: performance degraded or saturated when evaluated on time horizons longer than seen during training

Control and Optimization over Large-Scale Networks · Penn

Tried and failed

zero-shot policy transfer across network topologies applied to optimization algorithm parameter tuning. Outcome: did not generalise. Reason: policies trained on one graph structure fail on differently sized or structured graphs without retraining

Designing policy optimization algorithms for multi-agent reinforcement learning · Georgia Tech

Stacking deeper graph layers caused oversmoothing and feature collapse

8 theses · 6 institutions

Increasing the depth of graph convolutions and message passing layers led to severe oversmoothing where node representations collapsed into uniform vectors. This architectural deepening degraded discriminability across tasks including traffic forecasting, tracking, routing, and physical simulation.

Tried and failed

deep message passing in graph neural networks applied to molecular property prediction. Outcome: did not generalise. Reason: Excessive message passing steps caused feature collapse and over-smoothing of node embeddings.

Advances in Artificial Intelligence for Accelerated Discovery of Energy Storage Polymers · Georgia Tech

Tried and failed

graph attention networks for relational modeling applied to inter-agent interaction modeling. Outcome: worse than baseline. Reason: Suffered from oversmoothing and failed to outperform a simple multilayer perceptron.

Unified and Multimodal Learning for Gaze Prediction in Naturalistic Settings · EPFL

Tried and failed

deepening spatio-temporal graph convolutional networks applied to motion and action prediction. Outcome: worse than baseline. Reason: increasing graph convolution depth caused over-smoothing of node representations

On the motion and action prediction using deep graph models · UT Austin

Tried and failed

stacking deeper continuous graph ODE blocks applied to spatio-temporal traffic forecasting. Outcome: overfit. Reason: increasing neural ODE-GNN depth from 2 to 6 layers caused overfitting relative to a single block

Graph-based Multi-ODE Neural Networks for Spatio-Temporal Traffic Forecasting · Virginia Tech

Tried and failed

increasing graph convolutional neural network layer depth applied to multi-object tracking feature association. Outcome: worse than baseline. Reason: deeper graph convolutions cause feature over-smoothing and degrade node discriminability

A Graph Convolutional Neural Network Based Approach for Object Tracking Using Augmented Detections With Optical Flow · Virginia Tech

Tried and failed

increasing graph convolutional network depth applied to small graph path prediction. Reason: over-smoothing and diminishing returns on small 5-node graphs beyond four layers

Space Layout Optimization For Natural Ventilation Using Machine Learning Techniques · Harvard

Tried and failed

standard graph convolutional networks applied to fully connected graphs in routing problems. Outcome: no signal. Reason: Low-pass filtering caused oversmoothing on complete graphs, producing uniform edge predictions that failed to guide search.

DEEP UNSUPERVISED MODELS LEVERAGING LEARNING AND REASONING · Cornell

Tried and failed

deep single-stage graph neural network applied to flow simulation on network graphs. Outcome: worse than baseline. Reason: increasing network depth caused severe oversmoothing of node representations

Towards quasi real-time simulations of district heating networks for an optimal sustainable design and control · EPFL

Knowledge graph representations suffered from entity confusion and uninformative aggregation

8 theses · 8 institutions

Learning representations without negative sampling or structured relation constraints left models unable to distinguish similar yet distinct entities. Furthermore, min pooling and uniform sampling failed to retain discriminative features and produced disjoint subgraphs lacking semantic coherence.

Tried and failed

variational approximation with generic proposal distributions applied to graph neural network explanation. Outcome: did not converge. Reason: generic variational distributions fail to effectively approximate the complex posterior distribution of subgraph rationales

Learning-based search algorithm design · Georgia Tech

Tried and failed

non-negative representation learning without contrastive structured loss applied to knowledge graph completion. Outcome: worse than baseline. Reason: cannot sufficiently distinguish similar yet distinct entities without negative samples or structured relation constraints

Integrating structural and semantic understanding for robust knowledge graph construction: From knowledge graph completion to zero-shot entity linking · Iowa State

Tried and failed

graph neural network embeddings for clustering applied to pairwise difference latent type identification. Outcome: worse than baseline. Reason: Node embeddings produced false positives and false negatives, underperforming baseline hierarchical clustering.

Essays On Machine Learning And Labor Economics · Penn

Tried and failed

knowledge graph augmented retrieval generation applied to document question answering. Outcome: worse than baseline. Reason: vector search retrieved incorrect documents, compounding retrieval errors downstream

Loud Yet Invisible: A Humanist-Designed Pipeline for Unlocking the Early Modern Archive · HARVEST

Tried and failed

min pooling across edge-embedded graph dimensions applied to text graph neural networks. Outcome: worse than baseline. Reason: failed to retain discriminative features compared to max or average pooling

Improving Text Classification Using Graph-based Methods · Virginia Tech

Tried and failed

uniform edge sampling from knowledge graphs applied to generating coherent multi-relational subgraphs. Reason: sampled disconnected or disjoint triplets that lacked semantic coherence for natural text expression

Collaborative AI Agents in the Era of Large Language Models · EPFL

Tried and failed

LLM reranking with web search retrieval applied to knowledge graph completion. Outcome: worse than baseline. Reason: Noisy external search results and mismatch with closed-world benchmark evaluation criteria.

From graphs to truth: towards efficient knowledge graph fusion for factual verification · Imperial

Lost to a baseline

DistMult and ComplEx embedding models were outperformed by TransE across all evaluated triple types in the threat knowledge graph.

Facilitating decision-making in large distributed systems with selfish and adversarial actors · OpenBU

High computational overhead and memory bottlenecks limited scalability on large graphs

7 theses · 4 institutions

Scaling graph architectures to full-scale networks encountered out-of-memory errors and communication bottlenecks during distributed execution. High-dimensional tensor variants and complex attention mechanisms were rejected or beaten due to excessive compute costs without consistent performance gains.

Tried and failed

attentive fingerprint graph neural network applied to molecular property and activity prediction. Outcome: too slow. Reason: increased computational time without consistent performance gains over graph attention networks

Exploring Graph Neural Networks for Molecular Activity Prediction · Harvard

Considered and rejected

Considered and rejected: Rejected RDF-based triple stores and SPARQL querying for large-scale graph learning due to computational inefficiencies and poor storage performance relative to labeled property graphs.

MM-ADM: A model-based approach to multidisciplinary design to support automated decision-making · Georgia Tech

Considered and rejected

Considered and rejected: Rejected high-dimensional tensor GNN variants for MLN inference due to computational intractability on large knowledge graphs, choosing GNN with tunable embeddings.

Knowledge Reasoning with Graph Neural Networks · Georgia Tech

Tried and failed

monolithic centralized neural network for large graphs applied to large-scale grid power flow optimization. Outcome: infeasible cost. Reason: exceeded GPU memory and suffered large prediction errors at full system scale

Advances in Large-Scale Power System Operations: Reconstruction, Reliability, Learning · Georgia Tech

Lost to a baseline

Multi-GPU GraphSage scaling on ogbn-papers100M with small TT ranks (rank 8) achieves less than 2x speedup on 8 GPUs due to non-embedding compute dominance

Fast and compact neural network via Tensor-Train reparameterization · Georgia Tech

Lost to a baseline

For small-world network graphs (such as Barabasi-Albert, cond-mat-2005, loc-Brightkite), deterministic probing scaled poorly and was outperformed by adaptive Hutch++ due to the number of required colors scaling up with graph size.

Approximation of Matrix Functions Arising in Physics and Network Science: Theoretical and Computational Aspects · IRIS - SNS - prod

Lost to a baseline

DGL-GPU training was slower than PyG-GPU on the smallest graph dataset PPI for full-batch GraphSAGE.

Efficient Large-Scale Graph Neural Network Training · TXST Digital Repository

Feature smoothing and message passing broke down on heterophilic graphs

2 theses · 2 institutions

Standard feature and label propagation algorithms failed when applied to heterophilic graphs or networks with low homophily. Regularization techniques and neighborhood smoothing assumptions collapsed because connected nodes did not share similar labels or attributes.

Tried and failed

weight decay and dropout regularization applied to GNNs on heterophilic directed graphs. Outcome: did not generalise. Reason: regularization techniques failed to improve performance on heterophilic graph data

Deep learning on real-world graphs · Imperial

Tried and failed

label and residual propagation on graphs applied to heterophilic and negatively correlated networks. Outcome: worse than baseline. Reason: standard smoothing assumes positive homophily, failing when connected nodes have dissimilar labels

Modeling and Inferring Attributed Graphs · Cornell

Tried and failed

feature propagation for missing node attributes applied to heterophilous or low-homophily graphs. Outcome: worse than baseline. Reason: smoothing features across connected nodes fails when neighbors do not share similar attributes or classes

Deep learning on real-world graphs · Imperial

Left open by the authors

Problems the authors named and did not get to.

Left open

Evaluate 3D graph neural network architectures incorporating geometric features like distances, angles, and torsions for molecular activity prediction. Blocker: None

Exploring Graph Neural Networks for Molecular Activity Prediction · Harvard

Left open

Implement graph neural networks to encode drug compound chemical structures as molecular graphs instead of one-hot encodings for learning-to-rank models. Blocker: None

Robust learning to rank models and their biomedical applications · OpenBU

Left open

Develop graph neural network embeddings or structure-agnostic representations for chemical structures in Bayesian optimization. Blocker: Lacks specific target problem, chemical dataset, and concrete baseline metrics

Bayesian optimisation in chemical problems · Imperial

Left open

Develop graph neural network architectures to predict edges and their directionality over higher-order representations such as graph-products and tensorial adjacencies. Blocker: None

Structure-aware graph representation learning using graph neural networks · Iowa State

Left open

Implement and evaluate graph neural network architectures and graph normalizing flows for biomolecular conformation generation on protein structures. Blocker: None

Physically Interpretable Biomolecular Conformation Generation with A Deep Probabilistic Framework · Harvard

Left open

Implement multitask learning neural networks with parallel regression and classification heads to predict physicochemical properties from GC chromatographic data. Blocker: None

Machine Learning for Structure-Agnostic Chemical Analysis from Chromatographic Data · Virginia Tech

Left open

Pre-train or augment graph neural network molecular activity predictors using large synthetic chemical datasets to improve performance on low-data empirical benchmarks. Blocker: None

Exploring Graph Neural Networks for Molecular Activity Prediction · Harvard

Left open

Develop graph neural network and transformer surrogate models trained on expanded MOF adsorption datasets to improve prediction and extrapolation. Blocker: None

Efficient and Accurate Incorporation of Flexibility and Defects Into the Modeling of Adsorption in Metal-Organic Frameworks · Georgia Tech

Left open

Implement a Graph Neural Network to replace the embedding prediction network in KD-EMD for spatial and relational knowledge distillation. Blocker: None

From Symbolic Reasoning to Object Embeddings: Advanced Approaches of Knowledge Distillation in Compacted Neural Networks · Texas Tech

Left open

Develop graph neural networks and graph variational autoencoders to optimize subgraph sampling and motif search on connectome graphs. Blocker: Lacks specific target architectures, quantitative efficiency benchmarks, or concrete graph sampling algorithms specified in the thesis

Deep Learning Tools for Next-Generation Connectomics · MIT

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.