Chapter Four · failure evidence
What Model Pruning & Compression got wrong, from 60 dissertations
The records document empirical evaluations of model pruning, quantization, knowledge distillation, and low-rank compression across neural network architectures. Across these experiments, techniques frequently degraded model accuracy or failed to yield computational speedups when constrained by rigid structural patterns, capacity mismatches, or hardware execution limits. These records come from PhD theses at 20 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Knowledge distillation fails when teachers and students suffer from capacity gaps or misaligned supervision
Large capacity mismatches between teacher and student models lead to severe overfitting, while low-performing or overconfident teachers fail to provide useful representation guidance. Intermediate layer matching, rigid loss weighting, and distillation objectives in continual learning often penalize necessary updates and underperform standard cross-entropy training baselines.
Tried and failed
full model knowledge distillation applied to deep neural network compression. Reason: compounded gradient interference when distilling earlier layers led to diminishing accuracy returns
Tried and failed
layerwise hidden representation matching in knowledge distillation applied to cross-architecture language model compression. Outcome: worse than baseline. Reason: unfiltered layerwise intermediate state matching overly constrained student learning compared to output probability distillation
On Parameter Efficiency of Neural Language Models · Georgia Tech
Tried and failed
knowledge distillation loss with exemplar replay applied to continual learning with recurring classes. Outcome: worse than baseline. Reason: distillation penalizes updating representations when classes reappear across learning sessions with replay
Lifelong Machine Learning with Data Efficiency and Knowledge Retention · EPFL
Tried and failed
knowledge distillation using sample-dependent eigenbases applied to feature extractor compression. Outcome: did not generalise. Reason: sample-dependent eigenbases produce a sub-optimal representation space that fails to transfer across diverse inputs
Lightweight model for content-style balanced photorealistic style transfer · UT Austin
Tried and failed
knowledge distillation from large teacher models applied to quantized model compression on small datasets. Outcome: overfit. Reason: Teacher-student capacity mismatch caused severe overfitting on small training datasets
Energy-efficient Neuromorphic Computing for Resource-constrained Internet of Things Devices · Virginia Tech
Tried and failed
knowledge distillation from low-performing teacher models applied to compact neural network training. Outcome: no signal. Reason: student networks cannot acquire useful representations when the teacher model has poor performance
Tried and failed
distillation using robustly regularized teacher models applied to adversarially robust knowledge distillation. Outcome: worse than baseline. Reason: trade-off regularized teachers provided suboptimal supervision for student accuracy and robustness
Tried and failed
fixed loss weighting in knowledge distillation applied to multimodal cross-modal knowledge distillation. Outcome: overfit. Reason: static distillation balance weights caused severe overfitting during knowledge transfer
Effective deep leaning methodologies for salient object detection · Imperial
Tried and failed
knowledge distillation loss applied to continual learning for image classification. Outcome: worse than baseline. Reason: degraded performance and hindered the ability to benefit from repeated exposures to classes
Shape-biased representations for object category recognition · Georgia Tech
Tried and failed
knowledge distillation loss for guidance applied to long-term time series forecasting. Outcome: worse than baseline. Reason: distillation loss degraded guidance from the source teacher model over long prediction horizons
Artificial Intelligence for Data-centric Surveillance and Forecasting of Epidemics · Georgia Tech
Tried and failed
knowledge distillation applied to conditionally gated convolutional neural networks. Outcome: no signal. Reason: small difference between ground truth and teacher output made distilled loss ineffective
Algorithm-Accelerator Co-Design for High-performance and Secure Deep Learning · Cornell
Tried and failed
knowledge distillation of fine-tuned language models applied to clinical concept presence prediction. Outcome: worse than baseline. Reason: None
Tried and failed
knowledge distillation with a large teacher model applied to image classification on complex datasets. Reason: large capacity gap between teacher and student limited effective knowledge transfer
Tried and failed
knowledge distillation using higher-accuracy teacher model applied to compact object detection student network. Outcome: worse than baseline. Reason: teacher overconfidence degraded distillation quality compared to a smaller teacher model
Improving the Training of Compact Neural Networks for Visual Recognition · EPFL
Lost to a baseline
On CIFAR-10 distillation with 5000 samples/class, baseline cross-entropy student training reached 91.95%, beating FullGrad + Attention distillation (90.68%).
Gradient-based Methods for Deep Model Interpretability · EPFL
Lost to a baseline
On CIFAR-10 with standard training under full regularization (dcw), standard training (95.0% accuracy) beat Shrink & Perturb with distillation (94.7% accuracy).
On architectures and training techniques for neural networks · Oxford
Considered and rejected
Considered and rejected: Knowledge distillation: rejected in favor of domain-specific parameter pruning due to difficulties in distilling fine regulatory reasoning into small student models.
Evaluating Efficiency Gains and Security of LLM-Driven Test Generation for Computerised System Validation: A Compliance-Focused Analysis of Life Sciences Testing Processes · DSpace at Griffith College
Considered and rejected
Considered and rejected: Rejected traditional logit-based KL divergence distillation because fine-tuned teacher logits were noisy/imprecise and diluted binary classification signals.
Distilled model for contextual soft moderation on social media · OpenBU
Structured and filter pruning enforcing rigid constraints degrades accuracy and loses to simpler baselines
Imposing rigid N:M patterns, geometric filter constraints, or channel pruning inadvertently eliminates critical weights and destroys reasoning capabilities or representation quality. In several architectures and low-density regimes, these structured pruning heuristics underperform unstructured magnitude pruning and simple random pruning baselines.
Tried and failed
fixed N:M structured sparsity pruning applied to large language models. Outcome: worse than baseline. Reason: rigid ratio constraints inadvertently remove critical outlier weights, degrading model perplexity
Algorithm–Hardware Co-Design Of Digital Compute-In-Memory Architecture Supporting Flexible And Temporal N:M Sparsity · Georgia Tech
Tried and failed
structured filter pruning applied to convolutional neural network compression. Outcome: worse than baseline. Reason: degrades model accuracy significantly compared to fine-grained N:M structured sparsity patterns
Tried and failed
post-training pruning to structured N:M sparsity applied to large language models. Outcome: worse than baseline. Reason: structured sparsity constraints severely degraded performance on knowledge-intensive downstream tasks
Chasing efficiency in the era of large language models · UT Austin
Tried and failed
pseudorandom structured pruning with strict geometric constraints applied to convolutional neural networks on ImageNet. Outcome: worse than baseline. Reason: overly rigid geometrical constraints caused severe accuracy degradation at high sparsity levels
Efficient processing of vision kernels and deep neural networks on reconfigurable computing architectures · Iowa State
Tried and failed
structured weight pruning applied to large language models for reasoning. Outcome: worse than baseline. Reason: removing large fractions of structural weights destroys complex reasoning capabilities and output coherence
Enabling Small Language Models as Efficient and Capable Agents · Virginia Tech
Tried and failed
structured N:M magnitude-based weight sparsification applied to large language model compression. Outcome: worse than baseline. Reason: magnitude pruning alone severely degrades model perplexity without post-training fine-tuning or decomposition
Structured Sparsity-Aware Hardware-Software Co-Design for Deep Neural Network Acceleration · Georgia Tech
Tried and failed
activity-based structured state pruning applied to state space sequence models. Outcome: worse than baseline. Reason: aggressive pruning ratios cause non-linear accuracy collapse while latency gains plateau
Performance profiling and activity-based state pruning for efficient Mamba inference · Iowa State
Lost to a baseline
Structured pruning (e.g., pruning full convolutional filters/channels) loses 1 or more percentage points of accuracy compared to unstructured magnitude pruning at 60% weights remaining on ResNet-50 on ImageNet.
The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks · MIT
Considered and rejected
Considered and rejected: Rejected structured N:M sparsity patterns in favor of unstructured pruning because unstructured pruning performed better and proved more robust.
Lost to a baseline
Stat channel pruning on Resnet-20 performed worse than random pruning across most preservation rates.
Fine Granularity is Critical for Intelligent Neural Network Pruning · YorkSpace
Lost to a baseline
At very low channel preservation rates on Resnet-20, intelligent pruning methods (SNIP, SNat, Stat) were outperformed by random pruning at initialization.
Fine Granularity is Critical for Intelligent Neural Network Pruning · YorkSpace
Lost to a baseline
On DBSN node pruning at low preservation rates, intelligent pruning methods dropped below random pruning due to over-pruning the final bottleneck layer.
Fine Granularity is Critical for Intelligent Neural Network Pruning · YorkSpace
Left open
Investigate why structured pruning underperforms random initialization on larger networks trained with analog noise on MNIST. Blocker: None
Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech
Pruning at initialization or early in training underperforms gradual post-training pruning and random baselines
One-shot or early magnitude pruning methods inflict severe representation damage and can trigger catastrophic layer collapse. Across multiple benchmarks, post-training pruning and gradual iterative schedules consistently outperform initialization pruning, which often fails to beat random sparse initializations.
Tried and failed
iterative magnitude pruning without retraining applied to random feature models. Outcome: worse than baseline. Reason: fixed weights without re-optimization degraded representation quality compared to the unpruned minimal L2-norm solution
Adaptive and weighted optimization for efficient and robust learning · UT Austin
Tried and failed
uninformed magnitude pruning with weight regeneration applied to large language models. Outcome: worse than baseline. Reason: causes irreparable knowledge damage on complex tasks that weight regeneration fine-tuning cannot recover
Chasing efficiency in the era of large language models · UT Austin
Lost to a baseline
Magnitude pruning after training outperformed all initialization pruning methods including PHEW across high-sparsity regimes.
Leveraging sparsity in deep neural networks for training efficiency, interpretability and generalization · Georgia Tech
Lost to a baseline
Global random pruning and random reinitialization match or beat IMP at step 0 on standard ResNet-20 (88.6% and 88.8% vs 88.5% test accuracy at 16.8% density).
The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks · MIT
Lost to a baseline
Repeated full pruning and gradual pruning were beaten by simple random sparse initialization on large MNIST networks (2-3 hidden layers, 200 neurons/layer)
Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech
Considered and rejected
Considered and rejected: Rejected one-shot magnitude pruning in favor of iterative magnitude pruning because gradual pruning yields higher accuracy and sparser matching subnetworks.
The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks · MIT
Lost to a baseline
On Tiny-ImageNet, Initial (Weight) Magnitude Pruning suffered catastrophic accuracy drops due to layer collapse at lower network densities compared to PHEW, SynFlow, and SynFlow-L2.
Leveraging sparsity in deep neural networks for training efficiency, interpretability and generalization · Georgia Tech
Lost to a baseline
One-time pruning underperformed random sparsity initialization on small neural networks across multiple wait times
Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech
Lost to a baseline
On CIFAR-10 ResNet-50 pruning, 1-epoch early magnitude pruning performed worse than SNIP-MB and gradual pruning methods.
Theoretical model sparsity and compression fail to translate into practical latency reductions or hardware speedups
Unstructured weight pruning and compressed representations frequently fail to reduce execution time because sparse execution routines and metadata overheads dominate system latency. Additionally, complex pruning procedures or iterative factorizations introduce severe memory and computational overheads that exceed any downstream transmission or evaluation benefits.
Tried and failed
post-training channel pruning without fine-tuning applied to convolutional neural networks. Outcome: worse than baseline. Reason: severe accuracy degradation exceeding ten percentage points with negligible throughput improvement under hardware execution
Efficient AI model acceleration through tensor decomposition and dataflow architecture design · Imperial
Considered and rejected
Considered and rejected: Rejected unstructured pruning for learning compact hidden representations because pruned sparse weights do not reduce the number of features or inference overhead.
Achieving More with Less: Learning Generalizable Neural Networks With Less Labeled Data and Computational Overheads · Virginia Tech
Considered and rejected
Considered and rejected: Unstructured network pruning: rejected in favor of structured pruning to avoid requiring specialized sparse hardware accelerators and transmission of sparse structural metadata.
Deep joint source-channel coding for vision-based inference at the wireless edge · Imperial
Tried and failed
run-length encoding metadata compression applied to sparse tensor accelerator bandwidth optimization. Reason: bandwidth overhead is dominated by data payload rather than metadata encoding size
Systematic Modeling and Design of Sparse Deep Neural Network Accelerators · MIT
Tried and failed
weight pruning applied to bound propagation neural network verification. Outcome: too slow. Reason: Sparsity does not yield computational speedups in symbolic interval propagation routines despite reducing constraint complexity in SMT/MILP
Robust neural networks: verification, training, and repair · Imperial
Lost to a baseline
Embedded-ViT with 50% pruning at 128x128 input resolution on CPU achieved lower throughput (23.36 FPS) than unpruned DAE-Former baseline (31.32 FPS)
Weakly supervised and embedded semantic segmentation for Computer Aided Diagnostics · DSpace-CRIS at TU Wien
Lost to a baseline
The baseline accelerator from [29] is ×1.85 faster than the proposed bit-level pruning design due to its dataflow latency optimization.
Considered and rejected
Considered and rejected: Rejected dynamic pruning during local training because resource-constrained edge clients cannot train dense unpruned DNNs initially.
REFT: Resource-Efficient Federated Training Framework for Heterogeneous and Resource-Constrained Environments · Virginia Tech
Considered and rejected
Considered and rejected: Rejected alternating iterative SVD updates and outlier extraction during KV cache compression due to severe latency overheads unacceptable for generative inference.
On the Efficiency and Steerability of Self-Attention Mechanism of Large Language Models · Georgia Tech
Uniform and layer-agnostic pruning allocations cause severe representation collapse
Applying uniform pruning budgets or uniform low-rank matrix approximations across all layers ignores the varying sensitivity and compressibility of different architectural components. Indiscriminately removing weights from early residual blocks, skip connections, or dense language model layers disproportionately degrades critical representations and hurts downstream perplexity.
Tried and failed
post-training magnitude pruning on learned skip-connection weights applied to inter-layer depth-wise averaging modules. Outcome: worse than baseline. Reason: severely degrades model perplexity even at low pruning sparsity thresholds
Enhanced Architectures and Optimization Methods for Efficient Language Modeling · EPFL
Tried and failed
magnitude pruning applied to state space models. Outcome: worse than baseline. Reason: architecture was highly sensitive to parameter pruning with significant performance drops at moderate sparsity
Tried and failed
uniform layer-wise pruning budget allocation applied to deep neural network compression. Outcome: worse than baseline. Reason: fails to account for non-uniform layer compressibility and importance, yielding suboptimal accuracy-compression trade-offs
Tried and failed
simultaneous pruning of early residual blocks applied to deep neural network compression. Outcome: worse than baseline. Reason: removes critical low-level representations leading to severe performance degradation
Sparsity prior in efficient deep learning based solvers and models · UT Austin
Tried and failed
backward or parallel layer-wise pruning applied to neural network model compression. Outcome: worse than baseline. Reason: deep-to-shallow or simultaneous pruning degraded accuracy compared to forward layer-by-layer pruning
Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech
Tried and failed
uniform low-rank matrix approximation across all layers applied to deep neural network weight matrices. Outcome: worse than baseline. Reason: indiscriminate rank reduction harms critical representations compared to targeted layer-specific pruning
Discovering and Engineering the Computation Underlying Large Intelligent Agents · MIT
Tried and failed
tensor-train decomposition for neural network compression applied to dense language model layers. Outcome: worse than baseline. Reason: degrades broad, weak lower-level feature representation required for commonsense reasoning tasks
Generative language model compression with algebraic approaches · Imperial
Tried and failed
direct computation cost loss penalty applied to dynamic neural network channel pruning. Reason: disproportionately pruned higher-FLOP layers, causing severe layer-wise imbalance compared to threshold loss
Algorithm-Accelerator Co-Design for High-performance and Secure Deep Learning · Cornell
Quantization induces precision errors and disrupts weight ranking when combined with pruning
Aggressive sub-4-bit quantization and post-training integer rounding create substantial numerical errors that severely degrade model accuracy without extensive distillation. In addition, quantizing before magnitude pruning perturbs weight values, which disrupts the relative magnitude ranking required for effective parameter pruning.
Tried and failed
post-training quantization applied to large language models across task difficulties. Outcome: did not generalise. Reason: performance degradation was non-monotonic across varying task difficulty levels compared to pruning
Chasing efficiency in the era of large language models · UT Austin
Tried and failed
quantization before magnitude pruning applied to neural network model compression. Outcome: worse than baseline. Reason: quantization perturbs weight magnitudes, disrupting pruning order and degrading perplexity
Compressing DNNs Using Microscaling Formats with Sensitivity and Sparsity · EPFL
Tried and failed
post-training quantization without knowledge distillation applied to deep neural networks. Outcome: worse than baseline. Reason: aggressive low-bit term quantization severely degrades accuracy without multi-resolution distillation guidance
Systolic Architectures for Efficient Deep Neural Network Implementations with Assured Performance · Harvard
Tried and failed
extreme sub-4-bit iterative weight quantization hierarchy applied to deep neural network weight compression. Outcome: worse than baseline. Reason: drastic rounding errors at 1-bit and 2-bit precisions severely degrade network performance
Advancing efficiency and trustworthiness : from computer vision to multimodal large language models · UT Austin
Considered and rejected
Considered and rejected: Rejected post-training integer quantization (INT8/INT4) in favor of FP16/FP32 encoder pruning to avoid accuracy degradation and retraining/calibration overhead.
On-device Learning and Inference Optimization for Lightweight Neural Networks and Transformers on Microcontrollers · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Quantization-aware training (QAT) was considered for edge model compression but not used in the thesis, using post-training quantization instead.
Robust Machine Learning Against Faults in Micro-Controllers and Stragglers in Distributed Training on the Cloud · Virginia Tech
Considered and rejected
Considered and rejected: Rejected conventional bit packing because the wide dynamic range of ML values provides minimal compression across individual tensors.
Mitigating the Impact of Data Movement in Memory-Intensive Applications · ResearchWorks
Left open by the authors
Problems the authors named and did not get to.
Left open
Analyze why BERT pruning algorithms naturally prune more neurons in deeper layers than earlier layers. Blocker: None
Sparsity prior in efficient deep learning based solvers and models · UT Austin
Left open
Develop model compression techniques using highly sparse binary masks to reduce memory footprint and inference latency without degrading accuracy. Blocker: None
Algorithms for Efficient and Robust Distributed Deep Learning · EPFL
Left open
Develop pruning techniques to recover accuracy for deep neural network models with dual-side structured sparsity (HighLight_DSSO). Blocker: Lacks specific targets, pruning algorithms, or designated network architectures to evaluate
Systematic Modeling and Design of Sparse Deep Neural Network Accelerators · MIT
Left open
Tune training hyperparameters to quantify the workload required to recover accuracy after deep learning model pruning. Blocker: None
Tools for efficient Deep Learning · Imperial
Left open
Apply pruning, feature importance, and complexity analysis methods to CNN models representing binary nanophotonic structures. Blocker: None
Machine Learning Approaches for Knowledge Discovery in Nanophotonic Structures · Georgia Tech
Left open
Implement layer-wise pruning on 4D CNNs to evaluate temporal feature retention and parameter reduction in deep layers. Blocker: None
Development of a machine-learning platform for the autonomous analysis of 3D+T calcium imaging data · UT Austin
Left open
Extend the pruning and feature importance framework to CNN models trained on binary nanophotonic structure electromagnetic simulation data. Blocker: None
Machine Learning Approaches for Knowledge Discovery in Nanophotonic Structures · Georgia Tech
Left open
Apply advanced neural network pruning to reduce connection complexity between design and response spaces in the nanophotonic inverse design framework. Blocker: None
A New Paradigm for Knowledge Discovery and Design in Nanophotonics Based on Artificial Intelligence · Georgia Tech
Left open
Apply pruning, quantization, and edge hardware acceleration to the STFT-CNN power quality disturbance classification model. Blocker: None
Power Quality Disturbance Classification in IEEE 9-Bus System using STFT and Deep Learning · Texas Tech
Left open
Apply pruning and quantization to the autoencoder-compressed transformer summarization pipeline and evaluate size-accuracy trade-offs. Blocker: None
Efficient and Enhanced Text Summarization by Compressing and Data Augmentation for Transformers-Based Models · Scholarship at UWindsor Institutional Repository
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.