Chapter Four · failure evidence
What Reinforcement Learning got wrong, from 74 dissertations
Across numerous applications, reinforcement learning methods frequently struggle with reward specification flaws, training instability, and poor generalization across domain shifts. Researchers often observe that reinforcement learning policies are outperformed by simpler classical heuristics or reject the paradigm entirely due to high sample complexity and physical safety risks. These records come from PhD theses at 18 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Reward function misspecification leads to reward hacking and degenerate policy behavior
Agents trained with misconfigured or narrow rewards frequently exploit proxy signals such as format compliance or proximity incentives without completing the actual underlying task. In other instances, optimizing solely for isolated reward terms degrades essential qualities such as language fluency or fails to differentiate policy quality entirely.
Tried and failed
using temporal logic robustness as reinforcement learning reward applied to robot control with temporal logic specifications. Outcome: worse than baseline. Reason: robustness scores alone provide poor reward landscape and sparse feedback for reinforcement learning
Tried and failed
reinforcement learning with custom reward function applied to energy management optimization. Outcome: did not converge. Reason: downward trend in reward function over training time due to poor learning or unsuitable reward design
Probabilistic Approaches to Enhance Safety and Energy Management of Energy Transition Components and Systems · Texas Tech
Tried and failed
policy gradient reinforcement learning with reranking reward applied to dialogue response generation. Outcome: worse than baseline. Reason: optimizing reranking rewards via policy gradient produced shorter and less informative outputs than test-time reranking
Specific response generation and elaborative discourse structure · UT Austin
Tried and failed
reinforcement learning with single task reward classifier applied to language model generation. Reason: optimizing solely for task reward caused repetitive degenerative outputs without fluency regularization
Improving Data Efficiency and Availability for Alignment · Cornell
Tried and failed
format compliance reward in reinforcement learning applied to code generation agents. Reason: agents exploited format-compliance rewards to gain score without actually solving the underlying tasks
An Empirical Study of Reward Hacking in LLM Coding Agent Training · Georgia Tech
Tried and failed
reinforcement learning with domain-specific reward only applied to conditional text generation. Outcome: worse than baseline. Reason: optimizing solely for domain-specific metrics degraded natural language fluency and readability
Effective Modeling in Medical Imaging with Constrained Data · MIT
Tried and failed
reinforcement learning with misconfigured reward mechanism applied to traffic simulation control. Outcome: no signal. Reason: improper reward function prevented policy quality differentiation, causing persistent random exploration without learning
Bridging the gap: unifying transportation planning and operations through enhanced travel demand modeling · UT Austin
Tried and failed
adversarial inverse reinforcement learning with positive rewards applied to episodic goal-reaching tasks. Reason: positive rewards near the target incentivized hovering indefinitely rather than terminating the episode
Estimation and control of visitation distributions for reinforcement learning · UT Austin
Considered and rejected
Considered and rejected: Format-compliance reward shaping term; rejected because the policy optimized format compliance to collect guaranteed reward without solving the underlying coding tasks.
An Empirical Study of Reward Hacking in LLM Coding Agent Training · Georgia Tech
Considered and rejected
Considered and rejected: Rejected manually assigning constant rewards to door-opening behaviors in KeyCorridor because it causes reward hacking (repeatedly opening/closing doors).
Considered and rejected
Considered and rejected: Using raw task rewards in multitask NetHack, rejected because frequent Scout rewards caused the agent to ignore Score and Gold tasks
Reading to Learn · ResearchWorks
Tried and failed
learned reward model for reinforcement learning applied to mathematical reasoning benchmarks. Outcome: worse than baseline. Reason: answer-format drift degraded validation accuracy while increasing energy consumption compared to rule-based verification
No Free Lunch for Hungry Machines: A Systems Investigation of Reinforcement Learning · Harvard
Tried and failed
active preference learning with missing relevant features applied to reward function learning. Outcome: worse than baseline. Reason: feedback is misattributed to observed features due to omitted task-relevant features, degrading learned reward quality
Robot See, Robot Do: On the Development of Robust and Adaptive Imitation Learning for Robots · Virginia Tech
Tried and failed
resource-constrained reinforcement learning with factorized action spaces applied to neural architecture search. Outcome: worse than baseline. Reason: soft penalty rewards fail under factorized search spaces due to invalid independence assumptions across layers
Automated Machine Learning under Resource Constraints · Cornell
Tried and failed
reinforcement learning with standard reward formulations applied to discrete structural truss optimization. Outcome: worse than baseline. Reason: standard reward functions produced suboptimal designs with higher final weight than baselines
Machine Learning Applications in Structural Analysis and Design · Virginia Tech
Reinforcement learning optimization suffers from training instability and value estimation errors
Algorithms frequently experience training collapse, severe overfitting to high-dimensional state spaces, or divergence caused by negative entropy coefficients and uncompensated estimation errors. In addition, stale value targets, overestimation bias in actor-critic frameworks, and long planning horizons amplify stochasticity and impede convergence.
Tried and failed
regression-based Q-function training applied to constrained reinforcement learning. Outcome: worse than baseline. Reason: yielded suboptimal performance compared to distributional RL classification approaches
Risk-Aware Reinforcement Learning with Safety Constraints · MIT
Tried and failed
reinforcement learning on high-dimensional raw feature space applied to portfolio management and stock selection. Outcome: overfit. Reason: extremely high-dimensional state space with raw tabular features caused severe overfitting to training data
Tried and failed
reinforcement learning with negative entropy coefficient applied to reinforcement learning policy optimization. Outcome: unstable. Reason: a negative entropy coefficient severely penalized exploration, leading to training collapse and dropped returns
Computational Simulation and Machine Learning for Quality Improvement in Composites Assembly · Virginia Tech
Tried and failed
Adaptive reinforcement learning with dynamic policy shaping applied to interactive game policy learning. Outcome: worse than baseline. Reason: uncompensated initial estimation errors degraded performance under negative feedback
Learning to Teach: Models for Semantic and Adaptive Personalization of AI Tutors · MIT
Tried and failed
reward shaping in reinforcement learning applied to thermal dynamics control. Outcome: did not converge. Reason: configurations either failed to converge or resulted in significantly slower training compared to gradient modification baselines
Tried and failed
value-aware model learning without target value updates applied to continuous control model-based reinforcement learning. Outcome: worse than baseline. Reason: optimizing against stale value function estimates degrades dynamic model accuracy below standard maximum likelihood estimation
Leveraging Value-awareness for Online and Offline Model-based Reinforcement Learning · Georgia Tech
Tried and failed
on-policy reinforcement learning with stochastic exploration applied to reservoir history matching. Outcome: too slow. Reason: stochastic exploration caused the agent to take suboptimal actions, wasting significant compute before recovery
Machine Learning Based Algorithms for Improving Forecasting in Subsurface Energy Resources · MIT
Tried and failed
deep Q-learning for interactive control applied to human-robot social interaction. Outcome: did not converge. Reason: insufficient training episodes and unmodeled human response latency causing reward misclassification
Designing Socio-Spatial Interfaces for Embodied and Situated Social Interaction · Cornell
Tried and failed
Deep Deterministic Policy Gradient reinforcement learning applied to continuous control in power systems. Outcome: unstable. Reason: Q-value overestimation bias leading to degraded policy performance and training instability
Physics Informed Learning-Based Frequency Regulation and Virtual Inertia Control for Renewable Energy Generators in Power Systems · Carleton University Institutional Repository
Tried and failed
reinforcement learning fine-tuning with transfer learning applied to power distribution network design optimization. Outcome: did not converge. Reason: large shifts in the target performance boundary rendered pre-trained representations insufficient for optimization convergence
Design Optimization of Power Delivery Networks in Packaging · Georgia Tech
Tried and failed
increasing discount factor and prediction horizon applied to reinforcement learning with model predictive control. Outcome: worse than baseline. Reason: longer horizons amplified the effect of high system stochasticity, degrading controller performance
Human-in-the-loop energy-efficient building HVAC control · Georgia Tech
Tried and failed
fixed predefined weights in value factorization applied to multi-agent reinforcement learning credit assignment. Outcome: worse than baseline. Reason: static weights cannot adapt to state-dependent agent contributions compared to dynamically learned weight coefficients
Reinforcement learning underperforms simpler classical control baselines and heuristics
Learned policies regularly achieve worse performance than standard PID controllers, greedy search heuristics, and static portfolio baselines. In specialized applications such as power grid operations and clinical treatment, reinforcement learning resulted in higher operational costs and severe constraint violations.
Tried and failed
reinforcement learning for dynamic resource allocation applied to closed-loop perception control. Outcome: worse than baseline. Reason: RL controller achieved lower detection recall than a simpler model-based PID baseline
Closed Loop Perception for Resource Efficient Autonomous Systems · Georgia Tech
Tried and failed
potential-based reward shaping in reinforcement learning applied to robot manipulation and locomotion benchmarks. Outcome: worse than baseline. Reason: failed to outperform heuristic-only policies under finite data regimes
Lost to a baseline
Recommender agent trained by deep reinforcement learning demonstrated lower human performance improvement compared to baseline agents exhibiting human-like behavior.
Reciprocal human-machine learning in manufacturing · DSpace-CRIS at TU Wien
Tried and failed
reinforcement learning and iterative learning control applied to real-time cellular adaptive control. Outcome: worse than baseline. Reason: outperformed by Bayesian filtering methods for online parameter adaptation
Exploring solutions to the sensing and actuation limitations within cybergenetics · Imperial
Tried and failed
tree-structured reinforcement learning applied to branching heuristics in exact symbolic computation. Outcome: worse than baseline. Reason: Failed to outperform standard hand-crafted variable selection heuristics
Lost to a baseline
Under NEWS2 reward, the naive weighted random baseline beat all offline RL and SL policies on WIS (-3.78) and WISt (-3.78) on the full sepsis test set.
Dynamic treatment regime for electronic health record · Oxford
Lost to a baseline
Proximal Policy Optimization (PPO) reinforcement learning performed less stably at later iterations than the simpler greedy search heuristic.
Optimizing Brain Stimulation for Parkinson's Disease, Memory Enhancement, and Optogenetic Control · Georgia Tech
Lost to a baseline
Model-free PPO was beaten by the static Markowitz portfolio across substantial portions of the cumulative reward empirical CDF under alpha decay dynamics.
Reinforcement learning for sequential decision-making: a data driven approach for finance · IRIS - SNS - prod
Tried and failed
reinforcement learning for discrete sequential optimization applied to power grid unit commitment. Outcome: worse than baseline. Reason: RL operating costs were 2x to 13x higher with severe constraint violation and load shedding
Secure and cost-effective operation of low carbon power systems under multiple uncertainties · Imperial
Tried and failed
model-free reinforcement learning control applied to commercial refrigeration energy management. Outcome: worse than baseline. Reason: agents struggled across variable operating horizons and dynamic price profiles compared to simple PI control
Design and deployment of data-driven control retrofits for energy systems in commercial buildings · Imperial
Tried and failed
standard reinforcement learning with standard reward functions applied to multi-agent deliberative reasoning alignment. Outcome: worse than baseline. Reason: performed substantially worse than few-shot direct preference optimization
Toward Deliberative AI: Multi-Agent LLMs for Real-World Reasoning · Virginia Tech
Policies fail to generalize across domain shifts and sim-to-real transfer
Simulated training environments often fail to model physical real-world dynamics, biomechanical variations, or extreme demand shifts encountered during deployment. Consequently, trained policies suffer substantial performance drops and fail to transfer zero-shot to real robotic hardware or unseen benchmark programs.
Tried and failed
direct sim-to-real transfer of reinforcement learning policies applied to embodied robotic navigation. Outcome: did not generalise. Reason: unmodeled real-world physical dynamics mismatch between the simulation environment and physical hardware
4D audio-visual learning: a visual perspective of sound propagation and production · UT Austin
Tried and failed
history-conditioned adaptive reinforcement learning policies applied to zero-shot sim-to-real robotic manipulation. Outcome: worse than baseline. Reason: adaptive policies conditioned on history transferred worse to the real world than simpler reactive policies
Exploring sim-to-real transfer for learning-based robot manipulation · Imperial
Tried and failed
sim-to-real reinforcement learning without human domain randomization applied to physical human-robot interaction policies. Outcome: did not generalise. Reason: simulated biomechanical models failed to capture the complexity and variability of real human biomechanics
Robotic Caregivers -- Simulation and Capacitive Servoing for Physical Human-Robot Interaction · Georgia Tech
Tried and failed
regularizing reinforcement learning against a teacher policy applied to robot manipulation skill transfer. Outcome: did not generalise. Reason: biased initialization degraded downstream transfer performance on difficult tasks
Learning Motion Policies for Dexterous Manipulation with Geometric Fabrics · Georgia Tech
Tried and failed
information-regularized actor-critic reinforcement learning applied to predicting human performance in unchunked states. Outcome: did not generalise. Reason: fitted policy complexity penalties failed to robustly predict empirical performance benefits across task states
Policy compression: Acting with limited cognitive resources · Harvard
Tried and failed
reinforcement learning with strict zero slack time applied to real-time dynamic bipartite matching. Outcome: did not generalise. Reason: marginal performance gains over greedy baseline and poor zero-shot scale transferability
Integrating Econometric Behavioral Models into Transportation Network Optimization · Cornell
Tried and failed
reinforcement learning for compiler pass ordering applied to code optimization sequences. Outcome: did not generalise. Reason: optimization sequences learned from random programs overfitted and degraded performance on unseen benchmarks
Tried and failed
unregularized reinforcement learning applied to human-agent multi-agent coordination. Outcome: did not generalise. Reason: agents diverged from human play styles, reducing action prediction accuracy and cooperation
Building Strategic AI Agents for Human-centric Multi-agent Systems · MIT
Lost to a baseline
Customized vehicle automation using expert-coded rewards achieved 61.67% prediction accuracy on unobserved neutral trips, losing to the non-customized baseline of 65.33%
Toward Trust-calibrated Customized Vehicle Automation · ResearchWorks
Tried and failed
reinforcement learning integrated with integer linear programming applied to dynamic matching under demand surges. Outcome: did not generalise. Reason: performance declined and was worse than baseline during out-of-distribution transfer under extreme demand
Integrating Econometric Behavioral Models into Transportation Network Optimization · Cornell
Practitioners reject reinforcement learning due to high sample complexity and safety concerns
Reinforcement learning is repeatedly rejected in favor of imitation learning, search algorithms, or direct optimization because of slow convergence and extreme data requirements. Furthermore, unconstrained exploration poses dangerous physical risks to equipment and fails to ensure mandatory legal and operational constraints.
Considered and rejected
Considered and rejected: Reinforcement learning in the real-world, rejected due to risk of equipment damage from required negative/unsuccessful exploration trails
Perception Based UAV Path Planning for Fruit Harvesting · JScholarship
Considered and rejected
Considered and rejected: Full RL-based shared reward for formation control; rejected as its navigation accuracy and formation robustness were insufficient compared to hybrid position-based control.
Cooperative Payload Transportation by UAVs: A Model-Based Deep Reinforcement Learning (MBDRL) Application · Virginia Tech
Considered and rejected
Considered and rejected: Decided against Reinforcement Learning (DQN, PPO, SAC) due to lack of environment feedback, inability to handle expanding action spaces, and excessive data requirements.
Considered and rejected
Considered and rejected: Rejected pure Reinforcement Learning for multi-turn code generation, choosing imitation learning via one-step recoverable MDP reduction to avoid sparse reward exploration
Reasoning in the Wild · Cornell
Considered and rejected
Considered and rejected: Rejected relying solely on implicit RL or standard PPO reward shaping (distance travelled/collision avoidance) because it failed to enforce legal compliance such as stopping at red lights and yielding to emergency vehicles
Enhancing autonomous vehicle decision-making through scenario-based traffic rule integration · Imperial
Considered and rejected
Considered and rejected: Rejected policy gradient / reinforcement learning (REINFORCE) for policy optimization due to higher complexity, training instability, and inferior performance compared to Gumbel-Softmax.
Efficient deep learning for image and video understanding · OpenBU
Considered and rejected
Considered and rejected: Rejected multi-objective RLHF (Reinforcement Learning from Human Feedback) due to training instability, extensive hyperparameter tuning requirements, and computational inefficiency compared to MODPO.
Essays on Digital Content Strategies: Creation, Diffusion, and Monetization · Harvard
Considered and rejected
Considered and rejected: Standard Reinforcement Learning / MDP methods rejected because they cannot achieve better than O(sqrt(T)) regret and fail to exploit queueing/extreme-point network structure.
Learning-NUM: Utility Maximization in Stochastic Queueing Networks · MIT
Considered and rejected
Considered and rejected: Rejected Reinforcement Learning (RL) agents for decision modeling because they incur training overhead and require problem-specific DNN architectures compared to search algorithms.
Network Digital Twins: A Paradigm for Network Management Applications · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Rejected Policy Gradient reinforcement learning due to slow convergence requiring over 1,000 iterations to reach cumulative reward criteria
Preventing Breaks in Embodiment in Immersive Virtual Reality · EPFL
Learned reward models and off-policy evaluation methods introduce bias and noise
Off-policy estimators and inverse reward learning algorithms can severely misjudge policy quality when dynamics or observations are misspecified. In addition, learned reward models frequently inject noise or suffer from distribution shifts, leading agents to optimize against distorted supervisory signals.
Tried and failed
classical POMDP inverse reinforcement learning algorithms applied to reward learning under misspecified dynamics. Outcome: worse than baseline. Reason: algorithms suffer high estimation error when the forward agent's model dynamics or observation likelihoods are misspecified
Inverse Reinforcement Learning: A Microeconomics-Based Approach · Cornell
Considered and rejected
Considered and rejected: Rejected purely offline reinforcement learning / regression-based policy evaluation due to confounding and lack of penalization for suboptimal training decisions
Essays on Digital Transformation and Business Decision-Making · ResearchWorks
Lost to a baseline
Learned RL policy for sepsis management was inferior in true expected reward to the behavior policy despite outperforming it on Weighted Importance Sampling (WIS) and Model-Based off-policy evaluation.
Towards Rigorously Tested & Reliable Machine Learning for Health · MIT
Tried and failed
human feedback shaping with low-accuracy error signals applied to reinforcement learning convergence acceleration. Outcome: worse than baseline. Reason: decoding accuracy below threshold injected noise that confused the agent rather than guiding policy optimization
ON THE INTERPLAY BETWEEN BRAIN-COMPUTER INTERFACES AND MACHINE LEARNING ALGORITHMS: A SYSTEMS PERSPECTIVE · Georgia Tech
Lost to a baseline
Controlled decoding (CD) and value augmented sampling (VAS) guide with Q^{\pi_{ref}, 0} (unregularized Q-function) which fails to maximize reward (converging to reward 0.1 vs optimal 1) and incurs higher KL divergence compared to Q-sharp (Q♯).
Towards Safe, Efficient, and Steerable Reinforcement Learning · Cornell
Considered and rejected
Considered and rejected: Avoided evaluating QA-Feedback ChatGPT generations with trained reward models because T5-trained classifiers fail to generalize out-of-distribution.
Human-Centered Interactive Information Seeking · ResearchWorks
Considered and rejected
Considered and rejected: Rejected labeling dialogue data with learned reward model in favor of +-1 binary labels due to RM noise and inaccuracy
Towards a Theory and Practice of Open-Ended Reasoning with Generative Models · Georgia Tech
Considered and rejected
Considered and rejected: Rejected using learned reward models instead of value functions for state abstraction; value functions are smoother and capture long-term environmental effects.
Tried and failed
training reward models without grounding context applied to contextual evaluation preference modeling. Outcome: worse than baseline. Reason: lacking grounding context degrades performance to baseline non-contextual reward model levels
Text-Graph Encoders and Retrieval-Augmented Generation · EPFL
Left open by the authors
Problems the authors named and did not get to.
Left open
Learn neuro-symbolic low-level robotic skills using reinforcement learning instead of relying on expert demonstrations. Blocker: Lacks specific RL algorithm choices, reward formulations, and environment setups beyond general intent
Left open
Apply reinforcement learning fine-tuning on pre-trained imitation robot policies to achieve over ninety percent real-world deployment success rates. Blocker: Requires real-world physical robot hardware and physical experimental environments for deployment and evaluation.
Scaling robot learning with heterogeneous data from the real world, simulation, and the web · UT Austin
Left open
Deploy the language-guided reinforcement learning and adaptation models on physical robots using sim-to-real transfer and safe RL techniques. Blocker: Requires access to physical robotic hardware.
Using natural language to aid task specification in sequential decision making problems · UT Austin
Left open
Train graph neural networks using reinforcement learning for decentralized robot control instead of oracle imitation learning. Blocker: None
Left open
Transition FiLM-Nav to end-to-end reinforcement learning and safe real-world fine-tuning on robotic hardware. Blocker: Requires physical robot hardware and real-world testing environments for safe real-world fine-tuning
From Web to World: Harnessing Foundation Models for Intelligent Robotic Assistants in Real-World Environments · Georgia Tech
Left open
Evaluate differentiable MPC policies within offline reinforcement learning algorithms using static real-robot dataset benchmarks. Blocker: None
Learning Novel Strategies for Model Predictive Control by Leveraging Experience · ResearchWorks
Left open
Deploy and evaluate unsupervised model-based reinforcement learning agents on physical, real-world robotic systems to assess scalability. Blocker: Requires access to physical robotic hardware and real-world testing environments.
LEARNING TO ACT FROM DIVERSE DATA SOURCES VIA WORLD MODELS · Penn
Left open
Transfer learned reinforcement learning motion planning policies across varying robot kinematic chains and dynamic environments in simulation. Blocker: None
Left open
Deploy reinforcement learning policies safely onto physical collaborative robots in real-world industrial production environments under latency constraints. Blocker: Requires physical collaborative robot hardware and an industrial production testing environment.
Left open
Train reward models and apply RLHF to align a financial LLM's sentiment outputs for portfolio management. Blocker: None
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.