Chapter Four · failure evidence

What Zero-Shot Learning got wrong, from 93 dissertations

Across these thesis studies, zero-shot learning methods frequently underperform supervised alternatives, fail under distribution shifts, and struggle with specialized reasoning tasks. They also suffer from severe prediction biases and fail to capture subjective human qualities or fine-grained physiological and acoustic patterns. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Zero-shot methods lag behind supervised models and simple baselines across benchmarks

19 theses · 10 institutions

Zero-shot foundation models and prompting configurations consistently achieve lower accuracy, recall, and precision than fully trained or fine-tuned baselines across vision and language tasks. In several evaluations, zero-shot models even underperformed basic rule-based systems, random baselines, or few-shot prompts.

Lost to a baseline

On Multi30K German-to-image search, thesis English-only zero-shot model (31.4% R@1) lost to supervised MHA-D (40.3% R@1)

Learning and interpreting deep representations from multi-modal data · Oxford

Lost to a baseline

All zero-shot models achieved lower raw Recall@K across all cutoffs compared to fully trained models (IMP, GPSNet, Motifs, VCTree, OpenPSG)

Zero-Shot Scene Graph Relationship Prediction using VLMs · Virginia Tech

Lost to a baseline

UnifiedQA zero-shot answering on StrategyQA scored 58.95% accuracy, heavily underperforming GPT-3 zero-shot prompting (58.08% / 63.32% few-shot) and other reasoning baselines when not given gold supporting facts.

Incidental Supervision for Natural Language Understanding · Penn

Lost to a baseline

Vanilla zero-shot style transfer failed to return a valid response 25.4% of the time, compared to 0.6% for augmented zero-shot learning.

Understanding the Limitations of Using Large Language Models for Text Generation · Penn

Lost to a baseline

Zero-shot foundation model mPLUG-Owl achieved inferior or comparable precision/recall (0.20 precision / 0.20 recall on multiple groundings) compared to naive single/multiple baselines on the Single Answer Grounding Challenge.

Visual question answering : representing authentic use cases and ambiguity · UT Austin

Lost to a baseline

Zero-shot pretrained models (R-AUROC ~0.57-0.60) lost to many-shot ResNet-50 trained directly on downstream target classes (R-AUROC 0.68)

Uncertainties of Latent Representations in Computer Vision · Publikationssystem UB Tuebingen

Lost to a baseline

In-domain zero-shot German question formation for mBART failed to correctly transform the full sequence (0% sequence accuracy), underperforming mT5.

Emergent Syntactic Behaviors and Mechanisms in Neural Language Models · JScholarship

Lost to a baseline

Zero-shot CLIP OOD methods (MCM and Cosine) achieved 67.1% to 79.8% FPR on NINCO, failing to beat standard IN-1K fine-tuned classifiers

Out-of-Distribution Detection, Sharpness, and Unlearning: Advancing Robust and Trustworthy Deep Learning · Publikationssystem UB Tuebingen

Tried and failed

domain fine-tuned zero-shot language model applied to time series forecasting. Outcome: worse than baseline. Reason: exhibited performance degradation across all horizons and metrics compared to standard deep learning models

Deep Learning Methods for Built Environment Operational Management · Virginia Tech

Tried and failed

zero-shot prompting with abstract category definitions applied to program semantic reasoning. Outcome: worse than baseline. Reason: lacking concrete few-shot examples degraded performance relative to direct prediction

CodeSense: A real-world benchmark and dataset for code semantic reasoning · Iowa State

Lost to a baseline

Zero-shot and few-shot Llama 3 in-context learning achieved F1-scores of only 0.29–0.40, falling well behind the non-LLM BiLSTM baseline F1 of 0.67.

Surgical Site Infection (SSI) Identification Across Multiple Facilities and Surgery Types Using Multimodal Data and Deep Learning · ResearchWorks

Lost to a baseline

Zero-shot prompting on several LLMs (including Llama3 8B and Mistral 7B) achieved lower accuracy than the rule-based VADER sentiment analysis baseline

Leveraging social media data and large language models for understanding public health behaviours in the context of infectious diseases and vaccination · EPFL

Lost to a baseline

On C-GQA open-world compositional zero-shot learning with zero-shot CLIP, GloVe baseline beat FLM on unseen class accuracy (3.92% vs 2.62%) and AUC (0.20% vs 0.16%).

Learning under Limited Supervision for Better Generalization · Publikationssystem UB Tuebingen

Lost to a baseline

On ActivityNet-QA VQA, supervised finetuned JustAsk (FT) achieved 64.667% vs this method's zero-shot 61.168%.

Learning Generalizable Systems by Learning Composable Energy Landscapes · MIT

Lost to a baseline

DiffuCOMET-Fact (RA-F1 43.51) lost to Beam-COMET baseline (RA-F1 45.88) on out-of-domain MovieSummaries zero-shot testing.

Understanding Implicit Schemas of Reasoning · EPFL

Lost to a baseline

Zero-shot Kaleido 3B achieved 71.5% on ETHICS Commonsense, losing to few-shot ChatGPT (80.3%).

Steps Towards the Pluralistic Alignment of Language Models · ResearchWorks

Tried and failed

pretrained vision-language and self-supervised models for keypoint detection applied to deformable object semantic keypoint localization. Outcome: worse than baseline. Reason: zero-shot VLM and DINOv2 backbones lacked fine spatial localization precision compared to supervised ResNets

Deformable Object Manipulation with a Tactile Reactive Gripper · MIT

Lost to a baseline

On few-shot HICO classification (Few@1, Few@5, Few@10), zero-shot HTS ViT-B/32 (33.6, 36.3, 37.2 mAP) was beaten by supervised HTS* ViT-B/32 (49.6, 52.6, 52.9 mAP).

Crossing the Chasm: Bridging Perception and Generative Models for Enhanced Vision-Language Understanding · ResearchWorks

Lost to a baseline

Zero-shot and decoder-fit DNNs achieved lower accuracy (barely above 8.33% chance at 35% fragments) compared to humans (~50% accuracy).

Differences in Visual Perception in Humans and Deep Neural Networks · EPFL

Zero-shot models struggle with specialized domain knowledge and structured reasoning

17 theses · 8 institutions

Pretrained models fail when deployed on highly technical tasks including medical grounding, chemical reaction prediction, code synthesis, and temporal knowledge reasoning. Without domain-specific supervision or intermediate feedback, they produce ungrounded predictions, generate invalid program code, or trigger safety guardrails.

Tried and failed

zero-shot vision-language contrastive classification applied to fine-grained multi-class pathology subtyping. Outcome: did not generalise. Reason: lack of supervision causes zero-shot models to fail on complex, fine-grained multi-class classification tasks

Data-Driven General Purpose Foundation Models for Computational Pathology · MIT

Tried and failed

zero-shot text-prompted segmentation models applied to medical image finding grounding. Outcome: no signal. Reason: pretrained models failed completely on complex free-text clinical finding descriptions

AI Systems for Understanding and Grounding Radiology Reports · Harvard

Tried and failed

direct zero-shot diagnostic prompting of multimodal models applied to medical image classification. Reason: unengineered direct diagnostic queries triggered commercial safety guardrails, causing high refusal rates

Augmenting medical image classifiers with synthetic data across populations · Harvard

Tried and failed

zero-shot transformer reaction outcome predictor applied to unseen chemical reaction classes. Outcome: did not generalise. Reason: Models fail to infer transformation mechanisms for novel reaction types absent from training data

Predictive chemistry: machine learning for reaction deployment, reaction development, and reaction discovery · MIT

Tried and failed

direct zero-shot prompting of large language models applied to clinical note information extraction. Outcome: did not generalise. Reason: failed to adapt and performed poorly on advanced reasoning models without structured reasoning or fine-tuning

Machine Learning Approaches for Drug Combination Discovery · Cornell

Tried and failed

zero-shot sequence-to-sequence pretrained model classification applied to claim verification verdict classification. Outcome: worse than baseline. Reason: lack of task-specific fine-tuning severely limited classification capability compared to fine-tuned alternatives

Explainable Neural Claim Verification Using Rationalization · Virginia Tech

Tried and failed

instruction-tuned vision-language model applied to zero-shot visual question answering. Outcome: worse than baseline. Reason: prior instruction tuning biased the model toward coarse-grained answers rather than fine-grained predictions

Extracting Knowledge with Multimodal and Multilingual Intelligent Systems · Georgia Tech

Tried and failed

zero-shot transformer reranker on beam search candidates applied to neural code generation candidates. Outcome: worse than baseline. Reason: compounded overfitting from using an un-finetuned reranker

Neuro-symbolic program generation and execution for hybrid reasoning · Iowa State

Considered and rejected

Considered and rejected: Rejected relying solely on in-context prompting (zero-shot, few-shot, and Chain-of-Thought) for specialized legal construct classification due to poor performance and severe class imbalance errors.

Modeling Legal Constructs · Cornell

Tried and failed

zero-shot vision-language models applied to robot trajectory planning and motion reasoning. Outcome: did not generalise. Reason: models exhibit severe forward-motion and deceleration biases and fail at temporal motion reasoning

Building Intelligence that can Interact with the Physical World · MIT

Tried and failed

zero-shot code generation using pretrained LLM applied to high-level synthesis code generation. Reason: pretrained model lacked specialized domain knowledge and constructs required for valid high-level synthesis code

Harnessing Large Language Models Towards More Accessible Hardware Accelerator Design · Georgia Tech

Tried and failed

zero-shot large language model generation without exemplars applied to generating software vulnerability exploits. Outcome: no signal. Reason: the model lacked domain-specific context and test patterns needed to synthesize functional exploit payloads

Secure Coding Practice in Java: Automatic Detection, Repair, and Vulnerability Demonstration · Virginia Tech

Tried and failed

LLM zero-shot reward function code generation applied to dexterous robotic hand manipulation tasks. Reason: Generated reward code failed to execute consistently across multiple random seeds.

Personalized, Safe, and Interactive Robot Programming via Human Demonstrations · Georgia Tech

Tried and failed

zero-shot direct prompting of large language models applied to resolving semantic ambiguities in text. Reason: Failed to resolve complex relational and temporal dependencies without intermediate feedback or attributes

Transforming Free-Form Sentences into Sequence of Unambiguous Sentences with Large Language Model · Virginia Tech

Lost to a baseline

Zero-shot BLINK achieved 54.4% accuracy on NLM-Chem, missing domain-specific performance

Entity Linking in Low-Annotation Data Settings · JScholarship

Tried and failed

zero-shot in-context learning with large language models applied to entity linking disambiguation. Outcome: worse than baseline. Reason: struggled with fine-grained candidate entity disambiguation without fine-tuning

Information extraction with weak supervision · Iowa State

Tried and failed

zero-shot large language model reasoning applied to temporal knowledge graph link prediction. Outcome: did not generalise. Reason: models fail to understand structured graph entities and temporal relations without specialized fine-tuning

Temporal link prediction in the wild · Imperial

Models fail to capture subjective, affective, and nuanced human evaluations

14 theses · 7 institutions

Zero-shot prompting produces near-random accuracy or negative correlation when evaluating subjective qualities, emotional states, and rhetorical attributes. Across persona modeling, stance detection, and qualitative thematic extraction, unadapted models fail to match human consensus ratings and benchmark standards.

Tried and failed

zero-shot autoregressive language model generation and classification applied to logical fallacy prediction. Outcome: did not generalise. Reason: models fail to identify and distinguish unseen fallacy categories in zero-shot settings

Extending Provenance for Understanding Claims and Data Analyses · Penn

Tried and failed

Zero-shot prompting of large language models applied to thematic analysis across study groups. Outcome: did not generalise. Reason: Prompts alone struggled to synthesize generalized themes and perform cross-group comparisons across qualitative data.

Embodied Virtual Reality: The Impacts of Human-Nature Connection During Engineering Design · Virginia Tech

Tried and failed

zero-shot multimodal foundation model classification applied to image emotion recognition. Outcome: worse than baseline. Reason: general pre-trained embeddings lack domain-specific fine-tuning on nuanced emotion distributions

Three essays on online word-of-mouth and user behavior in online environments · Iowa State

Tried and failed

zero-shot large language model classification applied to persona relation extraction. Outcome: worse than baseline. Reason: Individual model predictions were unreliable compared to aggregated human majority voting.

Understanding Implicit Schemas of Reasoning · EPFL

Tried and failed

zero-shot large language model generation applied to automated customer review responses. Outcome: worse than baseline. Reason: generated text significantly impaired customer perceived helpfulness across varied sampling temperatures

Strategizing in Response to Environmental Uncertainty in the Hospitality Industry: A Data-Analytical Approach · Virginia Tech

Lost to a baseline

In conversational Turing tests, zero-shot ChatGPT performed poorly as an AI judge and was outperformed by one-shot prompted ChatGPT and trained SVM classifiers.

Finding Structure in Human Cognition: From Temporal Order Codes in Memory to Behavioral Signatures in Vision and Language · Harvard

Lost to a baseline

Zero-shot persona gpt-4o-mini performed worse than the theoretical random baseline (Qini -0.167 vs 0.000).

Language Models as Opinion Models: Techniques and Applications · MIT

Lost to a baseline

Zero-shot text2text gpt-4o-mini produced worse-than-random HTE prediction (Qini -0.052 vs baseline 0.183).

Language Models as Opinion Models: Techniques and Applications · MIT

Considered and rejected

Considered and rejected: Direct zero-shot classification with off-the-shelf VLMs without fine-tuning, rejected due to poor performance on affective tasks

Three essays on online word-of-mouth and user behavior in online environments · Iowa State

Considered and rejected

Considered and rejected: Rejected directly using LLM-generated zero-shot stance predictions in favor of extracting linguistic/subsequence rationales

Advancing stance detection and fine-grained content analysis for socially relevant domains · Leibniz Universität Hannover Repository

Tried and failed

zero-shot prompt classification of isolated short phrases applied to narrative frame identification. Outcome: no signal. Reason: isolated phrases lacked sufficient context, making direct frame elicitation too vague and unreliable

ON THE NARRATIVE CONSTRUCTION OF REALITY IN POLITICAL DISCOURSE: A THEORY, A METHOD, AND EMPIRICAL STUDY · Penn

Tried and failed

zero-shot LLM numerical attribute scoring applied to rhetorical attribute extraction from text. Outcome: no signal. Reason: LLM scores for nuanced subjective attributes failed to correlate with human benchmark ratings

Preaching to the Choir: An AI-Based Analysis of Religious Demand in U.S. Church Sermons, 2000-2023 · Harvard

Tried and failed

zero-shot vision-language classification with abstract keywords applied to visual art genre classification. Reason: abstract aesthetic keywords caused semantic leakage and false positive over-assignment across unrelated categories

Famous Views of Economics: An AI-Based Analysis of Japanese Woodblock Print Subject Matter and Economic Circumstances · Harvard

Tried and failed

zero-shot large language model prompting applied to nuanced text annotation tasks. Outcome: no signal. Reason: models achieved near-random performance on subjective constructs

More Than a Sum of Its Parts: Teamwork Through a Multidimensional Lens · Penn

Lost to a baseline

Zero-shot text2text Pythia-410m (untrained) produced an individual-dataset Qini (-0.063) that was worse than the ensemble baseline (0.183).

Language Models as Opinion Models: Techniques and Applications · MIT

Lost to a baseline

Unsupervised SimCSEbase and supervised FlanT5large achieved negative Spearman correlation (-3.0 and -0.43) on zero-shot Conditional Semantic Text Similarity.

Methods and Challenges In Inference Across Textual Sources · Penn

Lost to a baseline

Zero-shot GPT-4 achieved only 44% accuracy matching human ratings across rating constructs.

Procedural Justice for CS Hiring: Towards Fair and Equitable Resume Matching in AI Era · Virginia Tech

Severe domain shifts degrade zero-shot transfer across visual and physical environments

16 theses · 9 institutions

Vision and spatial models trained on standard distributions fail when evaluated on histopathology, satellite, aerial, aquatic, or simulated robotic environments. The lack of task-specific adaptation causes severe false negatives, unstable segmentations, and semantic confusion across disparate domains.

Tried and failed

zero-shot segmentation using pre-trained network applied to out-of-distribution histopathology images. Outcome: did not generalise. Reason: produced unstable nuclear segmentations and unreliable classification without fine-tuning and threshold calibration

Integrated Molecular and Histopathological Studies of Testicular Germ Cell Tumor Biology · Harvard

Tried and failed

zero-shot vision-language model classification applied to satellite imagery object detection. Outcome: no signal. Reason: poor zero-shot alignment on domain-specific satellite imagery led to zero true positives

Adaptive Learning for AI Systems: AI Methods for Scientific Domains under Limited Supervision · Cornell

Tried and failed

deep learning PDE solver direct zero-shot inference applied to out-of-distribution physical scattering fields. Outcome: did not generalise. Reason: test inputs had structures completely uncorrelated with the synthetic training distribution

Exploring Novel Modalities for Optical Diffraction Tomography · EPFL

Tried and failed

zero-shot transfer of task-finetuned conditional adapters applied to unseen visual generation tasks. Outcome: did not generalise. Reason: specialized finetuning destroyed zero-shot generalization capability on out-of-distribution conditioning tasks

Enhancing generative model efficiency and control for data generation, reinforcement learning, and robotics · UT Austin

Tried and failed

unprompted zero-shot foundation model segmentation applied to complex morphology cell microscopy images. Outcome: did not generalise. Reason: zero-shot segmentation failed on complex star-shaped developmental morphologies without task-specific prompting or fine-tuning

IDCC-SAM: Automated cell counting in immunocytochemistry using the segment anything model · Iowa State

Tried and failed

zero-shot pretrained cell segmentation model applied to unseen tissue histopathology classification. Outcome: did not generalise. Reason: the specific target cancer type was underrepresented in the pretraining dataset

Spatially aware deep learning for clear cell renal cell carcinoma characterization and discovery · Harvard

Tried and failed

zero-shot object detection using off-the-shelf models applied to floating aquatic debris identification. Outcome: did not generalise. Reason: Domain shift between standard training imagery and aquatic surface imagery caused near-zero detection accuracy.

Autonomous System for Identifying and Capturing Floating Waste · Georgia Tech

Tried and failed

zero-shot transfer of real-trained segmentation models applied to synthetic simulation environments. Outcome: did not generalise. Reason: Domain shift between real-world training imagery and synthetic rendering distributions.

Perception Enabled Planning for Autonomous Systems · Cornell

Tried and failed

zero-shot pre-trained instance segmentation applied to indoor robotic object detection. Outcome: did not generalise. Reason: distribution shift between pre-training dataset and deployment environment caused significant false negatives on visible objects

Autonomous exploration and object reconstruction with an MAV · Imperial

Tried and failed

zero-shot transfer of self-supervised image clustering applied to satellite imagery across geographic regions. Outcome: did not generalise. Reason: domain shift between geographic locations caused semantic confusion across visually distinct feature clusters

Characterising urban environments in Sub-Saharan Africa with satellite imagery and unsupervised deep learning · Imperial

Tried and failed

direct zero-shot cross-dataset transfer of object detectors applied to multi-object tracking benchmark evaluation. Outcome: did not generalise. Reason: domain shift between source and target datasets degraded detection accuracy without target fine-tuning

Improved 2D Camera-Based Multi-Object Tracking for Autonomous Vehicles · Virginia Tech

Tried and failed

zero-shot image-to-image translation GAN applied to cross-sensor aerial imagery. Outcome: did not generalise. Reason: sensor and domain distribution shifts between satellite training data and aerial test imagery

Deep Learning Emulators for Accessible Climate Projections · MIT

Tried and failed

Pre-trained video moment retrieval model applied to cross-domain video moment queries. Outcome: did not generalise. Reason: Severe zero-shot domain mismatch between pre-training distribution and evaluation moment queries.

User-centered Programmatic Data Labeling · Georgia Tech

Tried and failed

zero-shot transfer of pre-trained document layout models applied to long-form academic dissertations. Outcome: did not generalise. Reason: Distribution shift between short research paper training sets and complex long-form dissertation formats.

Analyzing and Navigating Electronic Theses and Dissertations · Virginia Tech

Tried and failed

pretrained cross-modal feature matching networks applied to zero-shot airborne 2D-3D data. Outcome: did not generalise. Reason: extreme domain shift between ground training data and airborne perspectives yielding negligible inlier ratios

Concurrent adjustment of active and passive optical sensors with GNSS and raw inertial data · EPFL

Tried and failed

zero-shot policy transfer across network topologies applied to optimization algorithm parameter tuning. Outcome: did not generalise. Reason: policies trained on one graph structure fail on differently sized or structured graphs without retraining

Designing policy optimization algorithms for multi-agent reinforcement learning · Georgia Tech

Signal and physiological variability cause zero-shot transfer to fail in acoustic and biological domains

11 theses · 9 institutions

Zero-shot models experience drastic performance drops when transferring across different human subjects, atypical vocal patterns, or unseen language phonetics. Generic pretrained representations fail to capture subject-specific functional features or subtle acoustic distinctions without individualized tuning.

Tried and failed

zero-shot cross-subject classification from decomposed signals applied to dexterous movement intent decoding. Outcome: did not generalise. Reason: inter-individual variability in physiological signal representations degraded generalization performance across subjects

Decoding peripheral neural correlates of dexterous movements · Imperial

Tried and failed

zero-shot cross-subject transfer of neural networks applied to fMRI brain response prediction. Outcome: did not generalise. Reason: models trained exclusively on other individuals failed to capture subject-specific functional representations without individualized readouts

Characterizing human vision through large-scale brain imaging and computational models · MIT

Lost to a baseline

Direct zero-shot Seq2Seq transfer yielded near or exceeding 100% WER, severely underperforming modular hybrid DNN-HMM systems.

Automatic Speech Recognition without Transcribed Speech or Pronunciation Lexicons · JScholarship

Tried and failed

zero-shot transfer learning with pretrained audio embeddings applied to idiosyncratic atypical vocalization classification. Outcome: no signal. Reason: generic pretrained representations failed to capture idiosyncratic acoustic patterns of atypical vocalizations

Foundations of Cognitive, Affective, and Communicative Systems for Neurodiverse Individuals · MIT

Tried and failed

zero-shot multimodal large language models applied to audio event extraction. Outcome: no signal. Reason: Models completely failed to extract audio events, resulting in zero F1 performance across all tests

M3EC: A Multimodal Multidocument Benchmark for Event Extraction and Coreference Resolution · Virginia Tech

Tried and failed

zero-shot cross-lingual transfer of sub-phonetic representations applied to unseen target language speech recognition. Outcome: did not generalise. Reason: source language lacked phonetic distinctions (like voicing contrasts) required to resolve target language syllables

Modeling of Language-Universal Speech Attributes for Multilingual Speech Recognition and Processing · Georgia Tech

Lost to a baseline

In zero-shot multilingual SLU intent classification, the proposed zero-shot Speech-LLM (16.24%–47.38%) lost to the text NLU baseline (70%–80%) and cascaded SLU across all evaluated languages.

Multilingual Spoken Language Understanding: Efficient Speech Dataset Collection, Architectural Exploration, and Zero-Shot SLU · IRIS - UNITN - prod

Lost to a baseline

Generic zero-shot AudioSet transfer learning achieved only 51.1% accuracy on self-talk, performing barely above chance

Interfaces and Models for Improved Understanding of Real-World Communicative and Affective Nonverbal Vocalizations by Minimally Speaking Individuals · MIT

Lost to a baseline

WaveNet slightly outperformed DiT in zero-shot SECS (0.307 vs 0.299) and CER (4.06% vs 4.43%) on GenerSpeech

Efficient Adaptation for Speech Technology · EPFL

Considered and rejected

Considered and rejected: Decided against using zero-shot automatic phoneme recognition models (e.g. XLS-R FAIR) because they learn surface-level acoustics without linguistic/phonemic contrasts.

Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data · ResearchWorks

Lost to a baseline

In zero-shot gaze prediction on whole signals, high-privacy autoencoder signals (AE-0.1, AE-0.2) achieved worse prediction error (2.27°, 2.52°) than the baseline no-prediction approach.

Toward Privacy-Preserving Eye Tracking with Applications in Cross-Platform and Extended Reality Environments · TXST Digital Repository

Models exhibit severe class prediction bias, label imbalance, and high false alarm rates

10 theses · 7 institutions

Zero-shot classifiers and natural language inference pipelines frequently collapse into majority-class predictions or over-predict neutral categories. Uncalibrated thresholding and prompt sensitivity also lead to elevated false positive rates and poor recognition of rare classes.

Tried and failed

NLI-based zero-shot classification applied to thematic categorization of domain-specific text. Outcome: did not generalise. Reason: the model exhibited severe label bias, classifying the vast majority of samples into a single category

Capturing Elusive Metrics in Economics · unevada

Considered and rejected

Considered and rejected: Using purely zero-shot NLP classification without deterministic adjustments; rejected because models favored brief answers and scored unrelated phrases.

Weaving AI into Society: The Co-Evolution of Artificial Intelligence and the social fabric of organizations · IRIS - LUISS - prod

Tried and failed

prompt-based zero-shot large language model classification applied to student text behavior classification. Outcome: worse than baseline. Reason: produced substantially lower AUC and more than double the false positives compared to embedding-based classifiers

MEASURING AND UNDERSTANDING STUDENTS’ SELF-REGULATED LEARNING IN TEXTUAL DATA IN COMPUTER-BASED LEARNING ENVIRONMENTS · Penn

Tried and failed

zero-shot prompt-based LLM classification applied to multi-class text categorization. Outcome: worse than baseline. Reason: None

Evaluating Human-LLM Alignment in ETD Subject Classification New Trends in Theory and Practice of Digital Libraries, TPDL 2025 · Virginia Tech

Tried and failed

zero-shot prompting of large language models applied to time series anomaly detection. Outcome: worse than baseline. Reason: produced high false alarm rates or generic text and code instead of precise anomaly intervals

Machine Learning Systems for Unsupervised Time Series Anomaly Detection · MIT

Tried and failed

Adding predicate confidence score to scoring function applied to parse tree simplification for event extraction. Outcome: worse than baseline. Reason: Model over-relied on learned language priors rather than generalising in the zero-shot setting

Towards Explainable Event Detection and Extraction · Virginia Tech

Tried and failed

zero-shot stance detection using general-purpose embedding models applied to political and congressional speech text. Outcome: did not generalise. Reason: lacked domain-specific adaptation, causing models to heavily over-predict neutral labels

Words Speak As Loudly As Actions: Deep Learning Methods For Stance-Based Ideal Points From Congressional Speeches · Harvard

Tried and failed

zero-shot LLM classification on rare classes applied to noisy text transcript classification. Outcome: data insufficient. Reason: low prevalence and high semantic ambiguity resulted in poor F1 scores

Local Television News in a Changing Media Environment · Penn

Tried and failed

zero-shot NLI for text classification applied to domain-specific technical requirements classification. Outcome: did not generalise. Reason: severe class prediction bias caused poor performance on specialized domain text

Standardization of Engineering Requirements using Large Language Models · Georgia Tech

Tried and failed

zero-shot co-training with soft prompting applied to weakly supervised question answering. Outcome: worse than baseline. Reason: extremely high label noise rate prevented pseudo-label refinement from improving over the initial zero-shot model

Learning from Weak Supervision: Theory, Methods, and Applications · MIT

Left open by the authors

Problems the authors named and did not get to.

Left open

Develop zero-shot or transfer learning AI prognostic systems for newly deployed engineering assets lacking operational or failure data. Blocker: The task is an open-ended research direction without specific target metrics, assets, or concrete methodologies proposed.

Advanced data-driven methods for prognostics and life extension of assets using condition monitoring and sensor data. · Cranfield

Left open

Evaluate generalized zero-shot learning for cross-modal IMU activity recognition by classifying both seen and unseen classes at test time. Blocker: None

Cross-modal learning from visual information for activity recognition on inertial sensors · Oxford

Left open

Develop and evaluate methods to resolve zero-shot classification failure of BioTrove-CLIP models on the Confounding-Species benchmark. Blocker: None

Curating the world’s largest biodiversity dataset for AI · Iowa State

Left open

Reduce the zero-shot scene graph relationship prediction performance gap between small and large vision-language models for resource-constrained environments. Blocker: None

Zero-Shot Scene Graph Relationship Prediction using VLMs · Virginia Tech

Left open

Improve zero-shot vision-language models for scene graph generation to match or exceed supervised recall benchmarks while preserving generalization. Blocker: None

Zero-Shot Scene Graph Relationship Prediction using VLMs · Virginia Tech

Left open

Evaluate the effect of representation dimension and dataset coverage on Proto-Successor Measure zero-shot RL performance. Blocker: None

Reinforcement learning beyond rewards : decision-making in the language of visitation distributions · UT Austin

Left open

Evaluate zero-shot, one-shot, and few-shot prompting on LLM annotation performance for pedagogical discourse moves in classroom transcripts. Blocker: Access to the annotated elementary classroom transcript dataset.

Comparing human and generative AI annotation performance: Understanding linguistic features of teacher talk in classroom discussions · Iowa State

Left open

Evaluate the zero-shot stance detection pipeline across additional domains and distinct NLP task settings beyond climate change. Blocker: None

Advancing stance detection and fine-grained content analysis for socially relevant domains · Leibniz Universität Hannover Repository

Left open

Evaluate alternative distance metrics, prototype encodings, and transductive zero-shot learning settings in Logic Tensor Networks with vision backbones. Blocker: None

Robust machine learning models for high dimensional data interpretation · IRIS - POLITO - prod

Left open

Apply few-shot, zero-shot, and one-shot learning methods to improve classification performance on tail classes in NSL-KDD, UNSW-NB15, and CIC-IDS-2018. Blocker: None

Modeling the Abnormality: Machine Learning-based Anomaly and Intrusion Detection in Software-defined Networks · unevada

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.