A dataset for studying how models interpret, follow, and generalize technical instructions derived from real research code, especially in research and agentic workflows.
Resource details →
A large corpus of research code referenced directly in scientific papers, designed to study how models reason about, modify, and execute real-world code in complex software systems.
Resource details →
A dataset of 16k AI safety-relevant papers from the ArXiv, enriched with structured metadata.
Resource details →
A collection of instruction-tuning datasets derived from scientific abstracts for studying how synthetic supervision shapes model behavior, overfitting, and generalization under distribution shift.
Resource details →
Question-answer datasets derived from ArXiv for evaluating retrieval and search behavior in technical domains, including retrieval failures, false confidence, and compounding RAG errors.
Resource details →
A filtered and LLM-enriched version of Allen AI's S2ORC corpus containing 1.1 million computer science papers with structured metadata for methods, models, datasets, metrics, compute usage, and limitations.
Resource details →
A dataset of 34GB of AI research data for supervised fine-tuning, designed to train models that can reason about and generate AI research.
Resource details →