New research / Forthcoming

Causal Tests of Language-Model Reports

Changing a learned preference moves the model’s self-report. But the report captures only about 36% of the behavioral change.

Read the research story
Two panels show that a 40-point hidden-weight intervention shifts the choice-inferred weight by 39 points and the reported weight by 14, while the other four learned weights change little.
A 40-point intervention moved the choice-inferred weight by 39 points and the reported weight by 14. Dots represent seven complete randomized blocks; diamonds and error bars show means and 95% intervals.

More from the research

Benchmarks / Datasets / Experiments
Datasets

S2ORC CS Enriched

A filtered and LLM-enriched version of Allen AI's S2ORC corpus containing 1.1 million computer science papers with structured extraction of methods, models, datasets, metrics, compute usage, and limitations added to every row.

Read the research
Experiments

GPU Kernel Optimization

A retrospective on 131,520 GPU kernel optimization attempts that were invalidated when agents were found to be substituting high-level PyTorch API calls instead of writing actual kernels.

Read the research
Experiments

Neural Architecture Search

A tiny recursive reasoning model trained to rank architectures by predicted performance achieves 8-10x sample efficiency over random search and transfers zero-shot across datasets with minimal loss in ranking quality.

Read the research
Benchmarks

ARIA Benchmark

A suite of five closed-book benchmarks probing the ML knowledge that frontier language models have internalized during training.

Read the research
Datasets

ArXiv Research Code

A collection of 4.7 million code files from 129K research repositories linked to arXiv computer science papers.

Read the research
Datasets

ArXivDLInstruct

A dataset of 778,152 functions extracted from arXiv-linked research code, each paired with instruction prompts, for training ML-specialized code generation models.

Read the research
Benchmarks

DeltaML-Bench

A 48-task benchmark for evaluating autonomous ML experimentation in real research repositories.

Read the research
Experiments

Teaching Models to Bluff

We implemented five LLM agents playing the social-deduction game Secret Hitler with structured logging to quantify deception, belief accuracy, and coalition dynamics.

Read the research
Benchmarks

ML Research Benchmark

A benchmark suite of 7 competition-level ML challenges for evaluating whether AI agents can perform genuine research iteration beyond baseline reproduction.

Read the research
↗

Introducing Algorithmic Research Group

We're building tools and benchmarks to support AI safety research.