ARIA Benchmark: How Much Machine Learning Do AI Models Actually Know?
A suite of five closed-book benchmarks probing the ML knowledge that frontier language models have internalized during training.
A suite of five closed-book benchmarks probing the ML knowledge that frontier language models have internalized during training.
A collection of 4.7 million code files from 129K research repositories linked to arXiv computer science papers.
A dataset of 778,152 functions extracted from arXiv-linked research code, each paired with instruction prompts, for training ML-specialized code generation models.
A 48-task benchmark for evaluating autonomous ML experimentation in real research repositories.
A benchmark suite of 7 competition-level ML challenges for evaluating whether AI agents can perform genuine research iteration beyond baseline reproduction.