Skip to content
ARGSTUDIO
  • Work
  • Research
  • Studio
  • Notes
  • Let’s talk ↗
← All tags

#benchmarks

ARIA Benchmark: How Much Machine Learning Do AI Models Actually Know?

March 01, 2026 research

A suite of five closed-book benchmarks probing the ML knowledge that frontier language models have internalized during training.

ArXiv Research Code Dataset: 129K Research Repositories

March 01, 2026 research

A collection of 4.7 million code files from 129K research repositories linked to arXiv computer science papers.

ArXivDLInstruct: 778K Research Code Functions for Instruction Tuning

March 01, 2026 research

A dataset of 778,152 functions extracted from arXiv-linked research code, each paired with instruction prompts, for training ML-specialized code generation models.

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

March 01, 2026 research

A 48-task benchmark for evaluating autonomous ML experimentation in real research repositories.

ML Research Benchmark: Can AI Agents Do Real ML Research?

January 01, 2025 research

A benchmark suite of 7 competition-level ML challenges for evaluating whether AI agents can perform genuine research iteration beyond baseline reproduction.

Independent thinking.
Things worth making.

Work with the studio ↗GitHub ↗Research ↗Notes ↗
ARG STUDIO✳
© 2026 ARG StudioDesign / Development / ResearchBack to top ↑