Algorithmic Research Group / Code & Benchmarks
ML Research Benchmark
A benchmark suite for evaluating AI agents on real machine learning research tasks, with 7 competition-level challenges from NeurIPS, ICML, and CoNLL plus task definitions, a baseline agent, and evaluation infrastructure.