Algorithmic Research Group / Code & Benchmarks
ARIA Benchmark
Five closed-book benchmarks probing the ML knowledge that frontier models have internalized during training, covering dataset recognition, model classification, and metric recall across the ML landscape.