Problem
Reproducing and improving on a published paper is slow, manual research work.
Approach
Researchers upload a published paper and define the metric they want to improve. Multiple agents analyse the paper, search GitHub and Hugging Face, test different ML approaches, and loop. Aria (Weights & Biases) traces every agent step so the team could see why a run got better.
Results
In one experiment accuracy went from 18% to 56%. Best Product and Best Use of Aria at CoreWeave Hacks: Agent Loops (W&B, TypeSafe AI, AGI House), Sep 2026. Built with Ali Amjad, Rikin Shah and Ahmad Mustafa.