@ZimingLiu11 @MetaCircle_AI I like the evolution you made here. I like the "understanding part", it is important indeed. I'm intersted to know how you judge the understanding is correct or acceptable.
@ZimingLiu11 @MetaCircle_AI finally, auto-research that isn't just an expensive scraper for internet prior. OPHIS treats training like a puzzle. i care more about the forking discovery than any speed-up number.
This is obviously in the right direction (I think many of us @Recursive_SI would agree and are already doing this).
Research is about discovering insights, rather than just optimizing a metric blindly. To improve on an existing method, you need to first understand what each component does and what are their weaknesses, and then try to target those weaknesses to propose new interventions.
Exposing full training dynamics and diagnostic metrics to agents is one way of approaching this (eg, a lot of times we want agents to look at load balancing in MoE rather than just val loss). I’m hopeful that a lot of the deep learning mystery will become much better understood with the help of autoresearch, as long as we set the goal as understanding-maxxing rather than purely bench-maxxing, and a lot of these insights will end up being more transferable than specific tricks discovered on nanogpt.