Researcher Urges Mech Interp Work on Model Eval Awareness
AI researcher Marco Mascorro replies to a post on open source alignment projects.
TLDR
Marco Mascorro, a roboticist and AI researcher who co-founded Fellow AI and served as a partner at a16z, posted a reply to @willdepue. He wrote that more mechanical interpretability research around model eval awareness seems pretty important, though he was unsure whether the work would fit a nanogpt budget or scale. The reply addresses a call for open source alignment research projects that labs could outsource at low cost.
Combined views
21.2K
3 Sources, first seen 29d ago