WorkspaceBench introduced to evaluate AI interpretability tools
Its creators describe WorkspaceBench as a test of how well interpretability tools reveal important intermediate variables during a model’s forward pass—the computation that produces an output.
TLDR
A WorkspaceBench creator raises concern about Astra’s capabilities without chain-of-thought and argues that interpretability can help. The team’s new evaluation is designed to measure how well interpretability tools surface the contents of the “global workspace,” described as the important intermediate variables in a model’s forward pass.
WorkspaceBench introduced to evaluate AI interpretability tools
Its creators describe WorkspaceBench as a test of how well interpretability tools reveal important intermediate variables during a model’s forward pass—the computation that produces an output.
