Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.
I'm a little confused by why/how this would work and can only speculate that this is co-designed with the Frontier Labs' architecture in mind. For attention, doing local sparsity makes little sense since tokens in the same neighbourhood will, by construction, have similar attention weights For the MLP, doing sparsity across the reduction dimension can only be designed with some form of QAT in mind
This sparsity idea is super interesting and out of left field. For every local group of 4 elements, we keep the 2 highest magnitude ones and store their indices - implicitly zeroing the rest Rubin then performs the MMA directly on the compressed tensor + indices which saves bw
Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.