Hilbert Space View Turns Neuron Pruning Into Geometric Projection
Reactions from ranked influencers
4 posts8/13 Once we are in this Hilbert space, compression becomes elegant geometry. Pruning a neuron means projecting a rank-1 operator to zero. Merging two neurons means projecting a rank-2 operator (the pair) into an optimal rank-1 parent.
9/13 HOPE progressively shrinks the network by picking the action with the lowest projection error. We even scaled this up to pruning a block of multiple layers (residual blocks) using the same geometric principles. But this goes beyond compression! We can use this for Continual and Transfer Learning.
10/13 For that, we partition the network using the Hilbert norm. Neurons with a relatively large norm are frozen to protect foundational knowledge (the "Core"), and the remainder are treated as a highly plastic "Slack".
13/13 It has been a long journey! HOPE opens up exciting open problems for the community, like extending this math to broader architectures or dynamically growing and shrinking networks during training. 🌱 Big thanks to my collaborator and mentor Peter L. Bartlett. Grateful to @sfrei_ @_vaishnavh @junokim_ai @miouantoinette @mc_mozer @brunorganised , Gil Shamir, and Alan Malek, for the amazing chats!
Combined views
323
4 posts, first seen 2h ago