Reaction
Could looped models be less susceptible to distillation attacks?
A user wonders whether reasoning in latent space, rather than executable chains of thought, might make a difference.
TLDR
A user asks whether looped models might be less susceptible to distillation attacks because more of their “reasoning” occurs in latent space rather than in executable chains of thought.
Combined views
725
1 Source, first seen ago