Reaction
Slicing as an approach to diffusion language model evaluation
A researcher calls the slicing clever but warns that the chosen model's biases remain.
TLDR
A researcher describes a diffusion language model evaluation approach that uses extensive slicing to handle high dimensionality. They consider it more robust than earlier model-based approaches, but say it still reflects the biases of the model used as a critic; using GPT-2 to judge modern LLM outputs is their example. In an August essay, they also argue that evaluations lack standardization and can be biased by surrogate models.
Combined views
2.4K
3 Sources, first seen ago
31 likes2 comments21 saves1 reposts