VoiceArena’s separate measures of task completion and naturalness
A post praising VoiceArena’s approach says humans and voice models appear much closer on task completion than on naturalness.
TLDR
VoiceArena evaluates task completion and naturalness separately using blind pairwise voting, according to a post discussing its design. The author praises the distinction for changing what the numbers reveal: humans and models are apparently much closer on task completion, but not on naturalness.
Combined views
6.3K
2 Sources, first seen 16d ago