Serving bugs may confound GPT-5.6 and GLM comparisons
A user comparing GPT-5.6 with GLM variants attributes “a scary amount” of the variance to model failures exposing hidden serving issues tied to chat templates.
TLDR
A user describes tool-call parser bugs as a bigger confound than they expected. In a follow-up reply, they say they are comparing environments that GPT-5.6 “aces near deterministically” but GLM variants fail. They attribute “a scary amount” of the variance to model failures exposing latent serving issues involving the chat template.
Combined views
5K
2 Sources, first seen 1d ago
Serving bugs may confound GPT-5.6 and GLM comparisons
A user comparing GPT-5.6 with GLM variants attributes “a scary amount” of the variance to model failures exposing hidden serving issues tied to chat templates.
TLDR
A user describes tool-call parser bugs as a bigger confound than they expected. In a follow-up reply, they say they are comparing environments that GPT-5.6 “aces near deterministically” but GLM variants fail. They attribute “a scary amount” of the variance to model failures exposing latent serving issues involving the chat template.