Concerns over Anthropic’s thinking-token claims and model reliability
One user highlights a read comparing Anthropic’s thinking-token delivery with its claims; another questions whether AI labs provide the reliable behavior needed to build products.
TLDR
A user recommends a read about the thinking tokens Anthropic actually delivers versus its claims, saying it explains why Fable felt materially worse in August than June. Another user says the paper left them feeling that AI labs had violated trust over reproducible, reliable model behavior, asking how to build a product under those conditions.
Combined views
2.5K
4 Sources, first seen 8h ago
Concerns over Anthropic’s thinking-token claims and model reliability
One user highlights a read comparing Anthropic’s thinking-token delivery with its claims; another questions whether AI labs provide the reliable behavior needed to build products.