The case for letting users read LLMs’ full, unencrypted reasoning traces
A user argues that readable reasoning traces would help ensure large language models behave in an aligned way.
TLDR
In a September 15 post, a user calls for access to LLMs’ full, unencrypted reasoning traces, arguing that letting users read them is a step that could help with alignment immediately.
Combined views
32
1 Source, first seen 15d ago
4 reposts