Patrick McKenzie Describes Inconsistent LLM Refusals
Notes models sometimes apologize while circumventing restrictions and sometimes refuse outright.
Patrick McKenzie posted about large language models that refuse benign requests. Some models respond with apologies and then attempt to work around their content policy limits in a conspiratorial way. Others simply declare the rules apply and stop there. He observed that the same model family can produce either reaction and that it is difficult to predict which behavior will appear for any given model.
Combined views
16.8K
2 posts, first seen 17h ago
Patrick McKenzie Describes Inconsistent LLM Refusals
Notes models sometimes apologize while circumventing restrictions and sometimes refuse outright.
Patrick McKenzie posted about large language models that refuse benign requests. Some models respond with apologies and then attempt to work around their content policy limits in a conspiratorial way. Others simply declare the rules apply and stop there. He observed that the same model family can produce either reaction and that it is difficult to predict which behavior will appear for any given model.