How training and watermarking could make AI text easier to detect
A user argues that post-training and synthetic data give AI models increasingly similar writing styles, making their output easier to detect. They also predict wider use of watermarking.
TLDR
A user argues that post-training imposes distinctive writing styles and that growing use of synthetic training data makes different models’ output more alike. They credit Pangram’s detection approach to spotting this recognizable “mainstream” AI style. The post also predicts that regulation such as the AI Act will drive watermarking: statistical patterns embedded during generation that can be detected without a classification model. The author expects watermarking to become an industry standard and identifying AI-generated content to become largely a solved problem.