RL-XAR is claimed to improve AI writing across three tasks
The thread’s author says RL-XAR trains a model using grading rubrics designed to score expert human writing above model-generated text. They report gains on scientific paper sections, story continuations and Wikipedia pages.
TLDR
The author says RL-XAR learns grading rubrics that favor expert human writing, then uses them to train a model. They report gains across three writing tasks and, in a small human evaluation of papers and stories, 89% and 95% win rates, respectively. The author says the strength of the AI judge and rubric optimizer matters to the method.
Combined views
164K
9 Sources, first seen 6h ago
