• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    CROCODIL Paper Tests Cross-Model Code Editing

    UT Austin researchers test how LLMs edit code written by other models.

    EL
    RP
    3 Sources, 25d ago, first seen 25d ago

    TLDR

    Elvis Saravia posted about the arXiv paper CROCODIL by Linghan Zhong, Aditya Thimmaiah, Milos Gligoric and Junyi Jessy Li. The work from UT Austin and Cisco Research measures editing by one model on code originally produced by a different model. Saravia noted that some models over-edit such code and that many repositories now contain commits from multiple models, which changes subsequent edits by each model.

    Combined views

    16.7K

    3 Sources, first seen 25d ago

    Combined views

    16.7K

    3 Sources, first seen 25d ago

    156 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    156 likes
    33 comments
    111 saves
    32 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    33 comments
    111 saves
    32 reposts

    3 Sources

    @omarsar0This is a weird behavior in coding models and something worth looking into. It turns that some models over-edit code that another models wrote. There is a high chance that your repo now has commits from more than one model, and that changes how each of them edits. Researchers measured what happens when one model edits code another model wrote. Different training data produces different stylistic preferences, and models make more edits, often excessive ones, on foreign code than on their own. CROCODIL is a post-training framework that reduces that behavior. A similarity reward penalizes large changes and an execution reward scores build and test success, and the two are multiplied rather than added. That product stops the policy from shrinking edits by simply failing the task. Paper: https://academy.dair.ai/papers/crocodil-cross-model-code-editing-with-llms-2609.03894
    @rohanpaul_aiWhen coding models edit each other’s work, they tend to over-edit, and this paper shows that training for minimal correct diffs works better than stricter prompts. A stricter prompt telling the model to make only minimal edits did not fix this consistently. The researchers instead post-trained Olmo3 7B with 2 signals: keep the edit small, but still build and pass tests. CROCODIL roughly halved its edit distance across implementations from every model, while improving build and all-test pass rates on every foreign implementor tested. The study is limited to Rust function edits. Still, if a team mixes coding models, it should benchmark cross-model editing and measure unnecessary diff size, not just whether the final code passes. – arxiv. org/abs/2609.03894 Title: "CROCODIL: Cross-Model Code Editing with LLMs"

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 Sources

    @omarsar0This is a weird behavior in coding models and something worth looking into. It turns that some models over-edit code that another models wrote. There is a high chance that your repo now has commits from more than one model, and that changes how each of them edits. Researchers measured what happens when one model edits code another model wrote. Different training data produces different stylistic preferences, and models make more edits, often excessive ones, on foreign code than on their own. CROCODIL is a post-training framework that reduces that behavior. A similarity reward penalizes large changes and an execution reward scores build and test success, and the two are multiplied rather than added. That product stops the policy from shrinking edits by simply failing the task. Paper: https://academy.dair.ai/papers/crocodil-cross-model-code-editing-with-llms-2609.03894
    @rohanpaul_aiWhen coding models edit each other’s work, they tend to over-edit, and this paper shows that training for minimal correct diffs works better than stricter prompts. A stricter prompt telling the model to make only minimal edits did not fix this consistently. The researchers instead post-trained Olmo3 7B with 2 signals: keep the edit small, but still build and pass tests. CROCODIL roughly halved its edit distance across implementations from every model, while improving build and all-test pass rates on every foreign implementor tested. The study is limited to Rust function edits. Still, if a team mixes coding models, it should benchmark cross-model editing and measure unnecessary diff size, not just whether the final code passes. – arxiv. org/abs/2609.03894 Title: "CROCODIL: Cross-Model Code Editing with LLMs"