• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    PhD Paper on IPA and Romanization Accepted to EMNLP2026

    Sachin Kumar announces Miuzhi's first PhD paper on IPA and romanization for cross-lingual transfer.

    TC
    SK
    2 Sources, 28d ago, first seen 28d ago

    TLDR

    Sachin Kumar, an assistant professor at Ohio State, posted that Miuzhi's first PhD paper was accepted to EMNLP2026. The work examines how alternative text representations such as IPA and romanization influence cross-lingual transfer. Earlier studies on IPA-based language models reported mixed outcomes, with performance sometimes matching or falling slightly behind text-based models. Most existing research on romanization receives limited coverage in the announcement. Tuhin Chakrabarty, an assistant professor at Stony Brook, reposted the update. The post presents the acceptance as confirmed and highlights the paper as Miuzhi's initial doctoral contribution in this area.

    Combined views

    2.5K

    2 Sources, first seen 28d ago

    Combined views

    2.5K

    2 Sources, first seen 28d ago

    37 likes
    37 likes
    3 comments
    11 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    11 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @shocheenExcited to share @miuzhi123's first PhD paper, accepted to #EMNLP2026! We study how alternative representations of text, such as IPA and romanization affect cross lingual transfer. Prior work on IPA based LMs has found mixed results, ranging from matching or slightly trailing text based models. And most work on romanization has used it as a post hoc adaptation strategy and reported improvements over using the original script. We conduct head to head comparisons by pretraining autoregressive multilingual LMs from scratch on text, IPA, and romanized text, across three scales and eight languages. (1) IPA often outperforms text. (2) Romanized pretraining performs best overall. (3) Both substantially reduce disparities in token counts across languages and the resulting gaps in compute, latency, cost, and context-window usage (4) Most surprising to me: romanized fine-tuning hurts when the base LM already covers the target script, and only helps when it does not. More details in the thread below 👇 And follow @miuzhi123 for more cool work coming soon!
    @TuhinChakrRT @shocheen: Excited to share @miuzhi123's first PhD paper, accepted to #EMNLP2026! We study how alternative representations of text, such…

    2 Sources

    @shocheenExcited to share @miuzhi123's first PhD paper, accepted to #EMNLP2026! We study how alternative representations of text, such as IPA and romanization affect cross lingual transfer. Prior work on IPA based LMs has found mixed results, ranging from matching or slightly trailing text based models. And most work on romanization has used it as a post hoc adaptation strategy and reported improvements over using the original script. We conduct head to head comparisons by pretraining autoregressive multilingual LMs from scratch on text, IPA, and romanized text, across three scales and eight languages. (1) IPA often outperforms text. (2) Romanized pretraining performs best overall. (3) Both substantially reduce disparities in token counts across languages and the resulting gaps in compute, latency, cost, and context-window usage (4) Most surprising to me: romanized fine-tuning hurts when the base LM already covers the target script, and only helps when it does not. More details in the thread below 👇 And follow @miuzhi123 for more cool work coming soon!
    @TuhinChakrRT @shocheen: Excited to share @miuzhi123's first PhD paper, accepted to #EMNLP2026! We study how alternative representations of text, such…