• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Kamal Gupta on Early Vision Token Injection Benefits

    Kamal Gupta argues for injecting vision tokens early during pre-training.

    1 Source, 24d ago, first seen 24d ago

    TLDR

    Kamal Gupta, previously at Tesla Optimus, posted that with a fixed vision-plus-text token budget, injecting vision tokens earlier in training improves results on both vision and text tasks, citing K2.5 and other works. He added that nearly all LLMs already add vision early in pre-training. The post includes a table titled Table 1: Performance comparison across different vision-text joint-training strategies, which indicates early fusion with a lower vision ratio yields better results.

    Combined views

    —

    1 Source, first seen 24d ago

    Combined views

    —

    1 Source, first seen 24d ago

    — likes
    — likes
    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @kamalgupta09Given a fixed vision+text tokens budget, earlier you inject the vision tokens into the training, the better your end model gets - not just at the vision tasks but also the text tasks (as shown in K2.5 and multiple other works). Nearly all LLMs today use vision early on during pre-training, and it would be silly to define a VLM as a model that uses LLM pretraining for the base model. I think it would be silly to define VLA as something that uses VLM pretraining for a base model. All future LLMs will have action data from early on in the pretraining. And they will be all VLAs.

    1 Source

    @kamalgupta09Given a fixed vision+text tokens budget, earlier you inject the vision tokens into the training, the better your end model gets - not just at the vision tasks but also the text tasks (as shown in K2.5 and multiple other works). Nearly all LLMs today use vision early on during pre-training, and it would be silly to define a VLM as a model that uses LLM pretraining for the base model. I think it would be silly to define VLA as something that uses VLM pretraining for a base model. All future LLMs will have action data from early on in the pretraining. And they will be all VLAs.