• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    A chess-playing robot arm reportedly uses ASCII camera frames for fine alignment

    Its builder says a text-only JEV model uses 100 × 100 character arrays to help align pieces and detect missed grabs.

    AM
    2 Sources, ,

    TLDR

    The builder says a Python process converts camera frames into 100 × 100 character arrays for JEV, a text-only model used for fine alignment and missed-grab detection. JEV does not run the whole system: Opus 5.5 handles high-level orchestration, while a chess engine and a trajectory planner handle other tasks. The builder says the arm has played through about 10 full games.

    Combined views

    76.4K

    2 Sources, first seen 6h ago

    Combined views

    76.4K

    2 Sources, first seen 6h ago

    107 likes
    6h ago
    first seen 6h ago
    107 likes
    10 comments
    23 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    10 comments
    23 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @andrewmccalipDo dumber things. I wired vision into JEV in a very hacky way and it worked great 😂 Last week’s rabbit hole was getting caught up on robot arms. Hardware, training pipelines, MuJoCo, sim2real, all of it. I wanted to understand how the pieces fit together and get enough of a working understanding to start building things. The last time I touched this stuff, my world was ABB RobotStudio, KUKA, and very deterministic machines doing exactly what you told them to do. It was all Faro laser trackers and immaculate localization with no vision in the loop. Coming back to it through vision language action models and learned policies has been wild. There are about six projects I want to build now, but somewhere in the middle of this, JEV FOMO hit. Anyone can make a 10 second latency vision+text API call to Astra. Training VLA models is also getting super easy, I had the arm running on TurboVLA and SmolVLA within 24 hours. That’s sorta table stakes. Can we possibly do it in a dumber way? When I searched around for JEV vision repos, I found a lot of people using other models as vision sidecars, or doing binned action outputs on existing 7B open source models. I wanted to see if I could wrangle extra capability out of the existing JEV model purely through harnessing. Side note, I’m bearish on fine-tuning and bullish on harnessing. Not enough people are single-model-to-rule-them-all-bitter-lesson pilled. So how do you force a low latency text only classifier to understand something useful about a visual scene? Not everything in the image. Just enough to answer a specific question about what’s happening. What information would you have to give it, and how ridiculous could the representation be before it stopped working? Turns out, the old ASCII text trick worked awesome. The goal was photons in, actions out. I ran a Python process to do simple color thresholding (no OpenCV cheating) and convert the camera frames into a 100 × 100 character array. I piped that into JEV and set up a training dojo in MuJoCo to iterate on all the tuning coefficients. A few million test cases later, covering all the annoying combinations (white rook on a white square, surrounded by eight pieces, etc.), I had a fairly robust system. The whole point was to provide fine alignment when pieces were off nominal in their positioning, and detect missed grabs. I should note that JEV is not running the entire pipeline. I’ve got Opus 5.5 as a high level orchestration machine, a stockfish chess engine, a classical kinematics trajectory planner, and a few other housekeeping tools running. Because this is a fairly bounded problem, we can utilize classical programming where it makes sense. The JEV engine gets turned on when we’re roughly positioned in open loop mode. This hybrid open-than-closed loop approach saves a ton of compute and tokens. It’s played through about 10 full games now. I’ve been pitting Opus 5.5 against Astra 6, because StockFish gets a bit boring. The models are worse, which adds a bit of variety to the games. Every night I start a game and let them duel it out. It feels very Westworld, having it whirring in the background while I’m doing research. Highly recommend. Anyways, full 60 second video below. It’s obviously the worst possible solution to the problem, but that was kind of the point. Anyone in SF want to talk VLAs next week?6h

    2 Sources

    @andrewmccalipDo dumber things. I wired vision into JEV in a very hacky way and it worked great 😂 Last week’s rabbit hole was getting caught up on robot arms. Hardware, training pipelines, MuJoCo, sim2real, all of it. I wanted to understand how the pieces fit together and get enough of a working understanding to start building things. The last time I touched this stuff, my world was ABB RobotStudio, KUKA, and very deterministic machines doing exactly what you told them to do. It was all Faro laser trackers and immaculate localization with no vision in the loop. Coming back to it through vision language action models and learned policies has been wild. There are about six projects I want to build now, but somewhere in the middle of this, JEV FOMO hit. Anyone can make a 10 second latency vision+text API call to Astra. Training VLA models is also getting super easy, I had the arm running on TurboVLA and SmolVLA within 24 hours. That’s sorta table stakes. Can we possibly do it in a dumber way? When I searched around for JEV vision repos, I found a lot of people using other models as vision sidecars, or doing binned action outputs on existing 7B open source models. I wanted to see if I could wrangle extra capability out of the existing JEV model purely through harnessing. Side note, I’m bearish on fine-tuning and bullish on harnessing. Not enough people are single-model-to-rule-them-all-bitter-lesson pilled. So how do you force a low latency text only classifier to understand something useful about a visual scene? Not everything in the image. Just enough to answer a specific question about what’s happening. What information would you have to give it, and how ridiculous could the representation be before it stopped working? Turns out, the old ASCII text trick worked awesome. The goal was photons in, actions out. I ran a Python process to do simple color thresholding (no OpenCV cheating) and convert the camera frames into a 100 × 100 character array. I piped that into JEV and set up a training dojo in MuJoCo to iterate on all the tuning coefficients. A few million test cases later, covering all the annoying combinations (white rook on a white square, surrounded by eight pieces, etc.), I had a fairly robust system. The whole point was to provide fine alignment when pieces were off nominal in their positioning, and detect missed grabs. I should note that JEV is not running the entire pipeline. I’ve got Opus 5.5 as a high level orchestration machine, a stockfish chess engine, a classical kinematics trajectory planner, and a few other housekeeping tools running. Because this is a fairly bounded problem, we can utilize classical programming where it makes sense. The JEV engine gets turned on when we’re roughly positioned in open loop mode. This hybrid open-than-closed loop approach saves a ton of compute and tokens. It’s played through about 10 full games now. I’ve been pitting Opus 5.5 against Astra 6, because StockFish gets a bit boring. The models are worse, which adds a bit of variety to the games. Every night I start a game and let them duel it out. It feels very Westworld, having it whirring in the background while I’m doing research. Highly recommend. Anyways, full 60 second video below. It’s obviously the worst possible solution to the problem, but that was kind of the point. Anyone in SF want to talk VLAs next week?6h