• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    NVIDIA publishes report on its open-source NeMo Data Designer tool

    DAIR AI describes a configuration-based workflow for generating synthetic data: define dataset columns, preview a few records, then adjust the setup before running at full scale.

    elvisEL
    NVIDIA AINA
    DAIR.AIDA
    6 Sources, ,

    TLDR

    DAIR AI highlights NVIDIA’s report on NeMo Data Designer, an open-source synthetic data tool. Its summary describes how a person or agent defines dataset columns in a configuration file, with options including text, code, structured output, images and embeddings. Users preview records before a full-scale run, while the runtime handles column dependencies, model endpoint calls and retries. Citing the paper’s Nemotron case studies, DAIR AI says about 9,000 JSON-schema tasks made with the tool raised Nemotron Nano v3’s score on JSONSchemaBench from 80.2% to 86.9%.

    Combined views

    14.7K

    6 Sources, first seen 20d ago

    Combined views

    14.7K

    6 Sources, first seen 20d ago

    220 likes
    20d ago
    first seen 20d ago
    220 likes
    20 comments
    115 saves
    19 reposts
    20 comments
    115 saves
    19 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    DAIR.AI@dair_aiBanger release from NVIDIA. They just published a report on NeMo Data Designer, their open-source synthetic data tool. It's a declarative config which makes a synthetic data pipeline easy to review, share and rerun. In NDD, a person or an agent defines each dataset column in a config file. Column types include generated text, code, structured output, images, embeddings and statistical samplers that control diversity. Plugins add new types. The workflow is built around previews. You generate a few records, check them, adjust the config and then run at full scale. The runtime handles column dependencies, calls to your model endpoints and retries. The paper includes Nemotron case studies. About 9K JSON-schema tasks made with NDD raised Nemotron Nano v3 from 80.2% to 86.9% on JSONSchemaBench. Paper: https://academy.dair.ai/papers/nemo-data-designer-an-extensible-framework-for-multimodal-synthetic-data-generat-2609.1769920d
    elvis@omarsar0RT @dair_ai: Banger release from NVIDIA. They just published a report on NeMo Data Designer, their open-source synthetic data tool. It's…20d
    Eric W. Tramel@fujikanaedaBuilding NeMo Data Designer was such a crazy ride. Now the team has released the tech report for what we grew it into! Long nights coding back at Gretel pre-acq before we had good agents to do this for us, then open-sourcing after the Nvidia acquisition. Now, the team reports over 25 Trillion tokens have run through Data Designer to date. We couldn't have done that without the trust Nvidia put in us back then to grow the open AI ecosystem. The features we built into Data Designer from the beginning to make it easy for humans to express high-quality synthetic data generation turned into the exact right call for 2026, as agents were able to pick up it up and immediately start building new datasets on their own. Gretel was early to see the value in synthetic data. To think that pre-2025 Gretel actually had to argue that synthetic generation any value at all! Now we see that synthetic data generation is the lifeblood of LLM training across all stages, pre/mid/sft/RL. AI Teaching AI is the present & future. 👏 @johnnypgreco @mvansegb @yevmeyer & team! https://arxiv.org/abs/2609.1769919d
    Cody Blakeney@code_starRT @fujikanaeda: Building NeMo Data Designer was such a crazy ride. Now the team has released the tech report for what we grew it into! L…19d
    xlr8harder@xlr8harderRemember where everyone was sure that synthetic data would ruin models19d
    NVIDIA AI@NVIDIAAI@dair_ai 💚19d

    6 Sources

    DAIR.AI@dair_aiBanger release from NVIDIA. They just published a report on NeMo Data Designer, their open-source synthetic data tool. It's a declarative config which makes a synthetic data pipeline easy to review, share and rerun. In NDD, a person or an agent defines each dataset column in a config file. Column types include generated text, code, structured output, images, embeddings and statistical samplers that control diversity. Plugins add new types. The workflow is built around previews. You generate a few records, check them, adjust the config and then run at full scale. The runtime handles column dependencies, calls to your model endpoints and retries. The paper includes Nemotron case studies. About 9K JSON-schema tasks made with NDD raised Nemotron Nano v3 from 80.2% to 86.9% on JSONSchemaBench. Paper: https://academy.dair.ai/papers/nemo-data-designer-an-extensible-framework-for-multimodal-synthetic-data-generat-2609.1769920d
    elvis@omarsar0RT @dair_ai: Banger release from NVIDIA. They just published a report on NeMo Data Designer, their open-source synthetic data tool. It's…20d
    Eric W. Tramel@fujikanaedaBuilding NeMo Data Designer was such a crazy ride. Now the team has released the tech report for what we grew it into! Long nights coding back at Gretel pre-acq before we had good agents to do this for us, then open-sourcing after the Nvidia acquisition. Now, the team reports over 25 Trillion tokens have run through Data Designer to date. We couldn't have done that without the trust Nvidia put in us back then to grow the open AI ecosystem. The features we built into Data Designer from the beginning to make it easy for humans to express high-quality synthetic data generation turned into the exact right call for 2026, as agents were able to pick up it up and immediately start building new datasets on their own. Gretel was early to see the value in synthetic data. To think that pre-2025 Gretel actually had to argue that synthetic generation any value at all! Now we see that synthetic data generation is the lifeblood of LLM training across all stages, pre/mid/sft/RL. AI Teaching AI is the present & future. 👏 @johnnypgreco @mvansegb @yevmeyer & team! https://arxiv.org/abs/2609.1769919d
    Cody Blakeney@code_starRT @fujikanaeda: Building NeMo Data Designer was such a crazy ride. Now the team has released the tech report for what we grew it into! L…19d
    xlr8harder@xlr8harderRemember where everyone was sure that synthetic data would ruin models19d
    NVIDIA AI@NVIDIAAI@dair_ai 💚19d