Baseten releases GLM-5.2-Vision-NVFP4 optimized for Nvidia B300
The release adds multimodal vision capabilities on Hugging Face.
Many users praised Baseten for publishing the GLM-5.2 Vision Model on Hugging Face because it delivers larger vision capabilities and shows local AI models continuing to improve.
No Digg Deeper questions have been answered for this story yet.
Most Activity
They really did it! I'm so happy to see this being published. Now we have GLM-5.2 with vision. Putting those B300s to good use. Thank you Baseten, I am porting this to the hybrid now. https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4
@0xSero How many more sparks do I need to add vision? 🙃
@volatilemarkts all baseten
@0xSero This is impressive. Well done.
@0xSero Have you tried OCRs as tools instead of giving it vision? What is the pros of this?
@0xSero omg finally, also nvfp4
@0xSero That's dope. Nice bigger vision models!
@0xSero Damn..vision was the one thing that was keeping glm in the cage..for my work i have to rely on 2 model only because it doesn't have vision capability
@0xSero What are the cons of this? Do we lose any quality?
@0xSero W Baseten 🔥
@0xSero you can just do that? 🤯
@0xSero Nice work! Will give it a test
@0xSero Congrats Baseten. And well done.
@0xSero That’s dope. But it’s pretty easy to just create work arounds too
https://x.com/part_harry_/status/2077610277571637435 great post I learned a lot from.