Sarvam’s former RL lead is starting something new after 2.5 years
The former reinforcement learning lead argues that scientific discovery needs faster, cheaper ways to verify new ideas if it is to scale with compute.
TLDR
Sarvam’s former reinforcement learning lead says they’re starting something new after 2.5 years at the company. They say their team launched 30B and 105B models in February, and argue that progress in chemistry, materials and biology will require verification that is faster, cheaper and not limited by the throughput of physical labs.
Combined views
1 Source, first seen 4h ago
10 reposts