Skepticism over rumors of failed R2 training on unproven hardware
One post argues that the team behind the rumored R2 model would have trained something small first, rather than risk a frontier model on unfamiliar hardware when its advantage was mastery of NVIDIA's stack.
TLDR
A post challenges rumors of a failed attempt to train a frontier model called R2 on unproven hardware. The author questions why a team whose advantage was, in their view, unmatched mastery of NVIDIA's stack would take that approach, arguing that it would have trained a small model first.