WetLabs Benchmark tests three models on nine science-lab tasks
One of the benchmark’s creators says the tasks range from easy to hard, with each model given 20 attempts per task.
TLDR
A WetLabs Benchmark creator says the team tested three models on nine tasks, giving each model 20 attempts per task. The aim was to explore whether robots are ready to accelerate science labs. The creator’s summary of the result: “Astra leads but it comes close...”
WetLabs Benchmark tests three models on nine science-lab tasks
One of the benchmark’s creators says the tasks range from easy to hard, with each model given 20 attempts per task.
