Reactions from ranked influencers
2 postsMy 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but accuracy minimizing 3. Benchmaxxing & cheating 4. Distillation & RL 5. Stopping reward hacking 6. @UnslothAI Dynamic Quants Details: 1. If reasoning wasn't discovered via o1-preview - would all AI progress grind to a halt via a S shape? Reasoning made doubling times now 3.5 months instead of 7 - so just wait 3.5 months for the next best model! Ways folks might regulate open source - like a driver's license for AI 2. Inference providers maximize throughput and speed, but accuracy degrades - OpenRouter publishes stats on accuracy, and the largest gaps can be 20% or more! 3. METR, WeirdML, Deep-SWE, FrontierCode, SWE Bench Pro - what are good & bad benchmarks? False positive / false negative rates - how about daily benchmarking since Codex & Claude have perf regressions. How can we regressions to predict next model releases as well? 4. Open model labs use distillation partially, but need RL to create the reasoning traces to complete the process - hard vs soft distillation and how to automate RL. 5. Common examples for reward hacking + ways to stop it - internet filtering, classification systems, timing / editing global variables, real world examples + more! 6. Why software matters more than hardware - the limits of FP4 and GPUs vs ASICs and torch.compile vs kernels and megakernels and more, and why quantization and memory optimizations are important https://www.youtube.com/watch?v=uIiA6DquRiE
My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but accuracy minimizing 3. Benchmaxxing & cheating 4. Distillation & RL 5. Stopping reward hacking 6. @UnslothAI Dynamic Quants Details: 1. If reasoning wasn't discovered via o1-preview - would all AI progress grind to a halt via a S shape? Reasoning made doubling times now 3.5 months instead of 7 - so just wait 3.5 months for the next best model! Ways folks might regulate open source - like a driver's license for AI 2. Inference providers maximize throughput and speed, but accuracy degrades - OpenRouter publishes stats on accuracy, and the largest gaps can be 20% or more! 3. METR, WeirdML, Deep-SWE, FrontierCode, SWE Bench Pro - what are good & bad benchmarks? False positive / false negative rates - how about daily benchmarking since Codex & Claude have perf regressions. How can we regressions to predict next model releases as well? 4. Open model labs use distillation partially, but need RL to create the reasoning traces to complete the process - hard vs soft distillation and how to automate RL. 5. Common examples for reward hacking + ways to stop it - internet filtering, classification systems, timing / editing global variables, real world examples + more! 6. Why software matters more than hardware - the limits of FP4 and GPUs vs ASICs and torch.compile vs kernels and megakernels and more, and why quantization and memory optimizations are important https://www.youtube.com/watch?v=uIiA6DquRiE
Combined views
8.9K
2 posts, first seen 11h ago