Kimi K3 Scores Competitively on WeirdML Benchmark
Kimi K3 (max) performs nearly as well as Opus 4.8 at similar cost on expanded evaluation.
TLDR
Reports position Kimi K3 (max) as a frontier model after results on the WeirdML benchmark. The v2 edition expands coverage to 19 tasks and adds API cost tracking across latest models. This places the system just behind the leading entry while remaining competitive on expense. Commentary from research engineers and AI developers frames the outcome as confirmation of strong baseline capabilities under greedy decoding. The evaluation continues to serve as a practical signal for comparing frontier systems on both accuracy and efficiency.
