Thinking Machines Lab’s Inkling scores an Elo of 836 on on our agentic knowledge work benchmark AA-Briefcase, ahead of DeepSeek V4 Flash but below leading open weights models including Nemotron 3 Ultra and GLM-5.2 Our new agentic knowledge work benchmark, AA-Briefcase, tests models on realistic tasks across thousands of input files, requiring deliverables such as spreadsheets, presentations, and UI mock-ups. Model performance is measured across three dimensions: binary rubric checks for ground-truth correctness, pairwise grading on analytical quality, and pairwise grading on presentation quality. The AA-Briefcase Elo is a single metric that combines results across all three dimensions Key Takeaways: ➤ Scores 19.3% on the AA-Briefcase rubric, below MiMo-V2.5-Pro (21.4%) but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%). Inkling’s rubric score is worst on tasks which include non standard “Other” file types (i.e., not Excel, PowerPoint, PDF, or Word), despite its native multimodal support ➤ Performs higher on Presentation than Analytical Quality with Elo scores of 863 and 764 respectively. Presentation and Analytical Quality are both measured using separate, independent pairwise checks from model submissions. Graders compare two submissions for the same task and pick the one that is more professionally presented (Presentation) and the one with deeper, better-structured analysis (Analytical Quality) ➤ Uses 52K output tokens on average per AA-Briefcase task and 5M output tokens for the full suite, slightly more than models with similar scores on AA-Briefcase. Inkling uses the most tokens for Excel deliverables, followed by Word, PDF, PowerPoint, and Other types ➤ Has one of the highest mean turns per AA-Briefcase task (81) with also one of the widest ranges, resulting in a much lower median (49). Despite one of the higher average turns per task, Inkling comparatively uses fewer tool calls per turn on average (0.5)
Inkling Model Scores Worst On Non-Standard Files In AA-Briefcase Benchmark
Users praise Mimo v2.5 pro as the GOAT for affordable steerable agentic models when reacting to Inkling's 836 Elo AA-Briefcase benchmark score.
No Digg Deeper questions have been answered for this story yet.
Most Activity
Despite one of the higher average turns per task, Inkling comparatively uses fewer tool calls per turn on average (0.5)
Inkling performs higher on Presentation than Analytical Quality with Elo scores of 863 and 764 respectively
Inkling scores 19.3% on the AA-Briefcase rubric, below MiMo-V2.5-Pro (21.4%) but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%)
Inkling’s rubric score is worst on tasks which include non standard “Other” file types (i.e., not Excel, PowerPoint, PDF, or Word), despite its native multimodal support
Inkling uses 52K output tokens on average per AA-Briefcase task and 5M output tokens for the full suite, slightly more than models with similar scores on AA-Briefcase. Inkling uses the most tokens for Excel deliverables, followed by Word, PDF, PowerPoint, and Other types
Inkling’s rubric score is worst on tasks which include non standard “Other” file types (i.e., not Excel, PowerPoint, PDF, or Word), despite its native multimodal support
Inkling has one of the highest mean turns per AA-Briefcase task (81) with also one of the widest ranges, resulting in a much lower median (49)
@ArtificialAnlys Good job lil guy
@ArtificialAnlys presentation > analysis is actually the most relatable benchmark result ive ever seen looks good, says nothing. we are all inkling fr 💀
@ArtificialAnlys I think Inkling's middle benchmark position is actually by design. The model's good enough to start with. The real value is Tinker and making it yours. What do you think companies actually prioritize, raw performance or customization?
@ArtificialAnlys 19% rubric score but it looks really clean doing it
@ArtificialAnlys Mimo v2.5 pro still the GOAT for a steerable, independent agentic model for cheap
@ArtificialAnlys Could you please also benchmark Laguna S 2.1
@ArtificialAnlys 81 turns average but median is 49 meaning some tasks got cooked for 200 rounds model said lemme just keep going and see what happens 😂
Inkling scores 19.3% on the AA-Briefcase rubric, below MiMo-V2.5-Pro (21.4%) but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%)