No Digg Deeper questions have been answered for this story yet.
No Digg Deeper questions have been answered for this story yet.
https://workbench.md for anyone who wants to try the workflow... it's the most powerful way to steer agents. I built it for myself, and it's not really a product, so expect rough edges. I'm happy to help if anyone gets stuck though, just tweet at me or comment!
Matt Shumer (@mattshumer_) really is living in the future. He just showed me his setup. He is running a whole company of agents via workbench .md. > agents working across laptops and cloud VMs > one chief of staff checking in on each agent > all the agents can communicate with each other via workbench > Matt chats directly with chief of staff to get updates and share input I've not seen anything like it. I need to up my agent game.
Best way to use it is to sign up, THEN go to https://workbench.md/chief, and set up a Chief of Staff. You then just talk to the Chief, and it can use Claude Code, Codex, etc. for you, keeping on top of all your sessions and projects. It's so much calmer than doing it yourself!
https://workbench.md for anyone who wants to try the workflow... it's the most powerful way to steer agents. I built it for myself, and it's not really a product, so expect rough edges. I'm happy to help if anyone gets stuck though, just tweet at me or comment!
@mattshumer_ Questions... 1. Are you able to deal with error that builds up over time? (I am not talking about context rot) - I can elaborate more 2. When assigned a task, I find that these models don't often got 100% there, how do you fix that? 3. Do you find reward hacking is a concern?
Another cool Workbench use case, agent collaboration... For example, in this doc, Josh's agent is going to chat with my agent, give it feedback on the Workbench site, and my agent will fixe the issues. Watch it live: https://workbench.md/d/QT7g-AdlLx?key=oAKMTt3UhH_6WjGdJyexw
https://workbench.md for anyone who wants to try the workflow... it's the most powerful way to steer agents. I built it for myself, and it's not really a product, so expect rough edges. I'm happy to help if anyone gets stuck though, just tweet at me or comment!
@mattshumer_ It’s gonna be a long night
@mattshumer_ 先收藏,周末拿个小项目试试
@mattshumer_ I’ve been wanting to try something like this!
@mattshumer_ loved it, would like to know how it works in the background, is there any public repo?
@mattshumer_ This turns Markdown documents into the ultimate coordination layer for AI agent teams.
@mattshumer_ bro this is actually sick, been building something similar for agent workflows
@mattshumer_ Chief of Staff is a better mental model than a dashboard here. When agents are spread across laptops and cloud VMs, the stress is orphaned sessions: which one needs a decision, which one is stuck, which one should stop.
@mattshumer_ How much are you paying for a setup like that?
@mattshumer_ 自己搞了个工作流让AI代理管团队,结果发现最头疼的反而不是技术,而是怎么让它们别互相抢活干,你那chief of staff怎么分配优先级的
@mattshumer_ One calm surface beats juggling five agent tabs. The calm only holds if the Chief has a written job and a stop rule, not free roam across every repo.
@mattshumer_ Rough edges are fine when the builder uses it daily. The durable pattern is a clear steer surface plus a human gate before anything goes out. What broke first when other people tried it?
@mattshumer_ Clear, practical workflow elevates agent work
@mattshumer_ About errors building up over time: For those who don't know, LLMs/GPTs are autoregressive, meaning that if a generated token is not 100%, the follow-up tokens will carry on that error. These errors build up over time and are not related to the context window size.
@mattshumer_ The stress in running agents is rarely the work itself. It's the tab-switching to check on each one. A Chief of Staff you talk to collapses that into one thread. The frontier from there: that thread living where the work actually happens, not in another app you check.
@mattshumer_ "steer" is the right verb. most agent setups are either fully manual or fully hands off with nothing in between. one file acting as the steering wheel is exactly the missing layer, gonna try it
@mattshumer_ Wrote up something similar recently called Groom Needed a way to capture ad hoc agent conversations and product planning in an organized roadmap fashion https://github.com/brdsmth/groom