Maksym Andriushchenko Posts Study on LLM Agent Time Awareness
Maksym Andriushchenko shares ongoing work testing agents on long-horizon benchmarks.
Maksym Andriushchenko posted a new LessWrong blog post asking whether LLM agents are time-aware. The post examines if agents can predict wall-clock time for tasks and estimate time already spent. Andriushchenko described the project as work in progress and said a full paper will follow with expanded related work. Replies from researchers such as Niloofar Mireshghallah and Andreas Kirsch discussed the results and suggested adding timestamps by default. Other posts noted similar failure modes observed in InferenceBench.
💥New blog post: Are LLM agents time-aware? Can they predict wall-clock time of tasks? Can they estimate time they spent? We study this on a range of tasks, including long-horizon ones (ProgramBench, PaperBench, DeepSWE, etc). Joint work with my MATS mentee @MOfengenden!
