Qwen Team Releases E-Commerce Bench for LLM Agents
Benchmark simulates a full year running multiple online stores to test LLM agents.
Qwen team researchers published the paper E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation. It presents a benchmark that places agents inside a simulated 365-day business operating multiple online stores. Agents must manage negotiations, respond to shocks, and sustain operations across long sequences with evolving conditions. The work targets tasks that require handling dynamic environments and long-range dependencies rather than chaining short actions. The paper is available on arXiv.
Combined views
11.5K
2 posts, first seen 1d ago
