τ-bench
A benchmark of tool-using agents talking to simulated customers under policy constraints.
τ-bench (Sierra, 2024) places an agent in retail and airline customer-service scenarios where it must call tools, follow domain policies and converse with an LLM-simulated user. It introduced the pass^k metric to measure consistency. At launch the best agent (GPT-4o) solved under 50% of tasks on one attempt and ~25% of retail tasks eight times in a row.