Blog

AI Agent Glossary

τ-bench

A benchmark of tool-using agents talking to simulated customers under policy constraints.

τ-bench (Sierra, 2024) places an agent in retail and airline customer-service scenarios where it must call tools, follow domain policies and converse with an LLM-simulated user. It introduced the pass^k metric to measure consistency. At launch the best agent (GPT-4o) solved under 50% of tasks on one attempt and ~25% of retail tasks eight times in a row.

Source: τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains — Yao et al., Sierra, 2024

See the numbers behind this term: AI Agent Statistics 2026.