All tools › Tools & Infrastructure › Terminal-Bench Terminal-Bench Tools & Infrastructure Terminal agent performance. GPT-5.4 leads at 77.3%. 🌐 Visit website ↗ 🔗 Similar tools OpenAI Codex CLI ★ 121.7k OpenAI Codex CLI is an open-source terminal coding agent written in Rust (Apache-2.0, 82K+… Coding Agents Harbor ★ 5.0k Harbor is an Apache-2.0 framework, developed by Stanford and the Laude Institute, for eval… MAI-Code-1.1-Flash MAI-Code-1.1-Flash is a production coding assistant that delivers higher-quality code than… SWE-bench Pro SWE-bench Pro is a repository-level software engineering benchmark that extends the Verifi… SWE-bench ★ 5.8k SWE-bench is a benchmark that evaluates LLM-based systems on their ability to resolve real… τ²-bench (tau2-bench) v1.0.1 ★ 2.0k τ²-bench v1.0.1 is Sierra Research's benchmark for evaluating tool-agent-user interaction …