Terminal-Bench-Mini-20
Agent benchmark with 20 tasks from Terminal-Bench 2.1: the agent gets a shell in a Docker container and has to do real work – build packages, configure services, hunt down bugs. Only what the task's verifier accepts counts.
Measurement conditions: reasoning effort
medium ·
one attempt per task (pass@1) · a single stream, no parallel requests · speculative decoding (MTP) active ·
60 minutes time limit per task · 163,840 tokens of context ·
agent Terminus-2 via Harbor. The reasoning effort strongly affects token count and run time and is therefore part
of each run's identity.
Tasks side by side
A cell shows whether the quant solved the task, how long the agent took and how fast the model generated tokens. Click to open the details: episodes, requests, time in the model, tokens, peak context, cause of failure – and the agent's full transcript with its reasoning, commands and terminal output.
passednot passedtime limit or abortpassed on a later attempt