Qwen3.8-Flash-Next on Strix Halo

Terminal-Bench-Mini-20

Agent benchmark with 20 tasks from Terminal-Bench 2.1: the agent gets a shell in a Docker container and has to do real work – build packages, configure services, hunt down bugs. Only what the task's verifier accepts counts.

Measurement conditions: reasoning effort medium · one attempt per task (pass@1) · a single stream, no parallel requests · speculative decoding (MTP) active · 60 minutes time limit per task · 163,840 tokens of context · agent Terminus-2 via Harbor. The reasoning effort strongly affects token count and run time and is therefore part of each run's identity.

Tasks side by side

A cell shows whether the quant solved the task, how long the agent took and how fast the model generated tokens. Click to open the details: episodes, requests, time in the model, tokens, peak context, cause of failure – and the agent's full transcript with its reasoning, commands and terminal output.

passednot passedtime limit or abortpassed on a later attempt

How it was run