Qwen 3.5 9B
Complete- Median TTFT
- 0.91s
- Median speed
- 17.7 tok/s
- Schema success
- 100.0%
- Peak memory
- 16.8 GB
Reviewed evidence · reference host: 24 GiB Apple Silicon
The dashboard preserves coverage, variability, and missing capabilities so a clean chart cannot outrun the evidence behind it.
Repeated-trial summary
Qwen and Gemma have complete three-run coverage. GPT-OSS remains useful but incomplete and is never silently treated as equivalent evidence.
Deterministic routing
A five-percentage-point margin is required. N/A means the needed evidence is unavailable, not that a model scored zero.
| Workload | Qwen | Gemma | GPT-OSS | Recommendation |
|---|---|---|---|---|
| coding | 100.0% | 80.0% | 100.0% | No meaningful difference |
| reasoning | 83.3% | 100.0% | N/A | Gemma |
| repository | 100.0% | 50.0% | N/A | Qwen |
| multimodal | 93.4% | 84.1% | N/A | Qwen |
| vision | 93.4% | 84.1% | N/A | Qwen |
| camera | 91.4% | 91.1% | N/A | No meaningful difference |
| rag | 99.9% | 95.0% | 94.1% | No meaningful difference |
| sports | 100.0% | 0.0% | N/A | Qwen |
| hallucination | 100.0% | 0.0% | N/A | Qwen |
| long context | 68.3% | 41.6% | 100.0% | GPT-OSSnormalized 16K throughput; review quality and grounding alongside speed |
Field notes
Each report records host conditions, method changes, caveats, and follow-up work needed before results become recommendations.
Four August 2026 checkpoints tested for runtime fit, memory behavior, and output quality.
Capability, reliability, and prompt-profile findings for GPT-OSS 20B on the reference host.
Coverage-aware distributions for latency, throughput, schema validity, and routing recommendations.
An investigation of channel markers, strict requests, and structured-output behavior.
The capability skip matrix, request profile, and completed-run status.