Reviewed evidence · reference host: 24 GiB Apple Silicon

Results without the victory lap.

The dashboard preserves coverage, variability, and missing capabilities so a clean chart cannot outrun the evidence behind it.

91observed repeated trialsAcross 3 model profiles and 13 core cases

Coverage before comparison.

Qwen and Gemma have complete three-run coverage. GPT-OSS remains useful but incomplete and is never silently treated as equivalent evidence.

Qwen 3.5 9B

Complete
39 / 39
Median TTFT
0.91s
Median speed
17.7 tok/s
Schema success
100.0%
Peak memory
16.8 GB

Gemma 4 12B

Complete
39 / 39
Median TTFT
0.91s
Median speed
13.5 tok/s
Schema success
40.0%
Peak memory
16.5 GB

GPT-OSS 20B

Incomplete
13 / 39
Median TTFT
0.67s
Median speed
21.7 tok/s
Schema success
100.0%
Peak memory
18.5 GB

A workload map, not a podium.

A five-percentage-point margin is required. N/A means the needed evidence is unavailable, not that a model scored zero.

WorkloadQwenGemmaGPT-OSSRecommendation
coding100.0%80.0%100.0%No meaningful difference
reasoning83.3%100.0%N/AGemma
repository100.0%50.0%N/AQwen
multimodal93.4%84.1%N/AQwen
vision93.4%84.1%N/AQwen
camera91.4%91.1%N/ANo meaningful difference
rag99.9%95.0%94.1%No meaningful difference
sports100.0%0.0%N/AQwen
hallucination100.0%0.0%N/AQwen
long context68.3%41.6%100.0%GPT-OSSnormalized 16K throughput; review quality and grounding alongside speed

Read past the aggregate.

Each report records host conditions, method changes, caveats, and follow-up work needed before results become recommendations.