Qwen 3.5 9B
RecommendedStrongest general local profile across repository, vision, sports, and grounded tasks.
Repeated-trial coverage
39 / 39Apple Silicon · 24 GiB · localhost only
A reproducible MLX lab for comparing one model at a time—complete with process safety, OpenAI-compatible APIs, realistic benchmarks, and evidence you can inspect.
Current lab snapshot
The useful question is not “which model wins?” It is which measured behavior fits the workload—and whether the evidence is complete enough to trust the answer.
Strongest general local profile across repository, vision, sports, and grounded tasks.
Repeated-trial coverage
39 / 39Excellent reasoning results, with weaker structured-schema reliability in this setup.
Reasoning route score
100%Fast and capable in observed text runs, but repeated-trial coverage remains unfinished.
Repeated-trial coverage
13 / 39