Apple Silicon · 24 GiB · localhost only

Measure local models where they actually run.

A reproducible MLX lab for comparing one model at a time—complete with process safety, OpenAI-compatible APIs, realistic benchmarks, and evidence you can inspect.

One model at a timeComparable runs without hidden contention
127.0.0.1 by designInference stays on the machine
Raw evidence retainedReports never replace measurements

Different strengths. Same hardware.

The useful question is not “which model wins?” It is which measured behavior fits the workload—and whether the evidence is complete enough to trust the answer.

Qwen 3.5 9B

Recommended

Strongest general local profile across repository, vision, sports, and grounded tasks.

Repeated-trial coverage

39 / 39

Gemma 4 12B

Measured

Excellent reasoning results, with weaker structured-schema reliability in this setup.

Reasoning route score

100%

GPT-OSS 20B

Incomplete

Fast and capable in observed text runs, but repeated-trial coverage remains unfinished.

Repeated-trial coverage

13 / 39