Registry-backed · no scattered model IDs

Five checkpoints. One constrained machine.

Every model is an explicit experiment profile with an exact checkpoint, runtime, capability boundary, and conservative memory posture.

01

Qwen 3.5 9B 4-bit

Best first candidate for general, coding, multimodal, and Codex compatibility tests.

mlx-community/Qwen3.5-9B-4bit
Download
5.98 GB
Context cap
16,384
Vision
Yes
Responses
Yes
02

Qwen 3.8 27B 4-bit with MTP

High-upside but tight 24 GiB fit; start at 2K with one sequence and the matching MTP drafter.

mlx-community/Qwen3.8-27B-4bit
Download
15.23 GB
Context cap
2,048
Vision
Yes
Responses
Yes
03

LFM 2.5 VL 3B OptiQ 4-bit

Lightweight screen, UI, grounding, and multi-image specialist; not a general RAG or autonomous tool-use model.

mlx-community/LFM2.5-VL-3B-OptiQ-4bit
Download
2.64 GB
Context cap
8,192
Vision
Yes
Responses
Yes
04

GPT-OSS 20B MXFP4-Q8

Text-only MoE; largest memory footprint here. mlx-lm lacks /v1/responses.

mlx-community/gpt-oss-20b-MXFP4-Q8
Download
12.1 GB
Context cap
16,384
Vision
No
Responses
No
05

Gemma 4 12B IT 4-bit

Multimodal instruction model; evaluate structured output and RAG carefully.

mlx-community/gemma-4-12B-it-4bit
Download
6.77 GB
Context cap
16,384
Vision
Yes
Responses
Yes
Profiles are not extra weights. gptoss-final, gemma-default, and gemma-strict are request profiles or aliases over the same cached checkpoints.