01
qwen · mlx-vlm
Qwen 3.5 9B 4-bit
Best first candidate for general, coding, multimodal, and Codex compatibility tests.
mlx-community/Qwen3.5-9B-4bit- Download
- 5.98 GB
- Context cap
- 16,384
- Vision
- Yes
- Responses
- Yes
Registry-backed · no scattered model IDs
Every model is an explicit experiment profile with an exact checkpoint, runtime, capability boundary, and conservative memory posture.
qwen · mlx-vlm
Best first candidate for general, coding, multimodal, and Codex compatibility tests.
mlx-community/Qwen3.5-9B-4bitqwen38 · mlx-vlm
High-upside but tight 24 GiB fit; start at 2K with one sequence and the matching MTP drafter.
mlx-community/Qwen3.8-27B-4bitlfm25 · mlx-optiq
Lightweight screen, UI, grounding, and multi-image specialist; not a general RAG or autonomous tool-use model.
mlx-community/LFM2.5-VL-3B-OptiQ-4bitgptoss · mlx-lm
Text-only MoE; largest memory footprint here. mlx-lm lacks /v1/responses.
mlx-community/gpt-oss-20b-MXFP4-Q8gemma · mlx-vlm
Multimodal instruction model; evaluate structured output and RAG carefully.
mlx-community/gemma-4-12B-it-4bitgptoss-final, gemma-default, and gemma-strict are request profiles or aliases over the same cached checkpoints.