Getting started · about 10 minutes before model download

From clean checkout to a local response.

The setup is intentionally explicit: dependencies first, model weights only when you ask, and every server bound to localhost.

01

Check the machine

You need an Apple Silicon Mac, macOS with Xcode Command Line Tools, Python 3.12, uv, and enough free storage for the selected checkpoint.

make info

02

Prepare the lab

This creates the Python environment and runs lint and tests. It does not download model weights.

git clone https://github.com/SecuritahGuy/mlx-local-lab.git
cd mlx-local-lab
make setup
make check

03

Download, then start

Downloads are explicit so starting a server cannot unexpectedly fetch multiple gigabytes.

make download MODEL=qwen
make model MODEL=qwen
make health
Memory first. Begin with the conservative 2K benchmark configuration, close memory-heavy apps, and stop if swap rises persistently.

04

Query the local API

uv run local-mlx chat "Write a robust Python retry function."

curl http://127.0.0.1:8080/v1/models

The library also exposes a runtime-neutral LocalLLM abstraction. See the Codex provider guide for Responses API compatibility.

05

Stop safely

make stop

The stop command verifies the recorded PID and process creation time before signaling it. It does not kill an unknown process merely because it occupies port 8080.