mlx-vlm-server¶
Packages
Package Apple MLX's vision-language inference server as a standalone Nix binary, with two build-time source patches that any Apple-Silicon user running local VLMs runs into sooner or later.
Apple-Silicon only — mlx is Metal-backed, so build and run on nix-darwin /
aarch64-darwin.
What it solves¶
mlx-vlm ships python -m mlx_vlm.server, an OpenAI-compatible endpoint for
vision-language models. Two things get in the way of using it as-is:
-
Text-only / image-only VLMs drag in torchvision. Loading a model calls
transformers'AutoProcessor, which eagerly builds anAutoVideoProcessor— and that pulls in a full PyTorch + torchvision closure. Plenty of capable VLMs never touch video, so this is pure closure bloat (and, on some models, an outright load failure). -
Reasoning models emit raw
<think>…</think>. Qwen3-style models stream their chain-of-thought inline incontent. Clients that expect an OpenAIreasoning_contentfield (or that simply should not show the scratchpad to users) get the reasoning dumped straight into the answer. Worse, many community MLX quantizations strip thechat_templatefromtokenizer_config.json, so the model never gets its<think>prompt at all.
The key insight / trap¶
Both fixes are build-time source patches on the packaged Python, not runtime config — the upstream server has no hook for either.
-
Video bypass (
postPatch,substituteInPlace mlx_vlm/utils.py): wrap the singleAutoProcessor.from_pretrainedcall so that, for the duration of that call only,AutoVideoProcessor.from_pretrainedreturnsNoneand the class check that would reject aNonevideo processor is relaxed. Both originals are restored immediately after. The processor loads torchvision-free; video is simply absent. The trap is that this is a targeted, self-reverting monkeypatch around one call — patch too broadly and you break real video models; patch too narrowly (or forget to restore) and you corrupt later processor loads. -
Reasoning split + chat-template recovery (
thinking-patch.py): a streaming state machine (ThinkingTagParser) that separates<think>…</think>intoreasoning_contenton both the streaming and non-streaming code paths, plus an_ensure_chat_templatehelper that re-hydrates a missingchat_template— first from a siblingtokenizer_config.json(offline), then from the base (non-quantized) model repo on Hugging Face. The parser buffers while undecided (reasoning is hidden anyway, so the latency is invisible) and handles both "model opens<think>" and "chat template already opened<think>, model only emits the closing tag" cases.
Because these are postPatch steps, the patched behavior is baked into the Nix
store path and reproduces exactly — no per-run flags, no drift.
-
Only patch what you actually patch.
mlx-vlm's text-only dependency,mlx-lm, needs no patching, so this recipe does not build it — it takespython3Packages.mlx-lmfrom nixpkgs. An earlier version re-vendored it from the PyPI sdist with a hand-copied dependency list anddoCheck = false, which bought nothing and cost a permanent version lag plus a manual hash bump per release. If you need a differentmlx-lm, override the nixpkgs derivation and pass it in; keep the vendoring for the package whose source you genuinely modify. -
sentencepieceis not a runtime dep in nixpkgs. Upstreammlx-lmdeclarestransformers[sentencepiece], but nixpkgs carriessentencepieceas a check input only. Models whose tokenizer is a SentencePiecetokenizer.modelwith no fasttokenizer.jsonfail to load without it, so the server env adds it back viaextraPythonPackages.
Usage¶
Import the file and build one of the three attributes it exposes:
let
mlx = import ./packages/mlx-vlm-server { inherit pkgs; };
in
mlx.mlx-vlm-server # the server binary
Or straight from the CLI:
nix-build -A mlx-vlm-server # OpenAI-compatible server binary
nix-build -A mlx-vlm # the patched Python package
nix-build -A mlx-lm # the text-only dependency (nixpkgs', re-exported)
Need a different mlx-lm than your nixpkgs ships? Override rather than
re-vendor:
import ./packages/mlx-vlm-server {
inherit pkgs;
mlx-lm = pkgs.python3Packages.mlx-lm.overridePythonAttrs (old: rec {
version = "0.31.4";
src = old.src.override {
tag = "v${version}";
hash = "sha256-...";
};
});
}
Run it — all args pass through to python -m mlx_vlm.server:
Then hit it as an OpenAI chat endpoint. Reasoning models return their
chain-of-thought in reasoning_content, with the user-facing answer in
content.
Serving it as a daemon¶
The binary is intentionally plain (a writeShellScriptBin wrapper, no baked-in
host/port), so wiring it into a service is trivial. Point your model cache at the
build via the standard Hugging Face env vars and launch one instance per model:
Caveats¶
- Version bumps are hash-paired.
mlx-vlmpinsversion+hashtogether (fetchPypi). Bump both; grab the new hash from the failing build'sgot:line ornix-prefetch-url --unpack.mlx-lmneeds none of this — it moves when your nixpkgs moves. substituteInPlace … --replace-failis a canary. If a futuremlx-vlmrelease rewrites theAutoProcessor.from_pretrained(...)call site, the build fails loudly instead of silently no-op'ing the patch. Same for the string anchors inthinking-patch.py, whichasserton every match — treat a build failure there as "upstream moved the code," and re-anchor the patch.- opencv: nixpkgs provides
opencv4, not the PyPIopencv-pythonwheel, so the wheel dependency is dropped viapythonRemoveDeps. - Metal at runtime. These build without a GPU but need Apple Silicon /
Metal to actually serve;
doCheck = falsebecause upstream tests want network and Metal.
Source¶
packages/mlx-vlm-server/default.nix
# mlx-vlm-server — a standalone Nix build of Apple MLX's vision-language
# inference server, with two build-time source patches that make text-only
# VLMs and reasoning models actually usable.
#
# What you get (import this file, then build one of the attrs it returns):
#
# nix-build -A mlx-vlm-server # OpenAI-compatible server binary
# nix-build -A mlx-vlm # the Python package (patched)
# nix-build -A mlx-lm # the text-only dependency
#
# The reusable value is `mlx-vlm` below: its `postPatch` (a) neuters
# torchvision video processing so text-only / image-only VLMs load without a
# heavyweight PyTorch stack, and (b) runs ./thinking-patch.py to split
# <think>…</think> reasoning into a separate `reasoning_content` field and to
# recover a chat_template that community quantizations strip.
#
# The text-only dependency, mlx-lm, is NOT re-packaged here: nixpkgs ships
# `python3Packages.mlx-lm`, tracks upstream, and runs the sandbox-safe half of
# its test suite. Pass `mlx-lm` if you need a different build — override the
# nixpkgs derivation rather than re-vendoring a fetchPypi copy.
#
# mlx-vlm itself IS built here, because the two source patches are the recipe.
# It is pinned by hash and builds offline once fetched — Apple Silicon only
# (mlx is Metal-backed), so build on nix-darwin / an aarch64-darwin box.
#
# Usage from a flake:
# let mlx = import ./packages/mlx-vlm-server { inherit pkgs; };
# in mlx.mlx-vlm-server
#
# Then run: mlx-vlm-server --model <hf-repo-or-path> --host 127.0.0.1 --port 8080
{
pkgs ? import <nixpkgs> { },
# Text-only MLX inference; mlx-vlm builds on top of it.
mlx-lm ? pkgs.python3Packages.mlx-lm,
# nixpkgs keeps sentencepiece as a check input only, while upstream mlx-lm
# declares transformers[sentencepiece]. Without it, models whose tokenizer is
# a SentencePiece `tokenizer.model` with no fast `tokenizer.json` fail to
# load. Set to [ ] for a smaller closure if none of your models need it.
extraPythonPackages ? [ pkgs.python3Packages.sentencepiece ],
}:
let
python3Packages = pkgs.python3Packages;
# ---------------------------------------------------------------------------
# mlx-vlm — vision-language inference + an OpenAI-compatible server.
#
# The two source patches in postPatch are the whole point of this recipe:
#
# 1. Video-processing bypass (the substituteInPlace on utils.py).
# transformers' AutoProcessor eagerly constructs an AutoVideoProcessor,
# which drags in torchvision. Many capable VLMs are text/image-only and
# you do not want a torch+torchvision closure just to load them. The
# patch monkeypatches AutoVideoProcessor.from_pretrained to return None
# for the duration of the AutoProcessor call, and relaxes the class check
# that would otherwise reject the None video processor — then restores
# both originals. Result: the processor loads, video is simply absent.
#
# 2. Reasoning + chat-template fixes (./thinking-patch.py), applied to
# server.py and prompt_utils.py. See that file for the details; in short
# it splits <think>…</think> into `reasoning_content` (streaming and
# non-streaming) and re-hydrates a chat_template that quantized repos drop.
# ---------------------------------------------------------------------------
mlx-vlm = python3Packages.buildPythonPackage rec {
pname = "mlx-vlm";
version = "0.4.2";
pyproject = true;
src = python3Packages.fetchPypi {
pname = "mlx_vlm";
inherit version;
hash = "sha256-MchLQyHI8XzssEV/oY1cBomCCmavGRkm0kDn35dWRT4=";
};
build-system = [ python3Packages.setuptools ];
dependencies = with python3Packages; [
mlx-lm
mlx
numpy
transformers
pillow
requests
fastapi
uvicorn
tqdm
datasets
soundfile
miniaudio
opencv4
];
# nixpkgs ships opencv as `opencv4`, not the PyPI `opencv-python` wheel;
# drop the wheel dep so the metadata check passes.
pythonRemoveDeps = [ "opencv-python" ];
postPatch =
let
# The eager AutoProcessor call, replaced by a torchvision-free version
# that temporarily disables the video processor. The leading indent on
# every line after the first matches the original call site's block.
old = "processor = AutoProcessor.from_pretrained(model_path, use_fast=True, **kwargs)";
new = builtins.concatStringsSep "\n" [
"from transformers.models.auto import video_processing_auto as _vpa"
" from transformers import processing_utils as _pu"
" _orig_vp = _vpa.AutoVideoProcessor.from_pretrained"
" _orig_check = _pu.ProcessorMixin.check_argument_for_proper_class"
" _vpa.AutoVideoProcessor.from_pretrained = classmethod(lambda cls, *a, **kw: None)"
" def _skip_none_check(self, name, arg):"
" if arg is None: return type(None)"
" return _orig_check(self, name, arg)"
" _pu.ProcessorMixin.check_argument_for_proper_class = _skip_none_check"
" processor = AutoProcessor.from_pretrained(model_path, use_fast=True, **kwargs)"
" _vpa.AutoVideoProcessor.from_pretrained = _orig_vp"
" _pu.ProcessorMixin.check_argument_for_proper_class = _orig_check"
];
thinkingPatch = ./thinking-patch.py;
in
''
substituteInPlace mlx_vlm/utils.py \
--replace-fail \
'${old}' \
'${new}'
${python3Packages.python.interpreter} ${thinkingPatch}
'';
doCheck = false;
};
# ---------------------------------------------------------------------------
# mlx-vlm-server — a plain, on-PATH binary. `python -m mlx_vlm.server`
# wrapped so nothing needs to know about the Python environment. All CLI
# args pass straight through to the module (--model, --host, --port, …).
# ---------------------------------------------------------------------------
pythonEnv = pkgs.python3.withPackages (_: [ mlx-vlm ] ++ extraPythonPackages);
mlx-vlm-server = pkgs.writeShellScriptBin "mlx-vlm-server" ''
exec ${pythonEnv}/bin/python -m mlx_vlm.server "$@"
'';
in
{
inherit mlx-lm mlx-vlm mlx-vlm-server;
}