Rapid-MLX

by raullenchaiVerified

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

3,530
Stars
401
Forks
Python
Language
8/23/2026
Added
View on GitHubDownload ZIP

⚠️ Third-Party Software Notice

This skill is third-party open-source software developed and hosted independently on GitHub. SkillTip is an informational directory and does not control or maintain the underlying repository. Any security checks displayed are automated and limited in scope. Review the source code before installing.

Read the Terms of Service

Installation

Add to your Claude Code skills directory:

# Add to your Claude Code skills
git clone https://github.com/raullenchai/Rapid-MLX

Getting Started

Guides for using skills like Rapid-MLX.

Security Report

Verified

Last scanned: —

{
  "status": "PASSED",
  "issues": []
}

README.md

Rapid-MLX — the fastest local AI engine for Apple Silicon

The fastest local AI engine for Apple Silicon.
Drop-in OpenAI / Anthropic API · up to 3× Ollama's throughput (measured) · Runs on any M-series Mac.

PyPI Homebrew core Python 3.10+ Apple Silicon License

CI GitHub stars Contributors Last commit Join the Rapid-MLX Discord Ask DeepWiki

rapidmlx.com · Docs · Model mirror · Desktop app · Discord


Works with your AI stack

Use Rapid-MLX as a local backend for agents, apps, or your own code. If a client accepts an OpenAI- or Anthropic-compatible endpoint, it can usually use Rapid-MLX without an adapter.

Codex CLI Claude Code OpenCode Qwen Code OpenHands Hermes Agent Aider Kilo Code DeepSeek Harness GitHub Copilot Factory Droid Kimi Code LangChain PydanticAI smolagents Any local-endpoint client

Five Tier-1 agents are exercised end-to-end on real weights before release. See the tested compatibility matrix for exact API coverage and setup status.


Install

Desktop — macOS (Apple Silicon)

The easiest way to chat locally, manage models, and use vision, files, voice, and image generation from one app.

CLI and server — macOS (Apple Silicon)

# Homebrew — prebuilt bottle from homebrew-core
brew install rapid-mlx

# Or the guided installer — detects RAM and recommends a starter model
curl -fsSL https://rapidmlx.com/install.sh | bash

Both install the same rapid-mlx CLI. Prefer uv or pip, or want to verify the installer before running it? See alternative install methods and install security.

The guided installer prints a serve command sized to your Mac (8–15 GB → lfm2.5-2.6b-4bit; 16–17 GB → qwen3.5-4b-4bit; 18–23 GB → qwen3.5-9b-4bit; 24–31 GB → bonsai-27b-2bit; 32 GB+ → qwen3.8-27b-4bit).


Quick Start (60 seconds)

1. Chat with a model right now:

rapid-mlx chat

Defaults to qwen3.5-4b-4bit. First run downloads the weights (~3 GB) with a progress bar and drops you into a REPL. Type /help for slash commands, /exit to quit.

2. Or serve it for use from other apps:

rapid-mlx serve qwen3.5-4b-4bit

Starts an OpenAI-compatible HTTP server bound to http://localhost:8000. Point any client that supports a local custom endpoint (Aider, LangChain, OpenCode, PydanticAI, your own scripts) at http://localhost:8000/v1; Claude Code / Anthropic SDK uses http://localhost:8000 (the Anthropic messages route lives at /v1/messages under the same host).

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"default","messages":[{"role":"user","content":"Say hello"}]}'
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
print(client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "Say hello"}],
).choices[0].message.content)

3. Or wire up your coding agent — one command:

rapid-mlx launch claude-code

With a server running (step 2), this patches Claude Code's local config (~/.claude/settings.json) to route at http://localhost:8000 — no manual env vars, no editing JSON by hand. You get a fully local Claude Code: $0 per token, nothing leaves your Mac. Swap in cline or continue-dev for the other IDE clients, or run rapid-mlx launch list to see what's detected on this machine.

Cursor: Cursor currently routes BYOK requests through its own servers, so its servers cannot reach a Rapid-MLX endpoint on localhost. Rapid-MLX therefore does not generate a Cursor localhost config. If you intentionally expose the server through a public HTTPS tunnel, set RAPID_MLX_API_KEY=your-secret for both rapid-mlx serve ... and rapid-mlx launch cursor --server-url https://your-public-host. This is no longer a fully local connection; never expose an unauthenticated server. Rapid-MLX rejects explicit local/private addresses but cannot verify reachability from Cursor's network, whose DNS view may differ from your Mac.

Vision / audio / video / diffusion models? Base install is text-only (~460 MB). Vision, audio (TTS, STT, voice cloning), video generation, embeddings, and DFlash speculative decoding ship as opt-in extras. → Optional extras

Not into the terminal? Rapid-MLX Desktop bundles the same engine inside a one-click Mac app.


Video generation

Run text-to-video or image-to-video locally through the OpenAI-compatible Videos API. Three backends ship — Wan 2.1 / 2.2, CogVideoX-Fun and LTX-2.3 — across 8 registered checkpoints. wan2.2-ti2v-5b-q8 is the recommended starting point: smallest of the Wan set, and TI2V means one checkpoint does both text-to-video and image-to-video.

Requires Python 3.11+ (the video runtime does not support 3.10; core text and audio still do) and ffmpeg for the final MP4 mux.

pip install 'rapid-mlx[video]'
brew install ffmpeg
rapid-mlx serve wan2.2-ti2v-5b-q8

Create and download a clip:

curl http://localhost:8000/v1/videos \
  -F model=wan2.2-ti2v-5b-q8 \
  -F 'prompt=A fox running through fresh snow, cinematic tracking shot' \
  -F seconds=1 \
  -F size=832x512

# Poll until GET /v1/videos/VIDEO_ID reports "status": "completed", then:
curl http://localhost:8000/v1/videos/VIDEO_ID/content -o output.mp4

The create call returns a job immediately. Poll GET /v1/videos/VIDEO_ID until status is completed. Add -F input_reference=@start.png for image-to-video.

Generation is serialized — one clip at a time — because two diffusion pipelines resident at once will exhaust unified memory. Expect minutes of compute per second of footage, not real time.

Every checkpoint, RAM requirement and tuning knob


Audio: speech, transcription, voice cloning

44 audio aliases behind the OpenAI-compatible /v1/audio/* endpoints — any OpenAI SDK works unchanged.

pip install 'rapid-mlx[audio]'

# Text to speech
rapid-mlx serve kokoro
curl http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"kokoro","input":"hello from rapid-mlx"}' --output hello.wav

# Transcription (Whisper / Parakeet / SenseVoice)
rapid-mlx serve whisper-large-v3-turbo
curl http://localhost:8000/v1/audio/transcriptions \
  -F file=@hello.wav -F model=whisper-large-v3-turbo

Beyond the basics, three things you may not expect to run locally:

  • Zero-shot voice cloning from a reference clip. indextts is the only one that takes the clip alone; qwen3-tts-clone, f5-tts-zh and chatterbox all require ref_text (the clip's exact transcript) paired with ref_audio, and the request is rejected before generation if it is missing.
  • Voice designqwen3-tts-voicedesign has no named speakers at all. Describe the voice you want in natural language via instructions (timbre, gender, age, accent, emotion, prosody) and it synthesises it.
  • Forced alignmentqwen3-aligner takes audio plus the transcript you already have and returns per-character timings. It never guesses at the words, so it cannot mis-hear them; that is what karaoke captions and beat-synced editing need.

Also: word-level timestamps on transcription, and local text-to-music at /v1/audio/music.

All 44 aliases across 13 families


Why Rapid-MLX

Apple-Silicon-nativePure MLX kernels — no llama.cpp fallback, no Metal shim. Continuous batching, prompt cache (radix + DeltaNet RNN snapshots), and a quantized live KV cache (int4/int8 on the continuous-batching cache + TurboQuant K8V4 codec) run at native MLX bandwidth on M1 → M4.
Drop-in OpenAI / Anthropic API/v1/chat/completions, /v1/responses (Codex CLI), /v1/messages (Anthropic SDK / Claude Code), /v1/embeddings, /v1/audio/*, /v1/videos — same wire as ChatGPT / Claude, no client adapter.
First-class ecosystem coverage12 agent CLIs and 3 Python frameworks are wire-verified against real weights every release (5 are Tier-1, re-verified on current binaries) — Codex CLI, Claude Code, OpenCode, Qwen Code, OpenHands, Hermes Agent, Aider, Kilo Code, DeepSeek Harness, GitHub Copilot, Factory Droid, Moonshot Kimi Code + LangChain, PydanticAI, smolagents.

Full feature breakdown


Use Cases

Chat in the terminalrapid-mlx chat qwen3.5-9b-4bitStreaming REPL, /help for slash commands, --think / --no-think to control CoT.
OpenAI server for your appsrapid-mlx serve qwen3.5-9b-4bitPoint Aider, LibreChat, Open WebUI, or LangChain at http://localhost:8000/v1.
Agent backendsrapid-mlx serve qwen3.6-35b-8bit &
rapid-mlx agents codex --setup && codex
10 agents auto-configure via agents <name> --setup once the server is up (12 wire-verified total, 5 Tier-1) — see Agent support.
Benchmark your Macrapid-mlx bench qwen3.5-9b-4bit --submitStandardized B=1 bench, opens a PR to publish your row on rapidmlx.com.

One-shot IDE setup with rapid-mlx launch <claude-code|cline|continue-dev>


Agent Support

All 12 agents below are wire-verified against real weights every release via their own integration-test cell. Of these, five are Tier-1Claude Code, Codex CLI, Hermes, Aider, and DeepSeek Harness — re-verified end-to-end against the current client binary every release, with one guardian per API wire (Anthropic /v1/messages, OpenAI /v1/responses, and /v1/chat/completions covered for tool-calling depth, reach, and DeepSeek's own harness protocol). The other seven are Tier-2: wire-verified in the matrix and configured on-demand. The first nine agents each ship a rapid-mlx agents <name> --setup config template (except Claude Code, which is one env-var), and Continue.dev gets the same one-command setup via rapid-mlx agents continue --setup (ten setup-capable clients in all, though Continue.dev is not part of the wire-verified matrix below); GitHub Copilot, Factory Droid, and Moonshot Kimi Code plug in through their own documented BYOK config (auth-gated, so the matrix cell is a wire smoke).

Tier-1 is not a label — it is a job that blocks the release. tests/integrations/agent_smoke.sh drives each of the five through the same real multi-step bug-fix task against a local 35B model and asserts the repo's own test suite goes green afterwards; if any one of them fails, the version cannot tag or publish.

Tier-1 (5): Claude Code · Codex CLI · Hermes · Aider — last re-verified end-to-end 2026-07-28 on current binaries (claude 2.1.211, codex 0.145.0, hermes 0.9.0, aider 0.86.2). DeepSeek Harness — promoted 2026-08-17, verified on dsh 0.1.0-rc.7 against qwen3.6-35b-8bit. Tier-2 (7): OpenCode · Qwen Code · OpenHands · Kilo Code · GitHub Copilot · Factory Droid · Moonshot Kimi Code.

Also compatible with OpenAI-compatible clients that allow direct local endpoints via http://localhost:8000/v1 — LibreChat, Open WebUI, and more plug in with a single URL change.

Full 12-agent + 3-framework matrix (test cells + xfail reasons)Codex CLI · Claude Code · OpenCode · Qwen Code · OpenHands · Hermes · Aider · Kilo Code · DeepSeek Harness · Copilot · Droid · Kimi Code


Choose Your Model

The installer and desktop app use the same RAM-tier recommendation catalog. Run rapid-mlx recipe to see its Smart and Fast picks for this Mac (--max-ram 32 simulates another tier; --json is machine-readable). If you want to shop the full catalog: rapid-mlx models lists every alias, rapid-mlx info <alias> shows the per-alias profile (parser, MoE / hybrid flags, KV codec eligibility, speculative-decoding gates).

This table is the same one the desktop app's picker reads, and the installer prints the matching line for your Mac — a CI test parses both files and fails if they drift apart. Measured rows use the standard ~8K prompt peak of the complete rapid-mlx serve process tree on an M2 Pro 32 GB Mac mini (the 32 GB+ row: M3 Ultra, 2026-08-18 — footprint is config-bound, speed reads lower on smaller chips).

RAMRecommendedPeak RSSOne-shot
8–15 GB MacBook Air / base Minilfm2.5-2.6b-4bit3.0 GBrapid-mlx serve lfm2.5-2.6b-4bit
16–17 GB MacBook Air / Proqwen3.5-4b-4bit6.0 GBrapid-mlx serve qwen3.5-4b-4bit
18–23 GB MacBook Proqwen3.5-9b-4bit8.7 GBrapid-mlx serve qwen3.5-9b-4bit
24–31 GB Mac Mini / MacBook Probonsai-27b-2bit13.0 GBrapid-mlx serve bonsai-27b-2bit
32 GB+ Mac Studio / MacBook Proqwen3.8-27b-4bit20.0 GBrapid-mlx serve qwen3.8-27b-4bit

Every Mac from 32 GB up gets the same pick, and that is the point: Qwen3.8-27B scores 52 on the Artificial Analysis Intelligence Index (2026-08-18) — GPT-5.6-class, the highest of any open-weights model we serve, ahead of the much larger 122B (33) and 35B (32) it replaces (the index scores the full-precision release; our 4-bit build's deltas are unmeasured — the standing caveat for every quantized pick here). Multi-token prediction is on by default (~40 tok/s decode, 8K prefill at ~324 tok/s, zero swap at every tier budget).

Full RAM tier map + serve flags per tierEvery alias, quant, and family (170 text + 2 text-diffusion + 2 image + 8 video + 44 audio aliases, 226 total) · interactive at models.rapidmlx.com


Alternative install methods

The two paths above cover most users — reach for these only if you already manage Python yourself.

Homebrew — Mac-native, one command, prebuilt bottle from homebrew/core
brew install rapid-mlx

Ships in homebrew-core since 0.10.12 — no tap, no trust prompt. Upgrade with brew upgrade rapid-mlx. If you previously installed from the legacy raullenchai/rapid-mlx tap, switch once: brew uninstall rapid-mlx && brew untap raullenchai/rapid-mlx && brew install rapid-mlx.

uv — isolated tool install, auto-manages Python
uv tool install rapid-mlx@latest

Don't have uv yet? curl -LsSf https://astral.sh/uv/install.sh | sh. Upgrade with uv tool upgrade rapid-mlx.

pip — requires Python 3.10+ (macOS ships 3.9)
python3.12 -m pip install rapid-mlx

If pip install rapid-mlx says "no matching distribution", your Python is too old. brew install python@3.12 first. Upgrade with pip install -U rapid-mlx.

For image-input / VLM models (Qwen-VL, true multimodal), install the vision extra: pip install 'rapid-mlx[vision]' — see Optional extras.

For the complete feature set — vision, chat, embeddings, and audio — install the [all] extra: pip install 'rapid-mlx[all]'. Audio alone is pip install 'rapid-mlx[audio]'; see Optional extras.


Command Reference

rapid-mlx --help                    # top-level command list
rapid-mlx <subcommand> --help       # per-subcommand flags

Covers chat, serve, share, agents (setup / test), bench, recipe, models, ls, pull, rm, alias, ps, info, connect, doctor, upgrade, telemetry, and launch.

Full CLI reference with every flag


Troubleshooting

Run the built-in self-check first:

rapid-mlx doctor

Top three things that go wrong:

  • Much slower than expected. Qwen3.5 / 3.6 default to thinking-on — add --no-think to skip chain-of-thought. → Slow tok/s
  • Out of memory. Model too big for your RAM — pick a smaller quant from Choose Your Model or the full tier map. → OOM guide
  • Tool calls arriving as plain text. Auto-recovery handles most cases; if not, set --tool-call-parser explicitly for your model. → Tool-call recovery

All troubleshooting entries (OOM, empty responses, slow TTFT, port taken, shell completion, HF cache, and more)


See it in action

Rapid-MLX demo — install, serve Gemma 4, chat, tool calling

Community & Support

Privacy: Anonymous telemetry is off by default and requires an explicit rapid-mlx telemetry enable. Prompts, completions, paths, IP addresses, and API keys are never collected. See what we do and don't collect.


Contributors

Every avatar here shipped something in rapid-mlx — model support, tool-call parsers, fixes, docs, and benchmark submissions. Thank you.

rapid-mlx contributors

Star History

Rapid-MLX GitHub star history through August 23, 2026

Acknowledgements

Rapid-MLX began as vLLM-MLX by Wayner Barrios, which is where this repository's history starts and where the engine's paged KV cache, prefix cache, and continuous batching were first built. It was renamed to Rapid-MLX in March 2026 and has been heavily modified since. Thank you.

It stands on Apple's MLX stack and the runtimes built around it:

  • MLX — Apple's array framework for Apple Silicon
  • mlx-lm — LLM inference, KV cache, quantization
  • mlx-vlm — vision-language models
  • mlx-audio — speech and audio models

Vendored third-party components and their licenses are listed in NOTICE; what the macOS app ships is enumerated in apps/rapid-mac/THIRD_PARTY.md.

License

Apache 2.0 — see LICENSE and NOTICE.

Frequently Asked Questions

What is Rapid-MLX?

Rapid-MLX is an open-source testing skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by raullenchai. The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider. It has 3,530 GitHub stars.

Is Rapid-MLX safe to use?

Yes. Rapid-MLX passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.

How do I install Rapid-MLX?

Clone the repository with "git clone https://github.com/raullenchai/Rapid-MLX" and add it to your Claude Code skills directory (see the Installation section above).

What programming language is Rapid-MLX written in?

Rapid-MLX is primarily written in Python. It is open-source under raullenchai on GitHub, so you can review or fork the full source.

Are there alternatives to Rapid-MLX?

Yes. SkillsLLM lists many other Testing skills you can browse and compare side by side. Open the Testing category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh Rapid-MLX against similar tools.

Comments (0)

No comments yet. Be the first to share your thoughts!

Claude-BugHunter

by elementalsouls

A Claude Code skill bundle for bug hunting and external red-team work - 82 skills, 15 slash commands, 681 disclosed-report patterns curated across 24 core vulnerability classes, plus enterprise identity + infrastructure attack matrices.

3,740578Python
Testing
View details

playwright-skill

by lackeyjb

Claude Code Skill for browser automation with Playwright. Model-invoked - Claude autonomously writes and executes custom automation for testing and validation.

3,014229JavaScript
Testing
View details

src-hunter-skill

by MyuriKanao

实战 SRC / 众测 / Bug bounty 漏洞挖掘 Claude Code skill — 19 个攻击类 playbook、305 个结构化 payload、263 个 WAF/EDR 绕过、2887 份 HackerOne 真实案例、88,636 WooYun 案例统计

60187
Testing
View details

100 field-tested Claude Code recipes for knowledge workers — prompts, steps, and 6 installable graded skills.

38048
Testing
View details

Claude Code Skill that turns any idea into a cinematic, model-ready video prompt — Sora · Kling · Veo · Seedance. 21 genre templates, 5-stage structure, eval-tested. Distilled from the AI short Hollywood director PJ Ace called "one of the best short films I've seen in years."

36869Python
Testing
View details

offensive-claude

by hypnguyen1209

Offensive security toolkit for Claude Code covering red team, exploit dev, AD attacks, EDR bypass, mobile pentest

34359Python
Testing
View details

Developers Also Liked

Based on votes and bookmarks from developers who liked this skill

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI Agentsai-agentsanthropicclaude-code
View details
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI Agentsai-agentsbrainstorming
View details

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP Serversapisai-tools
View details

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI Agentsai-agentsanthropicclaude-code
View details

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI Agentsclaude-codeai-tools
View details