lemonade

by lemonade-sdkVerified

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

5,441
Stars
473
Forks
C++
Language
8/23/2026
Added
View on GitHubDownload ZIP

⚠️ Third-Party Software Notice

This skill is third-party open-source software developed and hosted independently on GitHub. SkillTip is an informational directory and does not control or maintain the underlying repository. Any security checks displayed are automated and limited in scope. Review the source code before installing.

Read the Terms of Service

Installation

Add to your Claude Code skills directory:

# Add to your Claude Code skills
git clone https://github.com/lemonade-sdk/lemonade

Getting Started

Guides for using skills like lemonade.

Security Report

Verified

Last scanned: —

{
  "status": "PASSED",
  "issues": []
}

README.md

🍋 Lemonade: Refreshingly fast local AI

Discord PRs Welcome Latest Release GitHub downloads GitHub issues License: Apache Star History Chart

Lemonade Banner

Download | Documentation | Discord

Lemonade is the local AI server that gives you the same capabilities as cloud APIs, except 100% free and private. Use the latest models for chat, coding, speech, and image generation on your own NPU and GPU.

Lemonade comes in two flavors:

  • Lemonade Server installs a service you can connect to hundreds of great apps using standard OpenAI, Anthropic, and Ollama APIs.
  • Embeddable Lemonade is a portable binary you can package into your own application to give it multi-modal local AI that auto-optimizes for your user’s PC.

This project is built by the community for every PC, with optimizations by AMD engineers to get the most from Ryzen AI, Radeon, and Strix Halo PCs.

Getting Started

  1. Install: Windows · Linux · macOS · Docker · Source
  2. Get Models: Browse and download with the Model Manager
  3. Generate: Try models with the built-in interfaces for chat, image gen, speech gen, and more
  4. Mobile: Take your lemonade to go: iOS · Android · Source
  5. Connect: Use Lemonade with your favorite apps:

Claude Code  Firefox Chatbot  AnythingLLM  Dify  GAIA  GitHub Copilot  Infinity Arcade  n8n  Open WebUI  OpenHands

Want your app featured here? Just submit a marketplace PR!

Supported Platforms

PlatformBuild
Arch LinuxBuild on Arch
Debian Trixie+Build on Debian
DockerBuild Container Image
Fedora 43+Build .rpm
macOSBuild .pkg
SnapBuild Snap
Ubuntu 24.04+Build Launchpad PPA
Windows 11Build .msi

Using the CLI

To run and chat with Gemma:

lemonade run Gemma-4-E2B-it-GGUF

To code with Lemonade models:

lemonade launch claude

Multi-modality:

# image gen
lemonade run SDXL-Turbo

# speech gen
lemonade run kokoro-v1

# transcription
lemonade run Whisper-Large-v3-Turbo

To see available models and download them:

lemonade list

lemonade pull Gemma-4-E2B-it-GGUF

To manage model aliases for environment-independent naming and active-standby failover:

lemonade alias add production-llm Gemma-4-E2B-it-GGUF
lemonade alias list

# Instant active-standby failover to a different model target
lemonade alias add production-llm Qwen3-0.6B-GGUF
lemonade alias remove production-llm

To see the backends available on your PC:

lemonade backends

For hybrid setups, Lemonade can also route to any OpenAI-compatible cloud provider (Fireworks, OpenAI, OpenRouter, Together, …) alongside local models — see Cloud Offload. (Experimental.)

Model Library

Model Manager

Lemonade supports a wide variety of LLMs (GGUF, FLM, and ONNX), whisper, stable diffusion, etc. models across CPU, GPU, and NPU.

Use lemonade pull or the built-in Model Manager to download models. Custom GGUF/ONNX models can be pulled from Hugging Face or ModelScope, with their source retained for future updates.

Browse all built-in models →


Supported Configurations

Lemonade supports multiple inference engines for LLM, speech, TTS, and image generation, and each has its own backend and hardware requirements.

ModalityEngineBackendDeviceOS
Text generationllamacppsystemx86_64/ARM64 CPU, GPULinux
metalApple Silicon GPUmacOS
cudaNVIDIA GPUs (Turing or newer)**Windows, Linux
vulkanx86_64 CPU, AMD iGPU, AMD dGPU; ARM64 CPU/GPU (Linux)Windows, Linux
rocmAMD GPUs supported by ROCmWindows, Linux
cpux86_64 CPU; ARM64 CPU (Linux)Windows, Linux
flmnpuXDNA2 NPUWindows, Linux
ryzenai-llmnpuXDNA2 NPUWindows
vllm (experimental)rocmStrix Halo iGPU (gfx1151)Linux
Speech-to-textwhispercppnpuXDNA2 NPUWindows
metalApple Silicon GPUmacOS
vulkanx86_64 CPUWindows, Linux
rocmSupported AMD ROCm iGPU/dGPU families*Windows, Linux
cpux86_64 CPUWindows, Linux
moonshinecpux86_64/arm64 CPUWindows, Linux, macOS
Text-to-speechkokorometalApple Silicon GPUmacOS
cpux86_64 CPUWindows, Linux
openmoss (experimental)cudaNVIDIA GPUsWindows, Linux
vulkanVulkan-capable GPUsWindows, Linux
rocmAMD GPUs (ROCm via TheRock)Windows, Linux
Audio generationthinksound (experimental)cudaNVIDIA GPUsWindows, Linux
vulkanVulkan-capable GPUsWindows, Linux
rocmSupported AMD ROCm iGPU/dGPU families (ROCm via TheRock)Windows, Linux
acestep (experimental)cudaNVIDIA GPUsWindows, Linux
vulkanVulkan-capable GPUsWindows, Linux
rocmSupported AMD ROCm iGPU/dGPU families (ROCm via TheRock)Windows, Linux
Image generationsd-cppmetalApple Silicon GPUmacOS
cudaNVIDIA GPUs (Turing or newer)**Windows, Linux
vulkanVulkan-capable GPUsWindows, Linux
rocmSupported AMD ROCm iGPU/dGPU families*Windows, Linux
cpux86_64 CPUWindows, Linux
thenoise (experimental)rocmSupported AMD ROCm iGPU familiesLinux
3D generationtrellis (experimental)cudaNVIDIA GPUsWindows, Linux
vulkanVulkan-capable GPUsWindows, Linux
rocmSupported AMD ROCm iGPU/dGPU families (ROCm via TheRock)Windows, Linux
Text classificationonnxruntime (experimental)cpux86_64 CPUWindows
cpux86_64/arm64 CPULinux
cpuarm64 CPUmacOS

To check exactly which recipes/backends are supported on your own machine, run:

lemonade backends
* See supported AMD ROCm platforms
ArchitecturePlatform SupportGPU Models
gfx1151 (STX Halo)Windows, UbuntuRyzen AI MAX+ Pro 395
gfx120X (RDNA4)Windows, UbuntuRadeon AI PRO R9700, RX 9070 XT/GRE/9070, RX 9060 XT
gfx110X (RDNA3)Windows, UbuntuRadeon PRO W7900/W7800/W7700/V710, RX 7900 XTX/XT/GRE, RX 7800 XT, RX 7700 XT
** See supported NVIDIA CUDA platforms
Compute CapabilityArchitectureGPU Models
sm_75TuringRTX 20-series, GTX 16-series, T4
sm_80 / sm_86AmpereRTX 30-series, A100, A40
sm_89Ada LovelaceRTX 40-series, L40, L4
sm_90HopperH100, H200
sm_100 / sm_120BlackwellRTX 50-series, B100, B200

Project Roadmap

Lemonade's roadmap is defined by a set of working groups. Visit the landing page here to learn each group's goal and roadmap.

Integrate Embeddable Lemonade in Your Application

Embeddable Lemonade is a binary version of Lemonade that you can bundle into your own app to give it a portable, auto-optimizing, multi-modal local AI stack. This lets users focus on your app, with zero Lemonade installers, branding, or telemetry.

Check out the Embeddable Lemonade guide.

Connect Lemonade Server to Your Application

You can use any OpenAI-compatible client library by configuring it to use http://localhost:13305/v1 as the base URL. A table containing official and popular OpenAI clients on different languages is shown below.

Feel free to pick and choose your preferred language.

Python Client Example

from openai import OpenAI

# Initialize the client to use Lemonade Server
client = OpenAI(
    base_url="http://localhost:13305/api/v1",
    api_key="lemonade"  # required but unused
)

# Create a chat completion
completion = client.chat.completions.create(
    model="Gemma-4-E2B-it-GGUF",  # or any other available model
    messages=[
        {"role": "user", "content": "What is the capital of France?"}
    ]
)

# Print the response
print(completion.choices[0].message.content)

Click to learn more about the available APIs and how to embed Lemonade in your own application.

FAQ

To read our frequently asked questions, see our FAQ Guide

Contributing

Lemonade is built by the local AI community! If you would like to contribute to this project, please check out our contribution guide.

Maintainers

This is a community project with many maintainers, please see the maintainers list here to see their subject areas. You can reach us by filing an issue or joining our Discord. This project is sponsored by AMD.

Code Signing Policy

Free code signing provided by SignPath.io, certificate by SignPath Foundation.

Privacy policy: This program will not transfer any information to other networked systems unless specifically requested by the user or the person installing or operating it. When the user requests a model download or registry lookup, Lemonade may contact Hugging Face Hub (see their privacy policy) or ModelScope, according to the model source selected by the user or packager.

License and Attribution

This project is:

Frequently Asked Questions

What is lemonade?

lemonade is an open-source mcp servers skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by lemonade-sdk. Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk. It has 5,441 GitHub stars.

Is lemonade safe to use?

Yes. lemonade passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.

How do I install lemonade?

Clone the repository with "git clone https://github.com/lemonade-sdk/lemonade" and add it to your Claude Code skills directory (see the Installation section above).

What programming language is lemonade written in?

lemonade is primarily written in C++. It is open-source under lemonade-sdk on GitHub, so you can review or fork the full source.

Are there alternatives to lemonade?

Yes. SkillsLLM lists many other MCP Servers skills you can browse and compare side by side. Open the MCP Servers category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh lemonade against similar tools.

Comments (0)

No comments yet. Be the first to share your thoughts!

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP Serversapisai-tools
View details

Scrapling

by D4Vinci

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

75,9137,581Python
MCP Servers
View details

TrendRadar

by sansan0

⭐AI-driven public opinion & trend monitor with multi-platform aggregation, RSS, and smart alerts.🎯 告别信息过载,你的 AI 舆情监控助手与热点筛选工具!聚合多平台热点 + RSS 订阅,支持关键词精准筛选。AI 智能筛选新闻 + AI 翻译 + AI 分析简报直推手机,也支持接入 MCP 架构,赋能 AI 自然语言对话分析、情感洞察与趋势预测等。支持 Docker ,数据本地/云端自持。集成微信/飞书/钉钉/Telegram/邮件/ntfy/bark/slack 等渠道智能推送。

61,65224,883Python
MCP Servers
View details

context7

by upstash

Context7 Platform -- Up-to-date code documentation for LLMs and AI code editors

61,0602,938TypeScript
MCP Servers
View details

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

39,9393,219C
MCP Servers
View details

Developers Also Liked

Based on votes and bookmarks from developers who liked this skill

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI Agentsai-agentsanthropicclaude-code
View details
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI Agentsai-agentsbrainstorming
View details

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP Serversapisai-tools
View details

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI Agentsai-agentsanthropicclaude-code
View details

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI Agentsclaude-codeai-tools
View details