Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/build-container.yml
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ jobs:
id: meta
uses: docker/metadata-action@v5
with:
images: ghcr.io/lemonade-sdk/lemonade/build-environment
images: ghcr.io/${{ github.repository }}/build-environment
tags: |
type=raw,value=ubuntu${{ matrix.ubuntu_version }}
type=raw,value=latest,enable=${{ matrix.ubuntu_version == '26.04' }}
Expand Down
18 changes: 17 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -136,7 +136,7 @@ Lemonade supports multiple inference engines for LLM, speech, TTS, and image gen
</thead>
<tbody>
<tr>
<td rowspan="9"><strong>Text generation</strong></td>
<td rowspan="12"><strong>Text generation</strong></td>
<td rowspan="6"><code>llamacpp</code></td>
<td><code>system</code></td>
<td><code>x86_64</code>/ARM64 CPU, GPU</td>
Expand Down Expand Up @@ -185,6 +185,22 @@ Lemonade supports multiple inference engines for LLM, speech, TTS, and image gen
<td>Strix Halo iGPU (gfx1151)</td>
<td>Linux</td>
</tr>
<tr>
<td rowspan="3"><code>mlx-engine</code> (experimental)</td>
<td><code>metal</code></td>
<td>Apple Silicon Metal GPU</td>
<td>macOS</td>
</tr>
<tr>
<td><code>rocm</code></td>
<td>AMD ROCm GPUs (RDNA3/RDNA4)</td>
<td>Linux</td>
</tr>
<tr>
<td><code>cpu</code></td>
<td>CPU fallback</td>
<td>Linux, macOS</td>
</tr>
<tr>
<td rowspan="6"><strong>Speech-to-text</strong></td>
<td rowspan="5"><code>whispercpp</code></td>
Expand Down
4 changes: 3 additions & 1 deletion docs/assets/models.js
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ const RECIPE_PRIORITY = [
'flm',
'kokoro',
'llamacpp',
'mlx-engine',
'moonshine',
'onnxruntime',
'openmoss',
Expand All @@ -30,7 +31,8 @@ const RECIPE_DISPLAY_NAMES = {
acestep: 'ACE-Step',
onnxruntime: 'ONNX Runtime',
trellis: 'TRELLIS.2',
openmoss: 'OpenMOSS TTS'
openmoss: 'OpenMOSS TTS',
'mlx-engine': 'MLX Engine (Apple Silicon)'
};
/* END GENERATED: models-js-recipes */

Expand Down
19 changes: 19 additions & 0 deletions docs/dev/backends-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ the generator instead. Prose outside the markers is preserved. -->
| `flm` | FastFlowLM NPU | no | yes | npu |
| `kokoro` | Kokoro | no | no | cpu, metal |
| `llamacpp` | Llama.cpp GPU | yes | yes | cpu, cuda, metal, rocm, system, vulkan |
| `mlx-engine` | MLX Engine | yes | yes | cpu, metal, rocm |
| `moonshine` | Moonshine | no | no | cpu |
| `onnxruntime` | ONNX Runtime | no | no | cpu |
| `openmoss` | OpenMOSS TTS | yes | no | cuda, rocm, vulkan |
Expand Down Expand Up @@ -41,6 +42,9 @@ the generator instead. Prose outside the markers is preserved. -->
| `llamacpp` | vulkan | linux, windows | amd_gpu; cpu (arm64, x86_64) |
| `llamacpp` | rocm | linux, windows | amd_gpu (gfx103X, gfx110X, gfx1150, gfx1151, gfx1152, gfx120X, gfx942, gfx950) |
| `llamacpp` | cpu | linux, windows | cpu (arm64, x86_64) |
| `mlx-engine` | metal | macos | metal |
| `mlx-engine` | rocm | linux | amd_gpu (gfx110X, gfx1150, gfx1151, gfx120X) |
| `mlx-engine` | cpu | linux, macos | cpu (arm64, x86_64) |
| `moonshine` | cpu | windows | cpu (x86_64) |
| `moonshine` | cpu | linux | cpu (arm64, x86_64) |
| `moonshine` | cpu | macos | cpu (arm64) |
Expand Down Expand Up @@ -102,6 +106,14 @@ the generator instead. Prose outside the markers is preserved. -->
| `llamacpp_device` | `--llamacpp-device` | DEVICES | "" | Comma-separated list of accelerator devices to use (e.g. Vulkan0) |
| `llamacpp_args` | `--llamacpp-args` | ARGS | "" | Custom arguments to pass to llama-server |

#### `mlx-engine` — MLX Engine

| Option | CLI flag | Type | Default | Description |
|--------|----------|------|---------|-------------|
| `ctx_size` | `--ctx-size` | SIZE | -1 | Context size for the model |
| `mlx_backend` | `--mlx-backend` | BACKEND | "" | MLX backend to use (metal, rocm, cpu) |
| `mlx_args` | `--mlx-args` | ARGS | "" | Extra arguments passed to lemon-mlx-engine |

#### `moonshine` — Moonshine

| Option | CLI flag | Type | Default | Description |
Expand Down Expand Up @@ -268,6 +280,13 @@ the generator instead. Prose outside the markers is preserved. -->
| `nomic-embed-text-v1-GGUF` | 0.0781 | embeddings |
| `nomic-embed-text-v2-moe-GGUF` | 0.51 | embeddings |

#### `mlx-engine` — MLX Engine (2 models)

| Model | Size (GB) | Labels |
|-------|-----------|--------|
| `Qwen3-0.6B-MLX` | 0.42 | reasoning |
| `Qwen3-4B-MLX` | 2.3 | reasoning |

#### `moonshine` — Moonshine (3 models)

| Model | Size (GB) | Labels |
Expand Down
8 changes: 8 additions & 0 deletions docs/guide/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -408,6 +408,14 @@ The following options are available depending on the recipe being used:
| Option | Description | Default |
|--------|-------------|---------|
| `--openmoss BACKEND` | OpenMOSS TTS backend to use | Auto-detected |

#### MLX Engine (`mlx-engine` recipe)

| Option | Description | Default |
|--------|-------------|---------|
| `--ctx-size SIZE` | Context size for the model | auto |
| `--mlx-backend BACKEND` | MLX backend to use (metal, rocm, cpu) | Auto-detected |
| `--mlx-args ARGS` | Extra arguments passed to lemon-mlx-engine | `""` |
<!-- END GENERATED: cli-recipe-options -->
**Notes:**
- Unspecified options will use the backend's default values
Expand Down
3 changes: 3 additions & 0 deletions docs/guide/configuration/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,9 @@ Values set in the user's `config.json` always take precedence over these seeded
},
"log_level": "info",
"max_loaded_models": 1,
"mlx-engine": {
"backend": "auto"
},
"models_dir": "auto",
"moonshine": {
"args": "",
Expand Down
2 changes: 1 addition & 1 deletion docs/guide/configuration/custom-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ Supported registration flags:
|------|-------------|
| `--source SOURCE` | Remote registry for every checkpoint in this model: `huggingface` (default) or `modelscope`. |
| `--checkpoint TYPE CHECKPOINT` | Add a checkpoint entry. Repeat for multi-file models such as `main` + `mmproj` or `main` + `vae`. |
| `--recipe RECIPE` | Recipe to associate with the new `user.*` model. Common values: <!-- BEGIN GENERATED: recipe-values -->`llamacpp`, `whispercpp`, `moonshine`, `kokoro`, `sd-cpp`, `flm`, `ryzenai-llm`, `vllm`, `thinksound`, `acestep`, `onnxruntime`, `trellis`, `openmoss`, `collection.omni`<!-- END GENERATED: recipe-values -->. |
| `--recipe RECIPE` | Recipe to associate with the new `user.*` model. Common values: <!-- BEGIN GENERATED: recipe-values -->`llamacpp`, `whispercpp`, `moonshine`, `kokoro`, `sd-cpp`, `flm`, `ryzenai-llm`, `vllm`, `thinksound`, `acestep`, `onnxruntime`, `trellis`, `openmoss`, `mlx-engine`, `collection.omni`<!-- END GENERATED: recipe-values -->. |
| `--label LABEL` | Add a label to the new model. Repeatable. Valid labels include `coding`, `embeddings`, `hot`, `mtp`, `reasoning`, `reranking`, `tool-calling`, `vision`. |
| `--components MODEL [MODEL ...]` | Components for an omni collection (see below). Use with `--recipe collection.omni`. |

Expand Down
3 changes: 3 additions & 0 deletions src/cpp/resources/defaults.json
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,9 @@
},
"log_level": "info",
"max_loaded_models": 1,
"mlx-engine": {
"backend": "auto"
},
"models_dir": "auto",
"moonshine": {
"args": "",
Expand Down
7 changes: 3 additions & 4 deletions src/cpp/server/backends/mlx/mlx_server.cpp
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
#include "lemon/backends/mlx/mlx_server.h"
#include "lemon/backends/mlx/mlx.h"
#include "lemon/backends/backend_ops.h"
#include "lemon/backends/backend_utils.h"
#include "lemon/backend_manager.h"
Expand Down Expand Up @@ -112,7 +113,7 @@ void MlxServer::load(const std::string& model_name,
device_type_ = (mlx_backend == "cpu") ? DEVICE_CPU : DEVICE_GPU;

// Install mlx-engine binary if needed.
backend_manager_->install_backend(SPEC.recipe, mlx_backend);
backend_manager_->install_backend(mlx::spec()->recipe, mlx_backend);

// MLX identifies models by HuggingFace repo-id or a local directory path.
// The ModelManager resolves local paths when available; fall back to the
Expand All @@ -130,7 +131,7 @@ void MlxServer::load(const std::string& model_name,

port_ = choose_port();

std::string executable = BackendUtils::get_backend_binary_path(SPEC, mlx_backend);
std::string executable = BackendUtils::get_backend_binary_path(*mlx::spec(), mlx_backend);

std::vector<std::string> args;
// Positional model argument — pre-load mode.
Expand Down Expand Up @@ -222,8 +223,6 @@ json MlxServer::responses(const json& request) {
);
}

} // namespace backends

namespace {
class MlxOps : public BackendOps {
public:
Expand Down
Loading