EXL3 inference on Apple Silicon

Your models.Your Mac.

MLXL3 is an EXL3 inference engine for macOS, with a native Mac app and CLI. Your models run locally on Apple Silicon.

Apple Silicon M1–M5

macOS 26.2 or later

Open source MIT license

Start with
a conversation.

Choose a model. Add your documents. Keep your conversations.
A native Mac app, from the first message to the last token.

Native MLXL3 interface: conversations on the left, a Markdown response and MCP tool call in the center, message input below.
MLXL3 DesktopActual app interface · demo conversation

Your files, in context.

Attach a PDF, notes or code. Their text accompanies your message and stays saved with the conversation.

Readable as it streams.

Markdown, tables, LaTeX and highlighted code appear as the response streams. Copy code blocks in one click.

Connect your tools.

Enable MCP to connect the tools you choose. Reasoning, tool calls and the response each have their own place.

Pick up where you left off.

Your conversations are saved locally. Stop a generation, keep the partial response and return to your conversations.

The MLXL3 engine

EXL3 in memory.
Metal at work.

The MLXL3 EXL3 inference engine reads quantized weights directly. Rust orchestrates inference; MLX and custom Metal kernels run computations on Apple Silicon GPUs.

  1. EXL3Quantized weights
  2. RustNative runtime
  3. MLXCompute graphs
  4. MetalApple GPU

EXL3, read directly.

Inference does not rebuild a full dense copy of every weight. The engine works with the EXL3 representation.

Propose. Verify. Accept.

On supported Qwen models, MTP proposes tokens. The target model verifies them before delivery, using greedy decoding.

Read the MTP validation

Models,
by architecture.

Import an EXL3 checkpoint or search Hugging Face from the app. MLXL3 checks its structure before loading it.

Qwen3.5 / 3.6 / 3.8

Dense and MoE models with hybrid Gated DeltaNet attention. Text chat and MTP on supported checkpoints.

qwen3_5 · qwen3_5_moe
Liquid AILFM2 / LFM2 MoE

Recurrent-state models. The engine supports dense EXL3 weights and routed experts for text chat.

lfm2 · lfm2_moe
Gemma4

Text chat and native tool-call parsing. Image and audio inputs are not supported.

gemma4
Ling3 / Bailing V3

Text chat, hybrid architecture and MCP tool calls in Ling’s native format.

bailing_hybrid

The EXL3 format alone does not guarantee compatibility. Architecture, tensor shapes and engine requirements must also match.

The same engine.
In your terminal.

Inspect checkpoints, manage downloads and chat in streaming mode with the native mlxl3 client.

Read the CLI documentation
Register and run a model
# Replace the path with your EXL3 folder
mlxl3 register my-model /path/to/model

# Start a conversation
mlxl3 run my-model

The CLI client must be installed or built from source.

From download
to first message.

The app includes the Rust engine, MLX runtime and Metal assets. No Python or Homebrew installation is required.

  1. Install the app.

    Open the DMG, then drag MLXL3 Desktop into Applications.

  2. Choose a model.

    Download an EXL3 variant from Hugging Face in the app, or import an existing folder.

  3. Start a conversation.

    Load the model and send your first message. Model weights are not included in the app download.

The current distribution is signed ad hoc. If macOS blocks the first launch, allow the app in System Settings → Privacy & Security.

Before you start.

Can I run EXL3 models on macOS?

On an Apple Silicon Mac, install MLXL3 Desktop, download or import a supported EXL3 checkpoint, then load the model. The desktop app includes its native engine; no Python, Homebrew or CUDA installation is required. Follow the installation steps.

Which EXL3 models are supported?

Supported architectures include Qwen 3.5, 3.6 and 3.8 (dense and MoE), Liquid AI LFM2 and LFM2 MoE, Gemma 4, and Ling 3 / Bailing V3. The loader checks tensor shapes and engine constraints before loading. An EXL3 file alone does not guarantee compatibility. See model support for the current scope.

Is MLXL3 a port of ExLlamaV3?

MLXL3 is an independent Apple Silicon runtime. ExLlamaV3 provides the EXL3 format and numerical reference. MLXL3 executes inference through Rust, MLX and custom Metal kernels on Apple GPUs. The third-party notices document the code’s provenance and licenses.

Which Macs are supported?

An Apple Silicon Mac (M1–M5) running macOS 26.2 or later. Memory requirements depend on the model, quantization and context. The project’s performance measurements are run on M5.

Does everything stay on my Mac?

Inference runs locally and conversations are saved on your Mac. Downloads and updates use the network. Enabled MCP tools may send requests to the services you configure.

Which documents can I attach?

Text-based PDFs, TXT, Markdown, CSV, JSON and code files. MLXL3 extracts their text; it does not analyze images. Scanned or protected PDFs need preparation first. The model’s context window remains the limit.

Is the project open source?

MLXL3’s code is available under the MIT license. Models and third-party components retain their own licenses.

Contributors.

View on GitHub

Run your models.
Locally.

Download for Mac

Desktop 1.4.0 · Apple Silicon
macOS 26.2 or later