Model Gallery

Discover and install AI models from our curated collection

14 models available
1 repositories
Documentation

Find Your Perfect Model

Filter by Model Type

Browse by Tags

qwen3.8-27b-exl3-vllm-cpp
Qwen3.8-27B EXL3 3.5bpw served by vllm.cpp, LocalAI's C++ vLLM-style runtime. The published checkpoint generates on CUDA and measured 16.7 tokens/s on GB10 with the pinned revision and the limits configured here. This is the target-only setup. Use the DFlash2 variant for the measured speculative-decoding configuration. The entry downloads the complete revision-pinned repository, including its configuration, tokenizer, index, and safetensors shards.

Repository: localai

qwen3.8-27b-dflash2-exl3-vllm-cpp
Qwen3.8-27B EXL3 with its EXL3 DFlash2 companion, served by vllm.cpp. On GB10 this pinned pair measured 48.7 tokens/s at a seven-token draft budget, versus 16.7 tokens/s target-only, with token-identical greedy output. LocalAI stages both complete Hugging Face repositories before load. The content-addressed companion snapshot is passed to the backend as the draft model, while the speculative method and seven-token budget remain fixed.

Repository: localai

deepseek-v4-flash-spark-exl3-vllm-cpp
DeepSeek V4 Flash's Spark and GB10-oriented REAP-K216 EXL3 checkpoint, served by vllm.cpp. It needs CUDA and roughly 100 GiB for its large rank-sliced checkpoint. The repository is pinned to its latest recorded revision. vllm.cpp's existing runtime evidence measured the older 22f28d32b9b29b4352eaa380ff8c2c170b2847ab revision; this entry does not claim that the newer revision has passed the same end-to-end gate.

Repository: localai

deepseek-v4-flash-exl3-3bpw-vllm-cpp
Experimental non-Spark DeepSeek V4 Flash EXL3 3.0bpw checkpoint served by vllm.cpp. This is the complete, non-REAP layout and requires a large multi-GPU CUDA system. The publisher describes the artifact as structurally complete but has not passed end-to-end generation. Treat this entry as an integration target, not as a correctness- or performance-gated configuration.

Repository: localai

qwen3.6-27b-nvfp4-vllm-cpp
Qwen3.6-27B in NVFP4, served by vllm.cpp: LocalAI's own C++ port of vLLM, with no Python at inference time. This is the reference text-generation checkpoint the engine is gated on, token-for-token identical to vLLM's own greedy output over the 235-prompt correctness battery, and measured at or above vLLM's throughput at every concurrency from 1 to 32. The weights are PINNED to revision 890bdef7. That pin is load-bearing, not housekeeping: the same repository name was later re-quantized to FP8 W8A8 throughout, so an unpinned copy of this entry serves entirely different weights with no error and none of the measured behaviour above. Needs a Blackwell-class NVIDIA GPU (NVFP4 has no kernel on older architectures) and roughly 25 GB of weights plus KV cache. Tool calling and the thinking split are parsed inside the engine.

Repository: localaiLicense: apache-2.0

qwen3.6-27b-nvfp4-mtp-vllm-cpp
Qwen3.6-27B NVFP4 on vllm.cpp with MTP speculative decoding enabled. MTP (Multi-Token Prediction) drafts from a head that ships inside the target checkpoint's own mtp.* tensors, so there is no second model to download and no extra weights to manage. The verifier accepts roughly 85% of drafted tokens on prose and 92% on code, worth about 1.5x to 1.6x the decode throughput of the same weights with speculation off, and it holds that lead at concurrency 2, 4 and 8. Same weights and same revision pin as qwen3.6-27b-nvfp4-vllm-cpp; install that entry instead if you would rather not spend the extra memory. The speculative state (a doubled recurrent-state slot plus the draft cache and head) costs roughly 3.6 GB on top of the base footprint. MTP here is depth 1 by construction: the engine refuses num_speculative_tokens above 1 for this method.

Repository: localaiLicense: apache-2.0

qwen3.6-27b-nvfp4-dflash-vllm-cpp
Qwen3.6-27B NVFP4 on vllm.cpp with DFlash block-diffusion speculative decoding: the fastest configuration of this model the engine ships. Where MTP drafts one token at a time, DFlash drafts a whole 16-token block in a single non-autoregressive pass from a separate 3.5 GB drafter, then the target verifies the block in one step. At concurrency 1 that measures 2.9x the throughput of the same weights with speculation off, and at or above vLLM's own DFlash-on decode. Both checkpoints are installed for you: the target as a revision-pinned snapshot, the drafter into models/Qwen3.6-27B-DFlash, which is where the backend looks when speculative_config.model names it. The drafter shares the target's embed_tokens and lm_head, so the two are not independently swappable. Needs a Blackwell-class NVIDIA GPU and roughly 28 GB of weights in total.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-nvfp4-vllm-cpp
Qwen3.6-35B-A3B in NVFP4, served by vllm.cpp. A 35B mixture-of-experts model with roughly 3B parameters active per token, so it reads like a much larger model while costing about as much per token as a small one. This is the engine's gated MoE checkpoint: token-for-token identical to vLLM over the 315-prompt battery on both the synchronous and asynchronous paths, at 0.92x to 0.97x vLLM's throughput from concurrency 1 to 32. The architecture is a gated-delta-net hybrid, so automatic prefix caching is off by default here where it would be on for a dense model. That is the engine's own default and this entry does not override it. Needs a Blackwell-class NVIDIA GPU and roughly 23 GB of weights plus KV cache.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-nvfp4-mtp-vllm-cpp
Qwen3.6-35B-A3B NVFP4 on vllm.cpp with MTP speculative decoding enabled. The draft head ships inside the checkpoint's own mtp.* tensors, so there is no second model to download. On this model the speculative path is token-exact against speculation-off on both the synchronous and asynchronous schedulers. Same weights as qwen3.6-35b-a3b-nvfp4-vllm-cpp; install that entry instead if you would rather not spend the extra memory on speculative state. MTP is depth 1 by construction on this engine.

Repository: localaiLicense: apache-2.0

qwen3-coder-30b-a3b-vllm-cpp
Qwen3-Coder-30B-A3B on vllm.cpp: a coding and agentic-tool-use model, 30B total parameters with about 3B active per token, gated token-exact against vLLM on this engine. The tool-call parser is named explicitly rather than auto-detected, and that matters here. Qwen3-Coder's tool dialect is byte-identical on the wire to another family's, so template sniffing cannot separate the two and would fall back to the wrong parser. With qwen3_coder named, tool calls arrive as real tool_calls on the OpenAI response. This is the bf16 checkpoint, roughly 57 GB of weights, which is what the engine was gated on. Being bf16 rather than NVFP4 it does not need Blackwell on its own account, but LocalAI's CUDA images for this backend are currently built for Blackwell-family GPUs only, so on an older card use the CPU build.

Repository: localaiLicense: apache-2.0

qwen3-4b-vllm-cpp
Qwen3-4B on vllm.cpp, in bf16. The small end of the engine's gated dense family, which reaches parity with vLLM on every axis at concurrency 1. bf16 rather than NVFP4 on purpose: this is the entry that runs where the flagship NVFP4 checkpoints cannot, including Apple Silicon via Metal, Vulkan and plain CPU. Roughly 8 GB of weights, plus about 4.5 GB of KV cache at the context configured here. Tool calling and the thinking split are parsed inside the engine.

Repository: localaiLicense: apache-2.0

qwen3-0.6b-vllm-cpp
Qwen3-0.6B on vllm.cpp, in bf16. Roughly 1.4 GB of weights, which makes it the cheapest way to confirm a vllm-cpp install actually serves before committing disk and memory to one of the large checkpoints. It runs anywhere the backend does, CPU included, and it is a real chat model rather than a stub, so tool calling and the thinking split can be exercised on it too.

Repository: localaiLicense: apache-2.0

minimax-h3-fl2va-q4
MiniMax-H3 served by vllm.cpp, LocalAI's own C++ port of vLLM. It generates video AND audio jointly from a text prompt, so a clip comes back as an MP4 with a real soundtrack rather than a silent render: ask for speech in the prompt and the model lip-syncs it. This is the Q4_K_M quantisation of the FL2VA partition, which serves text-to-video (t2va) and first/last-frame conditioning (fl2va). Reference conditioning (ref2va) is a different checkpoint and is refused by this one. Roughly 40 GB of weights across five files, plus the two VAE configs that carry the latent statistics. The default canvas is 1344x768 at 124 frames and 24 fps, about 5.2 seconds. Generation is slow — measured at roughly 176 s per denoise step at that canvas on a 20-SM device, so the 50-step default is a multi-hour job. Muxing the finished frames needs ffmpeg on the host.

Repository: localaiLicense: other

minimax-h3-ref2va-q4
MiniMax-H3 served by vllm.cpp, LocalAI's own C++ port of vLLM. It generates video AND audio jointly, so a clip comes back as an MP4 with a real soundtrack rather than a silent render. This is the Q4_K_M quantisation of the Ref2VA partition, the one that takes REFERENCE conditioning: a reference image, a reference clip, or reference audio, prepended as their own blocks so the subject or style carries into the generated video. For plain text-to-video or first/last-frame conditioning use minimax-h3-fl2va-q4 instead - the two partitions are separate checkpoints and each refuses the other's tasks. Use this Q4_K_M build, NOT the NVFP4 Ref2VA weights: NVFP4 renders a multicolour patch grid, and it took three investigations upstream to establish that the fault is the quantisation rather than the reference path. On Q4_K_M the same code renders coherently. Roughly 40 GB of weights across five files, plus the two VAE configs that carry the latent statistics. The default canvas is 1344x768 at 124 frames and 24 fps, about 5.2 seconds. Generation is slow - roughly 176 s per denoise step at that canvas on a 20-SM device, so the 50-step default is a multi-hour job. Muxing the finished frames needs ffmpeg on the host.

Repository: localaiLicense: other