Model Gallery

Discover and install AI models from our curated collection

4 models available
1 repositories
Documentation

Find Your Perfect Model

Filter by Model Type

Browse by Tags

qwen3.8-27b-exl3-vllm-cpp
Qwen3.8-27B EXL3 3.5bpw served by vllm.cpp, LocalAI's C++ vLLM-style runtime. The published checkpoint generates on CUDA and measured 16.7 tokens/s on GB10 with the pinned revision and the limits configured here. This is the target-only setup. Use the DFlash2 variant for the measured speculative-decoding configuration. The entry downloads the complete revision-pinned repository, including its configuration, tokenizer, index, and safetensors shards.

Repository: localai

qwen3.8-27b-dflash2-exl3-vllm-cpp
Qwen3.8-27B EXL3 with its EXL3 DFlash2 companion, served by vllm.cpp. On GB10 this pinned pair measured 48.7 tokens/s at a seven-token draft budget, versus 16.7 tokens/s target-only, with token-identical greedy output. LocalAI stages both complete Hugging Face repositories before load. The content-addressed companion snapshot is passed to the backend as the draft model, while the speculative method and seven-token budget remain fixed.

Repository: localai

deepseek-v4-flash-spark-exl3-vllm-cpp
DeepSeek V4 Flash's Spark and GB10-oriented REAP-K216 EXL3 checkpoint, served by vllm.cpp. It needs CUDA and roughly 100 GiB for its large rank-sliced checkpoint. The repository is pinned to its latest recorded revision. vllm.cpp's existing runtime evidence measured the older 22f28d32b9b29b4352eaa380ff8c2c170b2847ab revision; this entry does not claim that the newer revision has passed the same end-to-end gate.

Repository: localai

deepseek-v4-flash-exl3-3bpw-vllm-cpp
Experimental non-Spark DeepSeek V4 Flash EXL3 3.0bpw checkpoint served by vllm.cpp. This is the complete, non-REAP layout and requires a large multi-GPU CUDA system. The publisher describes the artifact as structurally complete but has not passed end-to-end generation. Treat this entry as an integration target, not as a correctness- or performance-gated configuration.

Repository: localai