Model Gallery

Discover and install AI models from our curated collection

239 models available
1 repositories
Documentation

Find Your Perfect Model

Filter by Model Type

Browse by Tags

llm-jp-4-33b-thinking-q4
LLM-jp-4-33B-thinking is an Apache-2.0 Japanese and English reasoning model from Japan's National Institute of Informatics. Its dense Llama architecture has 33 billion parameters and a 65K-token context window. The model was aligned with supervised fine-tuning and DPO for multi-turn conversation and instruction following. This default entry uses the 20.2 GB Q4_K_M GGUF. The official 66.4 GB BF16 weights are available as a higher-fidelity variant.

Repository: localaiLicense: apache-2.0

llm-jp-4-33b-thinking-bf16
LLM-jp-4-33B-thinking in the official 66.4 GB BF16 GGUF format. This variant preserves the original model precision for hosts with enough memory.

Repository: localaiLicense: apache-2.0

Attention: Trust Remote Code is required for this model
wemm-embedding-2b
WeMM-Embedding-2B is Tencent's Apache-2.0 multilingual embedding model built on Qwen3.5. This entry serves the original bfloat16 safetensors with LocalAI's Transformers backend and produces 2,048-dimensional normalized embeddings for text retrieval, semantic search, and RAG. The upstream model can also embed images and videos. LocalAI currently exposes text input through its embeddings API for this backend.

Repository: localaiLicense: apache-2.0

Attention: Trust Remote Code is required for this model
wemm-embedding-4b
WeMM-Embedding-4B is Tencent's mid-sized Apache-2.0 multilingual embedding model built on Qwen3.5. This entry serves the original bfloat16 safetensors with LocalAI's Transformers backend and produces 2,560-dimensional normalized embeddings for text retrieval, semantic search, and RAG. The upstream model can also embed images and videos. LocalAI currently exposes text input through its embeddings API for this backend.

Repository: localaiLicense: apache-2.0

Attention: Trust Remote Code is required for this model
wemm-embedding-9b
WeMM-Embedding-9B is Tencent's largest Apache-2.0 multilingual embedding model built on Qwen3.5. This entry serves the original bfloat16 safetensors with LocalAI's Transformers backend and produces 4,096-dimensional normalized embeddings for text retrieval, semantic search, and RAG. The upstream model can also embed images and videos. LocalAI currently exposes text input through its embeddings API for this backend.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q4
IBM Granite 4.2 3B is a compact multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q8
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q4
IBM Granite 4.2 8B is a multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q8
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q4
IBM Granite 4.2 30B is the family's flagship multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q8
IBM Granite 4.2 30B in the higher-fidelity Q8_0 GGUF format. It is the family's flagship multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

hy-mt2-1.8b-q4
Hy-MT2-1.8B is Tencent's compact multilingual translation model. It follows translation instructions across 33 languages and supports tasks such as terminology control, style transfer, and structure-preserving translation. This default entry uses the 1.1 GB Q4_K_M GGUF. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: apache-2.0

hy-mt2-1.8b-q8
Hy-MT2-1.8B in the higher-quality 1.9 GB Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

nemotron-3.5-lightning-30b-a3b-q4
NVIDIA Nemotron 3.5 Lightning is a text-only hybrid Mamba-2, attention, and mixture-of-experts model with 30B total parameters and 3B active parameters. It targets reasoning, coding, tool use, multilingual chat, and long-context agent workflows, with a context window of up to one million tokens. This entry uses the official Q4_K_M GGUF. Automatic variant selection can choose the smaller NVFP4 build or the higher-quality Q8_0 build when it fits.

Repository: localaiLicense: openmdw-1.1

nemotron-3.5-lightning-30b-a3b-nvfp4
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official NVFP4 GGUF format. This is the smallest linked build and retains the model's reasoning, coding, tool-use, multilingual, and long-context capabilities.

Repository: localaiLicense: openmdw-1.1

nemotron-3.5-lightning-30b-a3b-q8
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official high-quality Q8_0 GGUF format for hosts with enough memory.

Repository: localaiLicense: openmdw-1.1

muse-glimmer-30b
Muse Glimmer is Meta Superintelligence Labs' Apache-2.0 dense 30B model for autonomous agentic work, coding, tool use, long-horizon reasoning, and multimodal understanding. It supports more than 100 languages, interleaved text and image input through its 1.8B-parameter perception encoder, and a 131K-token context window. This entry uses the publisher's higher-quality dynamic K-quant GGUF and official quantized vision projector. Automatic variant selection can use the smaller 17 GB quantization or a DFlash-accelerated build when it fits.

Repository: localaiLicense: apache-2.0

muse-glimmer-30b-dflash
Muse Glimmer's higher-quality dynamic K-quant GGUF with the official quantized perception encoder and DFlash drafter. DFlash proposes blocks of up to 16 tokens for the target to verify in parallel, accelerating output without changing model quality. Flash attention is enabled for this path.

Repository: localaiLicense: apache-2.0

muse-glimmer-30b-17gb
Muse Glimmer's smaller 17 GB K-quant GGUF with the official quantized perception encoder. It preserves the model's agentic, coding, tool-use, multilingual, and image-understanding capabilities for hosts with less memory than the dynamic quantization requires.

Repository: localaiLicense: apache-2.0

muse-glimmer-30b-17gb-dflash
Muse Glimmer's smaller 17 GB K-quant GGUF with the official quantized perception encoder and DFlash drafter. This is the lowest-memory published build that retains image understanding and block-speculative decoding. Flash attention is enabled for the DFlash path.

Repository: localaiLicense: apache-2.0

nemotron-3-embed-1b-q4
Nemotron-3-Embed-1B is NVIDIA's multilingual text embedding model for retrieval, semantic search, and RAG. This compact Q4_K_M GGUF produces 2,048-dimensional normalized embeddings and supports 36 languages. Prefix retrieval queries with `query: ` and documents with `passage: `.

Repository: localaiLicense: openmdw-1.1

Page 1 of many