Model Gallery

Discover and install AI models from our curated collection

122 models available
1 repositories
Documentation

Find Your Perfect Model

Filter by Model Type

Browse by Tags

glm-5.3-flash
# GLM-5.3-Flash 👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report. 📍 Use GLM-5.3-Flash API services on Z.ai API Platform. ## Introduction We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute. ## Serve GLM-5.3-Flash Locally ...

Repository: localaiLicense: mit

apodex-1.1-mini-q4
Apodex-1.1-mini is an Apache-2.0 Qwen3.5 mixture-of-experts model for long-horizon research, data analysis, coding, file work, and tool use. It activates about 3B of its 35.95B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the recommended Q4_K_M GGUF and F16 vision projector. An MTP-enabled build and a higher-quality Q8_0 model are available as variants.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q4-mtp
Apodex-1.1-mini with MTP speculative decoding enabled on the recommended Q4_K_M GGUF. The model carries its native MTP head, so it needs no separate draft model. The F16 vision projector supports multimodal prompts.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q8
Apodex-1.1-mini in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

glm-5.3-flash-q4
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the UD-Q4_K_XL GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

glm-5.3-flash-q8
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

llm-jp-4-33b-thinking-q4
LLM-jp-4-33B-thinking is an Apache-2.0 Japanese and English reasoning model from Japan's National Institute of Informatics. Its dense Llama architecture has 33 billion parameters and a 65K-token context window. The model was aligned with supervised fine-tuning and DPO for multi-turn conversation and instruction following. This default entry uses the 20.2 GB Q4_K_M GGUF. The official 66.4 GB BF16 weights are available as a higher-fidelity variant.

Repository: localaiLicense: apache-2.0

llm-jp-4-33b-thinking-bf16
LLM-jp-4-33B-thinking in the official 66.4 GB BF16 GGUF format. This variant preserves the original model precision for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-flash-next-q4
Qwen3.8-Flash-Next is Qwen's 125B-parameter, 6B-active experimental vision-language mixture-of-experts model. It targets agentic coding, reasoning, tool use, and long-context workloads with a native 262K-token context window. This default entry uses Unsloth's UD-Q4_K_XL GGUF and BF16 vision projector. The linked variant uses the higher-quality Q8_0 quantization.

Repository: localaiLicense: other

qwen3.8-flash-next-q8
Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.

Repository: localaiLicense: other

granite-4.2-3b-q4
IBM Granite 4.2 3B is a compact multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q8
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q4
IBM Granite 4.2 8B is a multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q8
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q4
IBM Granite 4.2 30B is the family's flagship multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q8
IBM Granite 4.2 30B in the higher-fidelity Q8_0 GGUF format. It is the family's flagship multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q4
Dirk is a Qwen3.8 27B vision-language model with a concise chat template for agentic coding, reasoning, tool use, and general knowledge tasks. It preserves the model's MTP head for speculative decoding and supports a 262K-token context window. This default entry uses the Q4_K_XL GGUF and F16 vision projector. A higher-quality Q8_K_XL build is available as a variant.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q8
Dirk in the higher-quality Q8_K_XL GGUF format, with MTP speculative decoding and the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

ling-3.0-flash-iq1
Ling-3.0-flash is InclusionAI's MIT-licensed hybrid reasoning MoE model with 124B total parameters and 5.5B active parameters per token. It targets coding, deep research, instruction following, and agentic workflows with a native 256K-token context window. This default entry uses the 36.5 GB AD-IQ1_M GGUF. A higher-quality 44.7 GB AD-IQ2_XS model is available as a variant.

Repository: localaiLicense: mit

ling-3.0-flash-iq2
Ling-3.0-flash in the higher-quality 44.7 GB AD-IQ2_XS GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: mit

carbon-3b-q4
Carbon-3B is Hugging Face's 3B-parameter genomic foundation model for DNA and RNA sequence generation, recovery, variant-effect prediction, and motif-perturbation analysis. It supports 32,768 tokens natively and uses a hybrid tokenizer with 6-mer DNA tokens. This default entry uses the Q4_K_M GGUF. A higher-quality Q8_0 build is available as a variant. Prefix DNA sequences with `` and use uppercase A, C, G, and T characters in groups of six.

Repository: localaiLicense: apache-2.0

Page 1 of many