Discover and install AI models from our curated collection
Repository: localaiLicense: apache-2.0
LLM-jp-4-33B-thinking is an Apache-2.0 Japanese and English reasoning model from Japan's National Institute of Informatics. Its dense Llama architecture has 33 billion parameters and a 65K-token context window. The model was aligned with supervised fine-tuning and DPO for multi-turn conversation and instruction following. This default entry uses the 20.2 GB Q4_K_M GGUF. The official 66.4 GB BF16 weights are available as a higher-fidelity variant.
Links
Tags
LLM-jp-4-33B-thinking in the official 66.4 GB BF16 GGUF format. This variant preserves the original model precision for hosts with enough memory.
Links
Tags
IBM Granite 4.2 3B is a compact multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.
Links
Tags
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.
Links
Tags
Repository: localaiLicense: apache-2.0
Carbon-3B is Hugging Face's 3B-parameter genomic foundation model for DNA and RNA sequence generation, recovery, variant-effect prediction, and motif-perturbation analysis. It supports 32,768 tokens natively and uses a hybrid tokenizer with 6-mer DNA tokens. This default entry uses the Q4_K_M GGUF. A higher-quality Q8_0 build is available as a variant. Prefix DNA sequences with `` and use uppercase A, C, G, and T characters in groups of six.
Links
Tags
Carbon-3B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.
Links
Tags
Repository: localaiLicense: mit
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It activates about 3B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.
Links
Tags
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.
Links
Tags
Tiel-Coder-35B-A3B is a 35B-parameter mixture-of-experts model for coding, reasoning, tool use, and vision tasks. This default entry uses the Q4_K_XL GGUF and BF16 vision projector.
Links
Tags
Tiel-Coder-35B-A3B in Q4_K_XL format with MTP speculative decoding and a BF16 vision projector.
Links
Tags
Tiel-Coder-35B-A3B in the higher-quality Q8_K_XL GGUF format, with the BF16 vision projector for multimodal prompts.
Links
Tags
Repository: localaiLicense: openmdw-1.1
NVIDIA Nemotron 3.5 Lightning is a text-only hybrid Mamba-2, attention, and mixture-of-experts model with 30B total parameters and 3B active parameters. It targets reasoning, coding, tool use, multilingual chat, and long-context agent workflows, with a context window of up to one million tokens. This entry uses the official Q4_K_M GGUF. Automatic variant selection can choose the smaller NVFP4 build or the higher-quality Q8_0 build when it fits.
Links
Tags
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official NVFP4 GGUF format. This is the smallest linked build and retains the model's reasoning, coding, tool-use, multilingual, and long-context capabilities.
Links
Tags
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official high-quality Q8_0 GGUF format for hosts with enough memory.
Links
Tags
Repository: localaiLicense: other
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the Q4_K_M GGUF quantization.
Links
Tags
Repository: localaiLicense: other
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the near-lossless Q8_0 GGUF quantization.
Links
Tags
Repository: localaiLicense: apache-2.0
# Parable-Granite-4.1-3B-Claude-Fable-5 Granite 4.1 3B fine-tuned on genuine Claude Fable 5 and GPT-5.5 agent traces (planning, tool use, reasoning from real agent sessions). Agent-flavored small model: terminal workflows, idiomatic code fixes, explanations. v2 recipe: completion-masked SFT, replay mix, seed-averaged weights. Published corpus and eval harness.
Links
Tags
Repository: localaiLicense: apache-2.0
Qwen3.6-35B-A3B Uncensored Genesis Hermes V6 is LuffyTheFox's multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry installs the Q8_0 GGUF together with its F16 multimodal projector for llama.cpp. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior. License: Apache-2.0.
Links
Tags
Repository: localaiLicense: apache-2.0
Qwen3.6-35B-A3B Genesis Hermes V7 is LuffyTheFox's Apache-2.0 multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry's own payload uses the model card's recommended APEX GGUF and the shared F16 multimodal projector. Automatic variant selection may instead choose Compact APEX, an MTP-enabled APEX build, or Q8_K_P based on serving features and available memory. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior.
Links
Tags
Repository: localaiLicense: apache-2.0
Qwen3.6-35B-A3B Genesis Hermes V7 in the smaller APEX Compact GGUF format, with the shared F16 multimodal projector. This build preserves the model's multimodal, reasoning, coding, and agentic capabilities for hosts with less memory than the recommended full APEX build.
Links
Tags
Qwen3.6-35B-A3B Genesis Hermes V7 in the full APEX GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared F16 multimodal projector.
Links
Tags