Model Gallery

Discover and install AI models from our curated collection

79 models available
1 repositories
Documentation

Find Your Perfect Model

Filter by Model Type

Browse by Tags

qwen3.8-27b-uncensored-hauhaucs-aggressive-mtp
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-turbo-fable-cold-fusion-735-882-heretic-uncensored-neo-coder-max-mtp
RELEASE #1 GGUFS [including detailed notes, how to use, benches and much more]: https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (release #1, others pending...) ( repo has 10+ other versions (and 3 branches) noted below that EXCEED the performance of all QWEN 27B models, including fine tunes. ) First, special thanks to Nightmedia for working on the first three stages prior to heretic'ing/post staging and benching everything (3 sections below). A number of my finetunes - both released and non-released - were used here as well as some third parties. Full details will be disclosed upon final release as the project shores up. THREE example generations [snippets] from STAGE1-PART2, STAGE1b-PART2 and STAGE2-rplus2 at the bottom of the page. Release(s) will be GGUFS first (linked here directly) then source code shortly thereafter [now released/open]. Some additional work and/ spawning of new branches from branch(es) below is still going on. NEW: Branch 3 added, see below. COMPLETED AND PENDING RELEASES: ...

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q4-mtp
Apodex-1.1-mini with MTP speculative decoding enabled on the recommended Q4_K_M GGUF. The model carries its native MTP head, so it needs no separate draft model. The F16 vision projector supports multimodal prompts.

Repository: localaiLicense: apache-2.0

glm-5.3-flash-q4
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the UD-Q4_K_XL GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

glm-5.3-flash-q8
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

dirk-qwen3.8-27b-q4
Dirk is a Qwen3.8 27B vision-language model with a concise chat template for agentic coding, reasoning, tool use, and general knowledge tasks. It preserves the model's MTP head for speculative decoding and supports a 262K-token context window. This default entry uses the Q4_K_XL GGUF and F16 vision projector. A higher-quality Q8_K_XL build is available as a variant.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q8
Dirk in the higher-quality Q8_K_XL GGUF format, with MTP speculative decoding and the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-27b-abliterated
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

tiel-coder-35b-a3b-q4-mtp
Tiel-Coder-35B-A3B in Q4_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

qwen3.8-27b-q4-mtp
Qwen3.8-27B with the official Q4_K_M model and Q4_0 MTP draft model. MTP speculative decoding can increase generation speed by proposing multiple tokens for the target model to verify.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-nvfp4-mtp
Qwen3.8-27B in a compact NVFP4 GGUF format with its MTP draft head embedded in the model file. This entry uses the medium tier, which keeps the NVFP4 backbone while using higher-precision output and embedding tensors. MTP speculative decoding proposes multiple tokens for the target model to verify.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-ridge
Qwen3.8-27B Ridge is a 3.69-bit mixed quantization that keeps the Gated-DeltaNet state path at Q8_0 and preserves the embedded MTP head. It reduces the model weights to 12.59 GB while retaining multimodal, reasoning, coding, tool-use, and long-context capabilities.

Repository: localaiLicense: apache-2.0

qwen3.5-9b-defiant-fable-mtp
Qwen3.5 9B Defiant Fable is an Apache-2.0 multimodal fine-tune for reasoning, coding, creative writing, and roleplay. It retains the 256K context window and vision support of Qwen3.5 while reducing refusals. This default entry uses the NEO-imatrix Q4_K_M build with multi-token prediction enabled for faster generation.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-cold-fusion-q4-mtp
Qwen3.8 27B Cold Fusion is an Apache-2.0 multimodal fine-tune for reasoning, coding, creative writing, and roleplay. This entry uses the publisher's NEO-imatrix Q4_K_M GGUF with multi-token prediction enabled. It supports vision through the shared BF16 projector and a native 256K context window.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-cold-fusion-q8-mtp
Qwen3.8 27B Cold Fusion in the higher-quality NEO-imatrix Q8_0 GGUF format. Multi-token prediction is enabled, and the shared BF16 projector provides vision support.

Repository: localaiLicense: apache-2.0

grug-27b-mtp
Grug 27B MTP is the Q4_K_M GGUF build with multi-token prediction enabled for speculative decoding, plus the shared vision projector.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7-mtp-apex
Qwen3.6-35B-A3B Genesis Hermes V7 in the full APEX GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared F16 multimodal projector.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7-mtp-apex-compact
Qwen3.6-35B-A3B Genesis Hermes V7 in the smaller APEX Compact GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared F16 multimodal projector.

Repository: localaiLicense: apache-2.0

qwythos-27b-v1-mtp
Qwythos-27B-v1 MTP is the Q4_K_M build with its native multi-token prediction head enabled for faster speculative decoding. It also includes the shared vision projector and supports tool use and long-context reasoning.

Repository: localaiLicense: apache-2.0

qwythos-9b-v2
Empero AI # Qwythos-9B-v2 — the new and improved Qwythos The next iteration of Qwythos: **all the reasoning of Qwythos-9B, with the looping behavior fixed.** v2 keeps the deep chain-of-thought, the uncensored research posture, and the 1M-token context of its predecessor, and cleans up the rough edges that showed up in real use. - 🔁 **Looping behavior eliminated** — repetition/degeneration under greedy or low-temperature decoding dropped from **6.7% → 0%**. You can serve it *without* leaning on `repetition_penalty` as a band-aid. - 🧠 **Reasoning fully preserved** — MMLU, GSM8K, GPQA, ARC and HumanEval are all held at (or above) the v1 level. This is a *hygiene* upgrade, not a capability regression. - 🧩 **MTP head restored** — the native multi-token-prediction module (dropped in the previous export) is back, so config and weights agree and speculative-decoding setups work. - 🪪 **Cleaner identity** — the model no longer prefaces unrelated answers with its identity; it introduces itself only when you actually ask. - 🔓 **Still intentionally uncensored** for research, cybersecurity, red-teaming, biology, chemistry, pharmacology and clinical work. - 📜 **St ...

Repository: localaiLicense: apache-2.0

tess-4-27b-mtp
Tess-4-27B with its Q4_K_M multi-token prediction draft enabled for speculative decoding. The main model verifies every proposed token, and the entry also includes the shared F16 vision projector.

Repository: localaiLicense: apache-2.0

Page 1 of many