Model Gallery

Discover and install AI models from our curated collection

145 models available
1 repositories
Documentation

Find Your Perfect Model

Filter by Model Type

Browse by Tags

s1-mini-q4
S1-mini by Superwhisper is a 0.6B English text normalizer for raw speech transcripts. It removes fillers and false starts, restores punctuation and capitalization, and formats spoken numbers, dates, currency, and email addresses as written text. This default entry uses the publisher's 462 MB Q4_K_M GGUF and greedy decoding. A higher-fidelity F16 model is available as a variant. Prefix the transcript with the styling, structure, and context control line documented on the model page.

Repository: localaiLicense: s1-mini-license

s1-mini-f16
S1-mini by Superwhisper in the publisher's 1.4 GB F16 GGUF format. This variant preserves full model fidelity for hosts with enough memory.

Repository: localaiLicense: s1-mini-license

nemo-parakeet-tdt-0.6b
NVIDIA NeMo Parakeet TDT 0.6B v3 is an automatic speech recognition (ASR) model from NVIDIA's NeMo toolkit. Parakeet models are state-of-the-art ASR models trained on large-scale English audio data.

Repository: localaiLicense: cc-by-4.0

moonshine-tiny
Moonshine Tiny is a lightweight speech-to-text model optimized for fast transcription. It is designed for efficient on-device ASR with high accuracy relative to its size.

Repository: localaiLicense: apache-2.0

whisperx-tiny
WhisperX Tiny is a fast and accurate speech recognition model with speaker diarization capabilities. Built on OpenAI's Whisper with additional features for alignment and speaker segmentation.

Repository: localaiLicense: mit

omnilingual-0.3b-ctc-q8-sherpa
Omnilingual ASR CTC 300M (int8) is a multilingual automatic speech recognition model supporting 1,600+ languages. Based on Meta's omniASR_CTC_300M architecture (Wav2Vec2 with CTC head), quantized to int8 for efficient inference. Uses the sherpa-onnx backend with ONNX Runtime.

Repository: localaiLicense: apache-2.0

streaming-zipformer-en-sherpa
Streaming English ASR: sherpa-onnx zipformer transducer (int8, chunk-16 left-128). Low-latency real-time transcription with endpoint detection via sherpa-onnx's online recognizer. English-only; for multilingual offline ASR see omnilingual-0.3b-ctc-q8-sherpa.

Repository: localaiLicense: apache-2.0

vibevoice-cpp
VibeVoice Realtime 0.5B (C++ / GGML, Q8_0) - native C++ port of Microsoft VibeVoice via the vibevoice-cpp backend. 24kHz mono TTS with a selectable precomputed voice prompt. Default voice prompt: en-Carter_man. This realtime variant does not accept raw Voice Library reference WAVs.

Repository: localaiLicense: mit

vibevoice-cpp-asr
VibeVoice ASR 7B (C++ / GGML, Q4_K) - long-form speech-to-text with speaker diarization. Returns per-speaker JSON segments with start/end timestamps. English-only. ~10 GB download.

Repository: localaiLicense: mit

qwen3-asr-1.7b
Qwen3-ASR is an automatic speech recognition model supporting multiple languages and batch inference.

Repository: localaiLicense: apache-2.0

qwen3-asr-0.6b
Qwen3-ASR is an automatic speech recognition model supporting multiple languages and batch inference.

Repository: localaiLicense: apache-2.0

lfm2.5-audio-1.5b-asr
LFM2.5-Audio-1.5B in ASR mode. System prompt `Perform ASR.` is prepended; output is capitalised and punctuated. Wire this entry as a transcription model on the /v1/audio/transcriptions endpoint.

Repository: localaiLicense: LFM-Open-License-v1.0

whisper-1
Port of OpenAI's Whisper model in C/C++

Repository: localaiLicense: mit

whisper-large-q5_0
Port of OpenAI's Whisper model in C/C++

Repository: localaiLicense: mit

whisper-medium
Port of OpenAI's Whisper model in C/C++

Repository: localaiLicense: mit

whisper-small-en-q5_1
Port of OpenAI's Whisper model in C/C++

Repository: localaiLicense: mit

whisper-small-en
Port of OpenAI's Whisper model in C/C++

Repository: localaiLicense: mit

whisper-large
Port of OpenAI's Whisper model in C/C++

Repository: localaiLicense: mit

whisper-large-turbo-q8_0
Port of OpenAI's Whisper model in C/C++

Repository: localaiLicense: mit

parakeet-cpp-tdt_ctc-110m
Hybrid TDT+CTC FastConformer, 110M. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime_eou_120m-v1
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M. Use with streaming transcription. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

Page 1 of many