Discover and install AI models from our curated collection
Repository: localaiLicense: s1-mini-license

S1-mini by Superwhisper is a 0.6B English text normalizer for raw speech transcripts. It removes fillers and false starts, restores punctuation and capitalization, and formats spoken numbers, dates, currency, and email addresses as written text. This default entry uses the publisher's 462 MB Q4_K_M GGUF and greedy decoding. A higher-fidelity F16 model is available as a variant. Prefix the transcript with the styling, structure, and context control line documented on the model page.
Links
Tags

S1-mini by Superwhisper in the publisher's 1.4 GB F16 GGUF format. This variant preserves full model fidelity for hosts with enough memory.
Links
Tags
NVIDIA NeMo Parakeet TDT 0.6B v3 is an automatic speech recognition (ASR) model from NVIDIA's NeMo toolkit. Parakeet models are state-of-the-art ASR models trained on large-scale English audio data.
Links
Tags
Moonshine Tiny is a lightweight speech-to-text model optimized for fast transcription. It is designed for efficient on-device ASR with high accuracy relative to its size.
Links
Tags
WhisperX Tiny is a fast and accurate speech recognition model with speaker diarization capabilities. Built on OpenAI's Whisper with additional features for alignment and speaker segmentation.
Links
Tags
Repository: localaiLicense: apache-2.0
Omnilingual ASR CTC 300M (int8) is a multilingual automatic speech recognition model supporting 1,600+ languages. Based on Meta's omniASR_CTC_300M architecture (Wav2Vec2 with CTC head), quantized to int8 for efficient inference. Uses the sherpa-onnx backend with ONNX Runtime.
Links
Tags
Repository: localaiLicense: apache-2.0
Streaming English ASR: sherpa-onnx zipformer transducer (int8, chunk-16 left-128). Low-latency real-time transcription with endpoint detection via sherpa-onnx's online recognizer. English-only; for multilingual offline ASR see omnilingual-0.3b-ctc-q8-sherpa.
Links
Tags
Repository: localaiLicense: mit

VibeVoice Realtime 0.5B (C++ / GGML, Q8_0) - native C++ port of Microsoft VibeVoice via the vibevoice-cpp backend. 24kHz mono TTS with a selectable precomputed voice prompt. Default voice prompt: en-Carter_man. This realtime variant does not accept raw Voice Library reference WAVs.
Links
Tags

VibeVoice ASR 7B (C++ / GGML, Q4_K) - long-form speech-to-text with speaker diarization. Returns per-speaker JSON segments with start/end timestamps. English-only. ~10 GB download.
Links
Tags
Qwen3-ASR is an automatic speech recognition model supporting multiple languages and batch inference.
Links
Tags
Qwen3-ASR is an automatic speech recognition model supporting multiple languages and batch inference.
Links
Tags
LFM2.5-Audio-1.5B in ASR mode. System prompt `Perform ASR.` is prepended; output is capitalised and punctuated. Wire this entry as a transcription model on the /v1/audio/transcriptions endpoint.
Links
Tags
Port of OpenAI's Whisper model in C/C++
Links
Tags
Port of OpenAI's Whisper model in C/C++
Links
Tags
Port of OpenAI's Whisper model in C/C++
Links
Tags
Port of OpenAI's Whisper model in C/C++
Links
Tags
Port of OpenAI's Whisper model in C/C++
Links
Tags
Port of OpenAI's Whisper model in C/C++
Links
Tags
Port of OpenAI's Whisper model in C/C++
Links
Tags
Hybrid TDT+CTC FastConformer, 110M. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.
Links
Tags
Repository: localaiLicense: cc-by-4.0
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M. Use with streaming transcription. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.
Links
Tags