Model Catalog

New & noteworthy speech and voice models curated for LA Studio. Ready to deploy locally or to cloud workflows.

Capabilities:

LA Studio Picks

All Models (15)

Qwen

Qwen3-ASR 0.6B

600M

Lightweight Qwen3-ASR model for offline multilingual speech transcription and audio Q&A, using the CrispASR runtime.

Speech to TextGGUFOffline
Qwen

Qwen3-ASR 1.7B

1.7B

Medium Qwen3-ASR model for offline multilingual speech transcription and audio Q&A, using the CrispASR runtime.

Speech to TextGGUFOffline
Qwen

Qwen3-TTS Base 0.6B

600M

Lightweight Qwen3-TTS base model for fast zero-shot voice cloning from reference audio. Requires 24kHz reference WAV and transcript.

Text to SpeechzhenGGUFVoice-Cloning
Qwen

Qwen3-TTS Base 1.7B

1.7B

High-fidelity Qwen3-TTS base model for zero-shot voice cloning from reference audio. Provides better prosody and quality than 0.6B.

Text to SpeechzhenGGUFVoice-Cloning
Qwen

Qwen3-TTS CustomVoice 1.7B

1.7B

Qwen3-TTS model specialized for preset speakers and style instructions. Optimized for 24kHz high-fidelity output.

Text to SpeechzhenGGUFHigh-Fidelity
Qwen

Qwen3-TTS 1.7B VoiceDesign

1.7B

Qwen3-TTS model specialized in zero-shot voice design from text instructions. Outputs high-fidelity 24kHz audio matching your style descriptions.

Text to SpeechzhenGGUFHigh-Fidelity
contextboxai

Kokoro Vietnamese

82M

Fine-tuned Vietnamese Kokoro TTS packaged for LA Studio with a native Windows CPU runtime, ONNX Runtime inference, and bundled vig2p sidecar phonemization.

Text to SpeechviONNXNative
hexgrad

Kokoro

82M

A compact open-weight TTS model with 82M parameters, installed in LA Studio through a Kokoro backbone GGUF plus a per-voice GGUF pack.

Text to SpeechenesGGUFSmall
k2-fsa

OmniVoice

1B

A quality-first local speech model family for multilingual text to speech, prompt conditioning, and zero-shot voice cloning with reference audio.

Text to SpeechGGUFHigh Quality
microsoft

VibeVoice

500M

Microsoft VibeVoice Realtime 0.5B packaged as CrispASR GGUF for low-latency local text-to-speech with selectable preset voice prompts.

Text to SpeechendeGGUFRealtime Speed
nvidia

Nemotron-3.5 ASR Streaming 0.6B

0.6B

NVIDIA's 0.6B FastConformer RNN-T ASR model, run locally through CrispASR v0.8.6 or later.

Speech to TextGGUFOffline
openai

Whisper.cpp

Multi-size

Native C/C++ inference for OpenAI Whisper models converted to the custom ggml format. Supports multilingual transcription, automatic language detection and translation to English, with optional timestamp, TinyDiarize and Silero VAD workflows in upstream whisper.cpp.

Speech to TextggmlOffline
openbmb

VoxCPM2

2B

Tokenizer-free 2B multilingual TTS model with 48kHz output, natural-language voice design, and zero-shot voice cloning via reference WAV.

Text to SpeecharmyGGUFMultilingual
pnnbao-ump

VieNeu-TTS v2 Turbo

300M

A lightweight, C++ bilingual Vietnamese-English TTS and voice cloning engine optimized for CPU execution.

Text to SpeechvienGGUFBilingual
pnnbao-ump

VieNeu-TTS v3 Turbo

131M

VieNeu-TTS v3 Turbo is a Vietnamese-English TTS model with preset voices, reference-audio cloning, and the v3_native C++ runtime pipeline.

Text to SpeechvienGGUFNative