VoiceStudio is a desktop application for voice cloning, video dubbing, dictation, and long-form audio generation that runs entirely on local hardware. It operates without an account, API key, subscription, or usage meter for the local workflow. The program supports zero-shot voice synthesis from short reference clips, voice design from textual descriptions, multi-speaker audiobook rendering, system-wide dictation, vocal isolation, speaker diarization, and batch processing of audio and video jobs. VoiceStudio includes a model catalogue for managing TTS, ASR, and LLM engines, remote model downloads to enrolled workers, GPU auto-detection across CUDA, MPS, ROCm, and CPU, an AI watermark system using AudioSeal, an MCP server for synthesis and transcription tools, diagnostics with self-checks and logs, and a local-first architecture where network-backed features are explicit opt-ins.

