Clips Kitty converts long videos and live streams into vertical clips formatted for Shorts, Reels, and TikToks. The processing happens on your Windows computer, so footage stays local unless you choose to publish a clip to your own YouTube channel. The application finds notable moments, crops to 9:16 with the speaker centered, adds synchronized captions, and writes titles, descriptions, and hashtags. It supports 19 languages with translation, subtitling, and dubbing, and includes an editor for correcting automated results.
The software evaluates each moment using five signals: transcript content, audio loudness, burst density for laughter or applause, visual activity, and on-screen reactions. It builds creator profiles over time to improve clip selection and titling, while keeping all data on your machine. A non-destructive editor stores trims, cuts, mutes, volume changes, fades, speed adjustments, hook text, music, and watermarks as operations applied at render time. An AI edit chat lets you describe changes in plain language, and the model proposes modifications that validated code applies to the project.
Clips Kitty requires Windows 10 or 11 with 16 GB of RAM, and an NVIDIA graphics card is recommended. It uses CUDA for detection and transcription, hardware video encoding through NVENC, AMF, or QSV, and runs the language model on GPU through Ollama. The model manager lets you download, remove, and switch AI models without using a terminal. Publishing to YouTube is optional and off by default, using your own Google API key. The project is free, open-source, and places no limit on the number of clips you create.
| Fuses transcript, audio, visual and structural signals | Score and extract high-value moments from video |
| Tracks active speaker with pose and audio analysis | Keep subjects framed in vertical clip output |
| Assigns steady per-camera crops for podcast footage | Produce clean vertical cuts from multi-camera recordings |
| Renders word-synced captions with editable transcript lines | Correct recognition errors before burning subtitles |
| Applies natural-language edit requests via validated operations | Adjust clips conversationally without destructive changes |
| Generates titles, descriptions and hashtags locally | Prepare publishing metadata without cloud services |
| Builds persistent per-creator knowledge across uploads | Improve title accuracy and recognise recurring content |
| Translates captions across nineteen languages with glossary | Protect brand names during multilingual publishing |
| Synthesises local dubbed audio tracks per language | Produce spoken translations without external TTS |
| Exports horizontal clips, highlights or trimmed streams | Repurpose long-form content beyond vertical format |