Clips Kitty


Clips Kitty — Free Download. AI Clip generator

Turn long streams and videos into ready-to-post Shorts, Reels, and TikToks, entirely on your own Windows machine. Clips Kitty is a free AI clip generator and video editor for creators: paste a YouTube, Twitch, or Kick link and it finds the best moments, crops them to 9:16 with the speaker kept centred, burns in word-synced captions, and writes titles, descriptions, and hashtags. No cloud AI. No subscription. No per-clip fees. Your footage never leaves your computer, unless you ask it to publish a clip to your own YouTube channel.

5.0(1 ratings)
File size: 0.794 MB
The latest version of Clips Kitty is: 1.1.4
Operating system: Windows
Languages: English
Price: $0.00 USD (Open Source (AGPL-3.0))
  • Multimodal clip detection. This core function analyses video content across multiple signal channels simultaneously. It scores every candidate moment on a scale from zero to one hundred by fusing five distinct data streams: the spoken words from the transcript, audio excitement measured through loudness spikes and burst density patterns, visual activity detected via scene cuts and motion analysis, on-screen reactions such as laughter or applause shapes, and hook-to-payoff structural strength. Each signal is normalised within the specific video being processed, meaning a quiet podcast and a loud gaming stream both produce meaningful relative peaks. Every clip that clears the quality threshold is preserved without any arbitrary numerical cap on output volume.
  • Speaker-aware face tracking. This function ensures the subject remains properly framed throughout each clip. It employs YOLOv8 pose detection to keep the primary speaker centred within the vertical frame. In footage containing multiple people, the camera follows whoever is actively speaking, determined by TalkNet active-speaker detection which analyses face and audio data together rather than relying on movement alone. When the active speaker changes, the framing cuts directly to the new speaker instead of panning across the scene, mimicking the decision pattern of a human editor. The cropping process only adjusts framing boundaries and never stretches or distorts the image.
  • Podcast mode. Designed specifically for multi-camera podcast recordings, this function assigns each camera shot its own steady crop focused on one individual. This approach means cuts land cleanly on a face without any panning motion between speakers. Within any single shot, the subject is selected by analysing mouth motion patterns, with a fallback to the most prominent face when mouth data is inconclusive. The result is a stable, professional-looking vertical format output from horizontally recorded multi-camera source material.
  • Editable burned-in captions. This function generates word-synced captions that are permanently rendered into the video during export. Caption styling is fully configurable, covering colour, font size, screen position, maximum words per line, letter casing, or complete deactivation. Because transcription errors are inevitable with any speech recognition system, the application provides line-by-line editing capability so mistakes can be corrected before the export process runs. Word-level timing data ensures captions appear on the correct syllable and clips begin at genuine sentence boundaries rather than mid-word.
  • AI edit chat. This function provides a conversational interface for modifying clips after the initial generation. The operator describes what needs changing in plain natural language, such as requesting a clip be five seconds longer or pointing out that a caption reads a wrong word. The language model interprets the request and proposes an edit operation, which is then applied by validated code rather than directly by the model. This separation between proposal and execution prevents invalid operations from corrupting the project. All edits remain non-destructive and are stored as operations applied at render time.
  • AI titles, descriptions, and hashtags. This function automatically generates publishing metadata for each clip. The local language model analyses the transcript and event timeline to produce a title, a description, and relevant hashtags suited to the clip content. Every generated field is presented as editable text before export, so the operator retains full control over the final metadata. The generation process runs entirely on the local machine without transmitting content to external services.
  • Creator Profiles. This function builds a persistent knowledge base for each creator across their uploaded videos. The system tracks recurring topics, series names, frequent collaborators, running jokes, and ongoing storylines. This accumulated context improves the accuracy of generated titles and helps the system recognise callbacks to earlier content. The profiling is deliberately conservative: a catchphrase must repeat across multiple videos before being recorded, knowledge that stops appearing in new content becomes dormant, and every score contribution derived from profile data is additive, capped, and can be disabled entirely by the operator.
  • Multilingual publishing. This function handles translation and caption generation across nineteen languages. Translated captions are displayed for review before any text is written to a file or burned into video frames. This review step means a poor translation can be corrected while still editable. A glossary system protects channel names, sponsor references, and community-specific phrases from being translated, preserving brand identity and in-group language across language versions.
  • AI dubbing. This function produces spoken audio tracks in target languages using local text-to-speech synthesis. Voice options can be auditioned per language before committing to a dub. The dubbing process runs on the local machine as part of the multilingual pipeline, and the resulting audio is integrated into the exported video file alongside the translated captions.
  • Long-form export. This function provides an alternative output path for horizontal 16:9 content using the same analysis engine that powers vertical clip generation. Output options include horizontal clips, X/Twitter-length cuts, a best-of highlight reel compiling top moments, or the complete stream with dead air segments removed. This path is opt-in and operates alongside the standard vertical clip workflow rather than replacing it.
  • Watermark and branding profiles. This function stores per-creator branding defaults that apply automatically during export. Configuration covers watermark image placement and appearance along with other branding elements. Once a profile is established for a creator, subsequent exports use those settings without requiring manual reapplication each time.
  • Model manager. This function provides an in-application interface for managing the local language model. Operators can download new models, remove existing ones, and switch between available models using graphical controls with progress bars. No terminal or command-line interaction is required for any model management task. This design keeps the AI component swappable as better models become available.
  • Publish to YouTube. This function enables direct upload from within the editor. Configuration options include title, description, tags, thumbnail selection, playlist assignment, audience settings, and visibility status. The operator can choose to upload immediately or schedule the publish for a later time. This feature is optional, disabled by default, and requires the operator to supply their own Google API key for authentication.
  • In-app feedback. This function allows operators to submit bug reports directly from the application. Diagnostic information is collected automatically and attached to the report, reducing the back-and-forth typically required to reproduce an issue. No account creation or sign-in is needed to send feedback.
  • Accessible UI. This function ensures the application remains usable across different operator needs. Keyboard focus navigation is supported throughout the interface. A reduced-motion mode respects system preferences for less animation. Font family and text size are adjustable to accommodate visual requirements.

Clips Kitty was created by ColinGPT9 and is developed as an independent project hosted on GitHub. The program has been available since 2024 and is written primarily in Python, with performance-critical components accelerated through CUDA and hardware video encoding APIs including NVENC, AMF, and QSV. The development approach prioritises local execution, keeping all processing on the operator's own hardware rather than relying on cloud services.

Alternatives to Clips Kitty:

Clip Squeezer — Free Download. Video converter

Clip Squeezer

Clip Squeezer is a compact desktop application for four routine video operations: reducing file size, converting between popular container formats, cutting unwanted segments from the beginning or end, and separating audio tracks from video files.
Price: Free   Size: 115 MB   Version: 0.3.0   OS: Windows, Mac OS
Clipforge Studio — Free Download. Local AI video clipping

Clipforge Studio

ClipForge Studio is a Windows desktop application that analyzes long-form videos with local AI to identify engaging moments and produce short clips for TikTok, Reels, and Shorts.
Price: $10   Size: 944.1 MB   Version: 1.0.0   OS: Windows