← Back to projects

AI Clipping

Local-first pipeline that turns long videos into edited short-form clips. Completed through V13.3, with subject-aware framing planned for V14.

Status

Completed

Year

2026

Core Tech

Python, FFmpeg, faster-whisper

01 / The Problem

Automated video cropping tools often select the wrong moments, make jarring cuts, and fail to capture the pacing of human-edited clips.

02 / The Solution

A pipeline that evaluates moments, framing, and pacing to generate short vertical clips from long-form video. It is currently on version 13.3, with continuous iteration based on comparing the output quality against previous versions.

03 / How It Works

The core pipeline follows a strictly decoupled architecture:

Long-form Video Input
faster-whisper Transcription
Scene/Moment Analysis
OpenCV & YOLO11 Framing
FFmpeg Clip Generation

04 / Engineering Decisions

Local-First Architecture

Built to run locally instead of relying on expensive cloud GPU APIs, utilizing faster-whisper and OpenCV for efficient local processing.

Iterative Evaluation

Instead of trusting a single LLM prompt to edit video, I iterate by comparing each generated clip to the previous version and keeping only what looks objectively better.

05 / Technology

  • Python
  • FFmpeg
  • faster-whisper
  • OpenCV
  • YOLO11

07 / Learnings

  • Subject-aware framing is incredibly complex due to erratic movement. Planning to integrate YOLO11 tracking in V14 to fix center-framing drift.

Ready to see it in action?