AI Podcast Clipper — Long-Form to Short-Form
PausedAn automated pipeline that turns full podcast episodes into publish-ready vertical clips — transcription, moment selection, speaker-tracked cropping, and burned-in captions — running on serverless GPU compute with durable job orchestration.
- Python
- PyTorch
- Whisper
- FFmpeg
- Modal
- Inngest
- Next.js
- TypeScript
- Tailwind CSS
- AWS S3
- Redis
- Docker
- Stripe
Overview#
Every podcast produces hours of material that never reaches an audience beyond the people who already subscribe. Clipping it into short vertical video is the highest-return distribution work available, and it is also the most tedious — which is exactly the shape of problem worth automating.
This platform takes a full episode and returns finished clips: transcribed, selected for the most compelling moments, cropped to follow whoever is speaking, captioned, and exported at the aspect ratio of the destination platform.
The Pipeline#
- Transcription — Whisper for accurate, word-timed transcripts across accents and languages, which everything downstream depends on
- Moment selection — the transcript analysed for the segments that stand alone: a complete thought, a strong opening, a payoff at the end. A clip that starts mid-sentence is worthless regardless of how good the content is
- Speaker detection and tracking — computer vision identifying the active speaker and following them through the frame
- Dynamic cropping — reframing 16:9 to 9:16 around the tracked speaker, rather than a static centre crop that cuts people out of shot
- Caption generation — word-timed captions burned in with styling, since most short-form video is watched muted
- Multi-format export — 9:16, 1:1, and 16:9 from the same source
- Thumbnail selection from the clip's own frames
Architecture#
The engineering constraint is that video processing is expensive, slow, and fails in the middle. The architecture is built around that rather than around the happy path:
- Serverless GPU compute through Modal — inference runs on GPUs that exist only for the duration of the job, so cost tracks usage rather than provisioned capacity sitting idle between episodes
- Durable orchestration through Inngest — the pipeline is a sequence of steps with independent retries, so a transcription failure does not discard completed work and a crash mid-render resumes rather than restarting
- Containerised processing for FFmpeg and model dependencies, keeping the environment identical between development and production
- S3 for media with lifecycle policies, since raw episode uploads are large and short-lived
- Redis caching on the paths that repeat
- Parallel episode handling, so a user uploading a back catalogue is not processed one at a time
Product#
- Drag-and-drop upload with support for common video and audio formats
- Live processing status driven from real pipeline state, with per-step progress — the difference between a user waiting patiently and a user assuming it broke
- Browser preview and adjustment — every generated clip reviewable and trimmable before export
- Custom branding — logo placement, caption styling, and colour themes applied across a batch
- Batch queueing for multiple episodes
- Subscription billing through Stripe, metered against processing minutes
This project is about production AI infrastructure rather than model work: orchestrating expensive, failure-prone, long-running jobs reliably enough that a user trusts the platform with a two-hour upload.