beiryu
  • Projects
  • Blog
  • About

© 2026 Khanh Dinh. I build software and write about the journey.

  • GitHub
  • LinkedIn
  • X

AI Podcast Clipper — Long-Form to Short-Form

Paused

An automated pipeline that turns full podcast episodes into publish-ready vertical clips — transcription, moment selection, speaker-tracked cropping, and burned-in captions — running on serverless GPU compute with durable job orchestration.

Jun 2024 – now· Side project· 0 views
GitHub Source
  • Python
  • PyTorch
  • Whisper
  • FFmpeg
  • Modal
  • Inngest
  • Next.js
  • TypeScript
  • Tailwind CSS
  • AWS S3
  • Redis
  • Docker
  • Stripe

Overview#

Every podcast produces hours of material that never reaches an audience beyond the people who already subscribe. Clipping it into short vertical video is the highest-return distribution work available, and it is also the most tedious — which is exactly the shape of problem worth automating.

This platform takes a full episode and returns finished clips: transcribed, selected for the most compelling moments, cropped to follow whoever is speaking, captioned, and exported at the aspect ratio of the destination platform.

The Pipeline#

  • Transcription — Whisper for accurate, word-timed transcripts across accents and languages, which everything downstream depends on
  • Moment selection — the transcript analysed for the segments that stand alone: a complete thought, a strong opening, a payoff at the end. A clip that starts mid-sentence is worthless regardless of how good the content is
  • Speaker detection and tracking — computer vision identifying the active speaker and following them through the frame
  • Dynamic cropping — reframing 16:9 to 9:16 around the tracked speaker, rather than a static centre crop that cuts people out of shot
  • Caption generation — word-timed captions burned in with styling, since most short-form video is watched muted
  • Multi-format export — 9:16, 1:1, and 16:9 from the same source
  • Thumbnail selection from the clip's own frames

Architecture#

The engineering constraint is that video processing is expensive, slow, and fails in the middle. The architecture is built around that rather than around the happy path:

  • Serverless GPU compute through Modal — inference runs on GPUs that exist only for the duration of the job, so cost tracks usage rather than provisioned capacity sitting idle between episodes
  • Durable orchestration through Inngest — the pipeline is a sequence of steps with independent retries, so a transcription failure does not discard completed work and a crash mid-render resumes rather than restarting
  • Containerised processing for FFmpeg and model dependencies, keeping the environment identical between development and production
  • S3 for media with lifecycle policies, since raw episode uploads are large and short-lived
  • Redis caching on the paths that repeat
  • Parallel episode handling, so a user uploading a back catalogue is not processed one at a time

Product#

  • Drag-and-drop upload with support for common video and audio formats
  • Live processing status driven from real pipeline state, with per-step progress — the difference between a user waiting patiently and a user assuming it broke
  • Browser preview and adjustment — every generated clip reviewable and trimmable before export
  • Custom branding — logo placement, caption styling, and colour themes applied across a batch
  • Batch queueing for multiple episodes
  • Subscription billing through Stripe, metered against processing minutes

This project is about production AI infrastructure rather than model work: orchestrating expensive, failure-prone, long-running jobs reliably enough that a user trusts the platform with a two-hour upload.

More projects

TryFinder

Live

B2B Prospect Search & Contact Discovery Platform

  • Side project
  • Next.js
  • React
  • TypeScript

KocerVPN

Live

Bandwidth-Sharing VPN Marketing Site

  • Side project
  • Vite
  • Tailwind CSS
  • Autoprefixer