ClipAI
A short-form video pipeline that ran itself — built, operated and shut down.
Problem
Turning one long recording into several short vertical clips is almost entirely mechanical: find the moments worth keeping, cut them, reframe to nine-by-sixteen, subtitle them accurately, then upload the same file to three platforms with three different APIs and three sets of rules. People pay editors for this, but the work is a pipeline problem wearing a creative costume.
Constraints
- Burned-in subtitles cannot be corrected after publishing, so the transcription had to be right the first time.
- Three platform APIs meant three upload flows, three sets of rate limits and three metadata schemas.
- Fully unattended, with no human between an uploaded recording and a published clip.
- With no editorial oversight in the loop, the entire quality bar had to live inside the selection step.
Decisions
-
Let a language model choose the segments from the transcript, rather than cutting on loudness or fixed intervals.
rejected Audio energy peaks or evenly spaced cuts — cheaper, and no model required.
Volume finds shouting, not substance. Whether a clip stands on its own is a semantic question: did the speaker finish a thought? Reading the transcript meant returning timestamps for self-contained moments rather than for the loudest two minutes, and it was the only part of the pipeline that could not have been written with a threshold.
-
Split orchestration from media work — n8n for scheduling, retries and credentials, Python for the actual cutting.
rejected One Python service doing everything, or pushing the entire pipeline into the workflow tool.
The workflow tool is good at retries, scheduling and secrets, and bad at video; Python is the exact reverse. Splitting along that seam gave the genuinely flaky parts — uploads and API limits — a retry mechanism I did not have to write, and left the media code testable in isolation.
-
Burn subtitles into the video instead of uploading a caption track.
rejected Ship a subtitle file and let each platform style and render it.
Burned-in captions look identical everywhere and cannot be switched off, which is what short-form viewers expect. It also made Whisper's accuracy the single point of failure, which is why word-level timestamps mattered far more than transcription speed.
-
Stop the project rather than add a human review step.
rejected An approval queue where a person signed off each clip before it published.
The whole premise was unattended publishing. The moment a human reviews every clip, this is a normal editing workflow with extra software in the middle, and a worse one than just editing. I would rather record that the premise did not hold than keep the project alive by changing what it was.
The hard part
The pipeline worked and the economics did not, and the failure mode was not a crash but a mediocre clip. Every stage ran unattended exactly as designed, and the output was a technically valid cut of a sentence that meant nothing without the context around it. Automated selection has no point of view, and no amount of prompt tuning gave it one, so the results were consistently publishable and consistently forgettable. The quality problem had a real solution — a reviewer — but that solution dissolved the product into a worse version of an editing tool, so the honest move was to stop.
Outcome
-
6 months
built and operated
-
3
platforms published to unattended
-
0
human steps in the publish path
-
Stopped
because the premise did not hold
Figures and revisions
source: Authored diagram, not a screenshotchecked: redrawn 11 Oct 2026
| plate | source | checked |
|---|---|---|
| FIG. 1 | Authored diagram, not a screenshot | redrawn 11 Oct 2026 |