Video production · Shotcut · OpenShot · Lightworks · AI dataset assets

I build the code-driven video pipelines that turn raw footage into clean training data and crisp research cuts.

A decade across Shotcut, OpenShot, Lightworks, ffmpeg and OpenTimelineIO. Most of my editorial is scripts, because timelines you can diff are timelines you can ship.

Editorial timeline mock showing three clips with cross-dissolves, dynamic lower-thirds and an audio track
Editorial timeline mock — rendered live from render_posters.py
Shotcut / MLT OpenShot Lightworks ffmpeg PySceneDetect OpenTimelineIO CMX-3600 EDL .cube LUT Whisper EBU R128 OpenCV Python

Selected projects

Five projects that together cover the full video-for-AI pipeline — from raw footage ingest through deduplicated datasets, editorial cuts and captioned delivery.

Scene grid of the synthetic dataset: color bars, moving gradient, and test pattern sampled at key timestamps

Dataset curation · ffmpeg · PySceneDetect · aHash

AI dataset video curation

The pipeline I run before any footage reaches a model. Scene detection via PySceneDetect, 8×8 average-hash deduplication, `ffprobe`-driven normalization to 1080p30 / yuv420p / AAC 48 kHz, and a diffable manifest.csv that reviewers can see in a PR.

Deterministic, no-GPU curation · drops exact + near-duplicate scenes · hook point for CLIP-based semantic dedup

Read the case study →
Editorial timeline with three video clips, cross-dissolves and a waveform track

Shotcut · MLT XML · dynamicText

Shotcut MLT project generator

Python emits a valid Shotcut .mlt with a pinned 1080p30 profile, three producers, a main playlist, a tractor with audio + video tracks, and dynamicText lower-thirds. Opens and edits just like a hand-built project — but lives in git.

MLT XML · reviewable cuts · parameter-sweep friendly · cross-machine deterministic

Read the case study →

OpenShot · OTIO · Interchange

OpenShot OSP + OTIO builder

Same source of truth, two extra formats: OpenShot's JSON .osp and an OpenTimelineIO .otio that round-trips through Resolve, Hiero and otiotool. No more "it won't open on my machine".

Case study →

Lightworks · CMX-3600 EDL · 33pt LUT

Lightworks EDL export

Finishing hand-off: CMX-3600 .edl, a producer-friendly shotlist.tsv, and a generated 33pt teal/amber .cube LUT so the look survives the round-trip to Lightworks or Resolve.

Case study →

Whisper · SRT · ffmpeg subtitles

Auto-captioning + burn-in

Whisper (when WHISPER=1) or a committed transcript feed the same SRT into ffmpeg's subtitles filter. Output honors a 10 % action-safe margin so captions never clip on mobile 9:16 crops.

Case study →

Code you can actually run

Every asset on this page is regenerated by GitHub Actions on every push — never pre-baked. Clone the repo, run one script, get the same MP4.

15-second editorial cut — three curated clips, xfade transitions, dynamic lower-thirds, 1080p30 H.264 · rendered live from render_demo.py.

Dataset curation — PySceneDetect + aHash

python scripts/python/curate_dataset.py

Scene-detects and deduplicates raw footage, writes a diffable manifest.csv, normalizes survivors to 1080p30 / yuv420p / AAC 48 kHz.

Shotcut project generator

python scripts/python/build_shotcut_mlt.py

Emits scripts/projects/portfolio.mlt with profile, producers, playlist, tractor and dynamicText lower-thirds — opens in Shotcut 22.06+.

OpenShot + OTIO builder

python scripts/python/build_openshot_osp.py

Writes a JSON .osp for OpenShot and an OpenTimelineIO .otio for Resolve / Hiero / custom tools.

Lightworks EDL + LUT

python scripts/python/export_lightworks_edl.py

CMX-3600 EDL, producer-friendly shotlist TSV, and a 33pt teal/amber .cube LUT applicable via ffmpeg lut3d=.

Auto-captioning + burn-in

python scripts/python/auto_caption.py

Whisper when WHISPER=1, otherwise a committed SRT. Burned in via subtitles filter with a 10 % action-safe bottom margin.

Rebuild the whole pipeline

python scripts/python/run_all.py

One-shot: synthetic clips → curation → MLT / OSP / OTIO / EDL → captions → final cut → posters. The same script CI runs.

Audio waveform and RGB parade scopes of the first curated clip
Audio waveform (1 kHz tone) and RGB parade — column-mean of the clip's middle frame. Both scopes regenerated in CI.

Stack

Editorial NLEs

Shotcut (MLT XML), OpenShot (JSON project), Lightworks (CMX-3600 EDL). Projects are generated, not clicked — every cut is reviewable in a PR.

Finishing & color

3D LUTs in .cube, applied via ffmpeg lut3d or any NLE. Teal / amber editorial look. EBU R128 loudness targets across the deliverables.

Interchange

OpenTimelineIO across tools, DNxHR / ProRes proxies for offline finishing, H.264 / H.265 for delivery. One source of truth drives every emitter.

Dataset curation

PySceneDetect for cuts, OpenCV + aHash for dedup, ffprobe + ffmpeg for normalization. Diffable manifest.csv and audit log of rejects.

Captions & accessibility

Whisper transcription when available, otherwise committed SRT. Burned-in ASS styling, 10 % action-safe margin. WCAG 2.2 1.2.2 in mind.

Code & CI

Python (numpy, matplotlib, opencv-python, scenedetect, opentimelineio), ffmpeg. Git, GitHub Actions. Every artefact is regenerated on push.

About

I'm a senior video production professional spanning editorial, motion, and technical supervision — most recently focused on video assets for AI research programs. I care about reproducible pipelines, honest delivery specs (EBU R128 is necessary but not sufficient), accessible captions done right, and timelines that still render the same MP4 a year from now.

Open to remote and contract engagements in video editorial, post supervision, or research-facing media engineering. The repository linked below is the living portfolio companion to my CV.