You streamed for 3 hours with OBS. The replay is there, the video is 12 GB. You know there are highlights in it, but finding them in 3 hours of video is monk's work. By hand, you scrub, you listen, you note timecodes, you forget half of them. With dozens of potential moments, you give up.

AI clipping apps exist for this. They analyze your replay, detect audio peaks, keywords, rhythm changes, and generate vertical clips with animated subtitles. It works, and it's fast. But they hear patterns, not meaning. A laugh can be nervous or delighted. "No way, I can't believe it" can be frustration or wonder. The app can't tell the difference.

An AI agent that reads your timecoded transcript can go further. Clipping tools already use AI agents, and they're good. But they work with what they have: a raw transcript. They lack the context only you have. Why is this moment funny and not awkward? Why is it the move of the century and not an ordinary failure? KontexVoX is the way to add that information by voice, while you watch the replay, so the AI — whether inside Captions, Opus Clip, or your local agent — finally has everything it needs to understand.

Prerequisites

Before you start, you need:

1. A timecoded transcript of your video, in SRT or VTT format. Generate it from your replay with Whisper, CapCut, or any automatic transcription tool.
2. An AI agent with file-system access: Claude Code, Codex CLI, or Gemini. Not a web chatbot.
3. The transcript exported to your computer.
Three hours of transcript is thousands of lines. A web chatbot can't handle that context length. And it doesn't generate a usable clips file. You need an agent that reads the file locally.

The timecoded video mode (coming)

KontexVoX today enriches photos: you look at an image, you speak, the description is engraved into the EXIF. For video, the principle is the same but the medium changes: you watch the replay, you comment at the moment it happens, and the app generates a timecoded file (SRT or JSON) linked to the video — not a per-image EXIF description. This mode is in development. In the meantime, you can transcribe your replay with any tool (Whisper, CapCut, Premiere) and run the prompt below. The result is the same.

What the apps hear vs what you see

Take three excerpts from a 3-hour live stream transcript:

01:12:34,200 → "And he tells me — no way, I can't believe it — he tells me he's never climbed. Ever. I'm crying." 01:47:05,800 → "WAIT. Wait wait wait. You see that? You see the move? That's the move of the century." 02:33:18,400 → "No but guys, this is… you're the best. Seriously. Thank you."

The AI in Captions or Opus Clip detects the peaks: the exclamation at 01:47, the laugh at 01:12, the emotion at 02:33. It generates three clips. That's fine. But it doesn't know why the move is "of the century" — it doesn't see the two failed attempts before, the arc of tension. It doesn't know "I'm crying" is laughter, not tears. It doesn't know the moment at 02:33 is a sincere thank-you after 3 hours of streaming, not a polite sign-off.

That context is yours. KontexVoX is the way to hand it to the AI. You watch the replay, you say out loud: "this is the move of the century, he failed it twice before, keep 30 seconds before the success." The AI — whether inside Captions, Opus Clip, or your local agent — receives not a raw transcript, but a transcript enriched with your human context. That's the missing information that separates a decent clip from a really good one.

The prompt

Read the timecoded transcript file (SRT or VTT) of this 3-hour live stream. Analyze the linguistic patterns that indicate highlights: laughter, emphasis ("that's good", "incredible", "wait", "no way"), tone changes, exclamations. For each identified moment, extract the start timecode and estimate the end timecode (usually 30-60 seconds after the peak). Generate a list of clips in Markdown with: clip number, start timecode, end timecode, transcript of the passage, and a brief note on why it's a highlight.

What the AI produces with this prompt

A list of clips in Markdown. Each clip has:

Its start and end timecode

The transcript passage word for word

An explanatory note: why it's a highlight (laughter, emphasis, exclamation, tone change)

You get a working document, not an edit. From this list, you go into Premiere, DaVinci or CapCut, extract the clips at the given timecodes, and publish. The scouting work, which would have taken 3 hours of viewing, is done in minutes.

The precision isn't perfect: the AI identifies candidates, but the final judgment is yours. Out of 20 proposed clips, you keep 10 or 12. It's still faster than rewatching everything.

The KontexVoX advantage (when video mode ships)

Clipping apps charge a premium subscription for volume: Captions is $25 a month, Opus Clip $19. For a creator streaming 3 hours a week, that's $300 a year. KontexVoX will offer the timecoded video mode for a few euros a month — a light subscription, not a pro one.

And it's not either/or. The idea is to combine both: you enrich your replay with KontexVoX (voice annotations that give the context and the why of each moment), then you hand it all to your favorite clipping app. Captions or Opus Clip receives not a raw transcript, but a transcript annotated by a human. The generated clips are better, better targeted, with the meaning the app alone couldn't guess. KontexVoX before Captions, not against Captions.

Recommended tool: Gemini

Gemini suits long-transcript analysis and pattern identification. On 3 hours of transcript, you need a model that handles long context while catching the nuances. Gemini handles long inputs well.

Frequently asked questions

Does the AI pick the right moments?

It picks candidates. Linguistic patterns ("wait", "no way", "incredible", laughter) are good indicators, but not infallible. Out of 20 proposed clips, some will be false positives. The final judgment is yours — but you judge in 30 seconds per clip, not 3 hours of rewatching.

Does it work on a replay without annotations?

Yes, with automatic transcription. Whisper or CapCut transcribe your replay's audio track. The AI spots emotional moments in the transcript text. It's less precise than if you'd annotated the moments by voice ("that's good" at the exact moment), but it already works well on exclamations and laughter.

Can I target a specific format (Shorts, Reels, TikTok)?

Yes. Add to the prompt: "Each clip must last between 15 and 60 seconds for a Short/Reel format." The AI adjusts the end timecode accordingly. For longer clips (YouTube excerpts), ask for 1 to 3 minutes.

Why pay for KontexVoX if I already have a clipping app?

It's designed to work together. You enrich with KontexVoX (voice annotations, human context), then hand it all to Captions or Opus Clip, which produce the clips. The clipping app receives an annotated transcript, not a raw one. The clips are better. The timecoded video mode is in development, the subscription will be light — a few euros a month. In the meantime, the prompt above already works with any transcript.