
Video mode is coming to KontexVoX: your voice transcribed in real time, synced to the timecode. Every word tied to the image. The editor — or the AI — will know exactly what to do. The app launches with the photo modes; sign up to hear about Video mode.
You come back from a shoot with 2 hours of footage.
Automatic trimming removes the silences. But it doesn't know which shot is good, which moment is strong, which take is best. Only a human knows that. And that human is you.
One gesture to keep, one to cut — sorting goes fast. Editors have their own tricks for that. But there's a moment when you want to say: "At 3 minutes 12 there's a good shot. I want this motion-design effect right here. I want us to go in this direction."
Saying it means not typing it. And it's tied to the exact moment in the video.
You review your video. You watch it in full. Along the way, you comment out loud.
"That's the wide shot of the entrance, you can see the building well."
"This one's good, the light is great, we keep it for the intro."
"That one's a miss, we cut."
"Right at this moment, she smiles — that's the highlight of the film."
That voice will be transcribed in real time, synced to the video's timecode. Every word, every phrase, every idea, tied to an image.
At the end, you'll have a video enriched with a complete timecoded transcription.
That transcription is:
A complete review of your footage, anchored in time. An editing brief, embedded in the file. A script usable by a human editor OR an AI agent.
Note: Video mode is in development — it will arrive after the photo modes launch, and it will be a paid PRO mode, offered as an add-on separate from the photo modes (see pricing). This page describes reviewing existing videos; Capture Mode will also let you film live while speaking, for instance filming each page of a binder while saying what it contains. A different use, closer to video inventory. More on the blog.
You review your footage and comment as you watch. At the end, your selection is done. No need to rewatch everything. Send it all to your AI agent, which pre-edits from your timecoded instructions.
Instead of writing a separate editing doc that drifts out of sync with the footage, you review the video and speak. The editor receives the video with the instructions embedded, at the right timecode. No more guesswork.
The timecoded transcript works with AI agents (Claude Code, Codex, Gemini). The AI reads the instructions, identifies the shots to keep, and generates a preliminary cut or an EDL for Premiere/DaVinci.