I didn't set out to build a photo app.

I just externalized the way I work. The rest followed almost on its own.

The real start

Honestly, if you ask me where KontexVoX comes from, the honest answer isn't "the photos". And it isn't my father's cluttered workshop either, even though everyone assumes that's it.

The real start is the way I think. I've always been obsessed with tidying, organizing, optimizing. Except tidying is time-consuming. And anything time-consuming, we put off. We procrastinate. Me first.

So at some point I got fed up with procrastinating, and I found another way. That's where it all started. Not from a photo. From a problem I had with myself.

Speaking instead of writing

For about a year now, I've stopped writing to the AI. I talk to it. Speech-to-text, all the time.

Why — because writing is slow, and above all because when I speak I unspool my thinking without losing the thread. I have tons of ideas, they go in every direction, and saying them out loud lets me get it all out. Then the AI agent reframes, reorganizes, synthesizes. My raw thinking plus an agent to shape it — that's how I work on just about everything now.

And doing that over and over, one thing jumped out at me: context is the base of everything. An AI agent is only good if you give it context. Otherwise it guesses, and it guesses badly. Sounds simple put that way, but it's really the heart of the story.

Quick note in passing, because it matters to me: when I say "agent", I'm not talking about a chat where you ask a question. An agent acts, it does things for you. It's not the same.

The click

At some point I wanted my assistant to sort my files, my photos, all of it. A real Jarvis, with a digital brain holding my data.

Except I ran into something concrete: it consumes a lot. When you hand a raw image to an agent, it "sees" pixels before it understands what it's looking at. That's heavy, it costs tokens, and tokens are money.

The real problem was right there. And the solution was right in front of me: instead of letting it guess, if I describe the image myself — by voice, so transcribed into text — then I help it understand, and it can do the job. The sorting, the classifying especially. Basically I do half the way for it, and it does what it does best.

That's the founding gesture. Add context, with your voice, so an agent can finally do something useful. All of KontexVoX was born from that intuition.

The spark

And then came the concrete moment. My mother hands me her phone: over 21 GB of photos, storage full, she can't take a single new one.

I offer to move everything to my Mac to free up space. And I tell her something like "it'd be nice to do something with them, it's a shame to have so many photos that'll just sleep". She answers "I don't have the time". And she's right. Sorting and contextualizing thousands of photos by hand — nobody has the time for that.

So I thought: I need a tool that adds context to a photo, and the fastest way is the voice. I speak, and it writes straight into the photo's metadata. I built a first version on Mac, not on a phone.

And then another idea landed: what if we did it like dating apps, with the swipe? That way you enrich your photos on your commute, relaxed, and in the evening you hand it all to your AI agent. The next day you have a beautiful album, a real story, to print and give to your family. Because an album on a shelf tells you something. It keeps a trace.

That's how Classic Mode was born.

The discovery

Capture Mode, though — I didn't imagine it in my head. I discovered it by testing.

When I pitched the idea around me, I started seeing all the other uses: tidying, inventory, the workshop. And since my classic prototype was already running — thrown together in a day, plenty of technical debt but it worked — I went to test in the field.

I took about a hundred photos in a workshop, and I came back to describe them at my desk. Except that's when I realized something: describing a photo after the fact is nothing like telling the story at the moment you take it. Cold, you've already forgotten half. Context wants to be captured hot.

So I concluded it had to be a mode of its own. A mode where you photograph and describe in the same gesture, on the spot. That became Capture Mode.

One idea, many fields

And that's where it's great: once you get the principle, you see it everywhere. It's never more than a single gesture — adding context hot, by voice — applied to different fields.

There's the memory side: keepsakes, family albums. And there's the work side: the inventory an accountant I know has to produce every year for the tax office, the rental condition report, the insurance claim, the architect who today takes photos and annotates with a stylus on the side when voice would be ten times faster, the company audit.

And that — the audit — I know what I'm talking about. My first company, In Safety We Trust, I ran it, and I did safety and security audits. Concretely: take dozens of photos on site, then go back and write a report. No AI back then. Today I know that by contextualizing my photos at the moment I take them, I could have delegated half that report to an agent. I lived the problem before I had the solution.

Same with video. I've done editing, a channel, hours logging footage after a shoot. And the work before handing off to the editor — what is it? Watching the rushes and noting the timecodes you want to keep. Instead of writing it, you say it, by voice, locked to the timecode. And you can even do a first pass with an agent before involving the editor. It's not there to replace the editor — it's there to take the time-consuming part off his plate, so he spends his time on motion design, on what has value.

What it really is

Okay, and I have to be honest, because otherwise you wonder "but what's new about this?". The answer is: nothing's new. And that's exactly right.

KontexVoX doesn't do magic. It's a faster workflow, that's it. Speaking is fluid, writing is slow — that's all. It's really the same principle as the dictaphone that transcribes a meeting to give you the minutes. Except here it's applied to your photos, and soon your videos.

And I have to be clear about who does what. KontexVoX produces the raw material: your enriched photos, with your description transcribed into text, engraved in the EXIF metadata. No audio file, ever — "Vox" is for voice, but the result is text. The final deliverable — the inventory spreadsheet, the report, the pre-edit — an AI agent generates that afterward, from this material and a good prompt. And the prompts, we give them to you, for free, on the blog.

That's it. It's not revolutionary. It's just a simpler, faster way to do something you already do badly, or don't do at all because it bores you. And honestly, that's exactly why it works.

And now

Start with your own photos, or go see the prompts we share to turn all this into albums, spreadsheets, reports.

Join the waitlist See the prompts (free)