LinkedIn post

Minutes from audio files – fully automated, fully local and sovereign

Andreas Pelzner2 min read
← Back to the blog overview

Translated from the German original. Read the original in German.

For months I have been working intensively with AI-driven workflows built on OpenClaw in my day-to-day work as a consultant. One topic that keeps coming up: meeting minutes. A time sink, tedious, but indispensable.

So I developed a complete minute-taking stack – my AI assistant and I, built on OpenClaw. And I can tell you: it works. With good quality. And sovereign.

🎙️ The recording

Either the whole meeting is recorded, or I dictate into the voice memo app on my iPhone afterwards. I send the recording via Telegram to my personal AI assistant – and from then on everything runs automatically.

What happens next:

1. The assistant receives the audio and starts the workflow. No button, no form. I say “Minutes, please”, add the metadata in the recording – location, participants, occasion, date and format if needed, in a few keywords – done.

2. The audio stays on my Mac. Transcription runs locally with faster-whisper large-v3 on the CPU (M4 Pro). Not a single byte goes to the cloud. Data protection without compromise.

3. Diarisation via pyannote separates the speakers automatically – who spoke when. In the minutes the different speakers appear as Participant 1 to x. At the end you have to replace these with the actual names by hand, especially for verbatim minutes.

4. Post-processing by a local LLM (qwen3:30b, running locally via Ollama) turns the raw transcript into clean minutes – either verbatim minutes or summary minutes.

5. Rendering in Word. The finished minutes land as a .docx directly in my Nextcloud, based on ready-made document templates. Done.

How was it developed?

Entirely through vibe coding. I give the instructions, the assistant develops. The architecture guidelines for all my OpenClaw workflows are set out in a cookbook – that is the guard rail.

I don’t write the code myself. I describe what I need, and the system builds it. That is now my productive everyday reality.

What is impressive:
The quality still surprises me. Local models can do this. The transcript is clean, the structure is right, and speaker attribution works even with several people. And: ~20 minutes’ processing time for 15 minutes of audio – perfectly acceptable for minutes that would otherwise have taken me an hour to write.

My conclusion:

AI workflows don’t have to run in the cloud. They don’t have to be expensive. But they do have to be thought through – from input to filing. And they have to fit the way you work.
This is not a product. This is my personal stack. And this is what the future looks like: composable AI with vibe coding.

If this interests you, I am happy to compare notes.

Andreas PelznerManaging Director, WE SUCCESS Consulting GmbH · LinkedIn
Book an initial consultation