Architecture & the processing pipeline
A detailed look at what BananaScribe is built from and what actually happens to a recording between import and export. Written so a technical reader can follow the data, not just the marketing version.
01The stack
| Backend | Python, FastAPI, served by uvicorn. |
|---|---|
| Speech-to-text | faster-whisper (CTranslate2 runtime), model chosen per import — from tiny up to large-v3/distil-large-v3. |
| Diarization | pyannote.audio, an open-source speaker-diarization toolkit. |
| Text clean-up | A local language model served by Ollama (punctuation, capitalization, homophone fixes — text only). |
| Window | pywebview — a native window rendering a local web page, no browser install required, no external window-manager dependency. |
| Media handling | ffmpeg/ffprobe — stream probing, demuxing, resampling, denoising, format conversion. |
| System monitoring | psutil (CPU/RAM) and pynvml/nvidia-smi (GPU load + VRAM) for the on-screen gauges. |
| Export | Plain text natively; python-docx for Word; fpdf2 for PDF. |
There's no Docker layer and no container runtime involved — GPU passthrough for a single-machine Windows tool adds overhead without adding anything; faster-whisper gets CUDA/cuDNN access directly through its own Python wheels.
02Hardware: GPU or CPU, auto-detected
BananaScribe is not built against one specific graphics card. Both the
compute device and the compute precision are set to auto:
CTranslate2 picks CUDA if a usable GPU is present and falls back to CPU
otherwise, and separately picks the best precision the hardware
actually supports. Only one speech model is ever kept loaded on the GPU
at a time — switching models unloads the current one first — and GPU
inference is never run in parallel across streams, since sharing one
GPU's memory across concurrent inferences degrades throughput rather
than improving it. The default model offered in the UI (small)
is picked specifically because it stays usable even with no GPU at
all; heavier models are available but optional.
03The file lifecycle, step by step
-
Import
The user picks files (or an MXF folder/ZIP) and import options: model, language, whether to transcribe each audio stream separately, whether to diarize, downsample rate, denoise on/off.
-
Probe
ffprobeidentifies the audio streams in each file. -
Demux
In multi-stream mode,
ffmpegsplits each stream into its own file — each then becomes an independent sub-job with its own progress and checkpoint. -
Prepare
Per stream: downsample (if set) and apply the denoise filter (if enabled), then cut into chunks on silence boundaries.
-
Transcribe
A single sequential GPU worker processes chunks in order: transcribe, append the text to the stream's output, update the checkpoint. See chunking & checkpoints for why this is sequential rather than parallel.
-
Diarize (optional)
If enabled, speaker turns are attributed and merged into one chronological timeline. See diarization.
-
Clean up (optional)
A local LLM pass corrects punctuation, capitalization, and homophones on the finished turn-based transcript. See text clean-up.
-
Export
Available in the viewer (one tab per file/stream) and exportable as text, Word, or PDF — per job, or merged across several finished jobs into one combined script.
04MXF import & multi-track handling
Production sound is frequently delivered as MXF takes with one physical
track per lavalier microphone, rather than one mixed-down track.
Importing a folder or ZIP of MXF takes scans every
take and track first — reading tags and, where a track's character
name isn't embedded in the MXF metadata, falling back to a sound-report
CSV to map file names to characters. Scanning alone does no
transcription work; the user then picks which mics to actually process
(searchable by name) before anything is queued for the GPU. Tracks are
tagged as either an isolated character mic (ISO) or a
pre-mixed track (MIX), which matters for the bleed
suppression and diarization steps below.
05Cross-track bleed suppression
A lavalier mic on one person still picks up some of what nearby people say — "bleed." For isolated per-character tracks from the same take, BananaScribe runs a gating pass that compares tracks moment to moment and attenuates a track when another track in the same take clearly dominates at that instant (i.e. someone else is talking and this person's mic is just catching the spill). The intent is that a line gets transcribed once, on the track of the person who actually said it, rather than duplicated across everyone else's mic too. This step only applies to isolated per-character tracks pulled from the same take — it has nothing to compare against on a single standalone file.
06Diarization
"Diarization" is the step that answers who was speaking, as opposed to transcription, which answers what was said. BananaScribe uses it in two different situations:
- On a single mixed track (e.g. a
MIXtrack, or any non-MXF file with more than one speaker) — pyannote.audio segments the audio by speaker and labels each segment, since the source audio itself doesn't tell you who's who. - Across isolated per-character tracks — the speaker for a segment is already known (it's whichever mic's track it came from); diarization here is about merging every track's turns into one correctly ordered, correctly attributed timeline rather than detecting speakers from scratch.
The first use of a given diarization/transcription model requires a one-time download of that model's publicly published weights — see the network dependency note on the Security page for exactly what that does and doesn't involve.
07Text clean-up pass
After transcription and diarization produce a turn-based transcript (speaker, timestamp, text), an optional pass through a local language model fixes punctuation, capitalization, and words the speech model mis-heard as a similar-sounding alternative. This step only ever touches text that the earlier steps already produced — it has no access to raw audio, and it runs through a local model server that only listens on the machine's own loopback address (see Security).
08Chunking, checkpoints & pause/resume
Each stream is cut into chunks on silence boundaries (a target length, a hard maximum if no silence occurs for a while, and a minimum so a too-short trailing piece gets merged into its neighbor rather than left as its own chunk). Chunks are processed in order by a single sequential GPU worker, with the checkpoint recorded at the chunk level rather than the whole-file level. That makes pause genuinely safe (the current chunk finishes, nothing is left half-written) and resume cheap (skip every chunk already marked done, continue from the first one that isn't) — including resuming automatically if the app itself was restarted mid-job.
GPU inference is deliberately not parallelized across
chunks or streams on one GPU — running several inferences against the
same CUDA context at once causes contention, not speed. Instead,
throughput comes from batching multiple speech segments within
a single chunk into one GPU pass (faster-whisper's
BatchedInferencePipeline), which scales with available
VRAM without any manual multithreading.
09Live microphone mode
Alongside importing existing files, BananaScribe can transcribe straight from a microphone. The browser captures audio locally and watches it for pauses; each time a pause is detected, the in-progress recording segment is finalized and uploaded for transcription while a new segment keeps recording on the same mic stream without interruption. Functionally this reuses the same chunk/checkpoint machinery as file import — a live session is just a stream whose chunks arrive incrementally instead of all at once.
10What's actually parallel
| ffmpeg (demux/resample/denoise/convert) | Yes — CPU-bound, runs across available cores. |
|---|---|
| Cutting into chunks | Yes — CPU-bound. |
| GPU inference across chunks/streams | No — one sequential worker per GPU, by design (see above). |
| GPU inference within one chunk | Yes — via batched inference. |
| Dashboard display | Every stream is shown as its own progress "lane," even though the GPU itself processes them one at a time. |
11Storage layout
Each job gets its own folder under the app's local storage directory, holding the original input, per-stream working audio and chunks, a state file used for progress/pause-resume/crash-recovery, and the growing transcript text. This folder is what persists across restarts (see Storage & retention for deletion behavior) and what gets written during export.