BananaScribeuser guide

User guide

Day-to-day use: importing a take, choosing what to transcribe, tracking progress, searching across mics, reviewing a transcript, and exporting.

On this page
  1. App layout
  2. Importing plain files
  3. Importing an MXF take
  4. The Dashboard
  5. Searching & bulk actions on mics
  6. The Viewer & player
  7. Exporting a transcript
  8. Combining several jobs into one script
  9. Live microphone transcription

01App layout

The top menu bar has File (import, export, close tab), Edit (copy, select all), View (switch between Dashboard and Viewer), Window (open tabs), Help, and — only for admin accounts in shared mode — Admin (user management, see Administration). Below that sits a tab switcher for the two main screens:

02Importing plain files

File > Import Files... opens a dialog with:

Import Files dialog
Choose Files...Choose Sound Report CSV (fallback)
ModelLanguage
✓ Transcribe streams separately✓ Diarize speakers
Downsample HzDenoise
Transcribe audio streams separately
If a file has more than one audio track (e.g. a multi-channel recording), each track becomes its own independent sub-job with its own progress, rather than being mixed down into one.
Diarize speakers
Runs speaker attribution on the result — see How It Works — diarization for what this does differently on a single mixed track vs. separate tracks.
Sound Report CSV (fallback)
Optional. Maps a file name to a character name using the same CSV format as MXF import, for cases where the file name itself matches a row in that report.
Downsample / Denoise
Optional pre-processing. Leave at defaults unless you have a specific reason to change them.

Click Import and the job appears on the Dashboard immediately and starts processing according to the queue.

03Importing an MXF take

File > Import MXF Folder/ZIP... is the workflow for production sound delivered as MXF takes with one physical track per microphone. It's a two-stage flow, deliberately:

  1. Point at the source

    Choose a folder or a ZIP of MXF takes, and optionally a sound-report CSV (only needed for tracks whose character name isn't already embedded in the MXF's own tags).

  2. Scan

    Clicking Scan reads every take and every track's metadata and lists them all immediately — this step does not transcribe anything yet, it just inventories what's there, tagging each track ISO (an isolated per-character mic) or MIX (a pre-mixed track).

  3. Select & transcribe

    On the Dashboard, the scanned tracks sit in a pending list. Use the mic search (see below) to find specific people/characters, check the ones you actually want transcribed, and click Transcribe Selected. Nothing runs on the GPU for a track until it's explicitly selected this way.

This scan-then-select design exists because a take folder can contain far more mic tracks than anyone needs transcribed for a given purpose — scanning is cheap, transcription is the expensive step, so the app never commits GPU time to a track without an explicit choice.

04The Dashboard

At the top, live gauges show CPU, GPU, and memory load. Below that:

The mic search box filters by name across every job and pending track, with autocomplete suggestions as you type. Two matching modes, both usable at once:

Plain textMatches as a substring — plant matches every "Plant Mic," useful for a whole family of related names.
"Quoted text"Matches exactly — "Hero" matches only a mic literally named Hero, not "Other Hero."
Comma-separatedCombine several terms as an OR — e.g. Anthony, "Hero", plant.

Once filtered, four buttons act only on the matching tracks:

Pause Matching Resume Matching Check Matching Uncheck Matching

This is the fast path for "everyone named X across every take" without clicking through each job individually.

06The Viewer & player

Switch to Viewer (or View > Viewer) to read a transcript. The sidebar lists jobs; opening one adds a tab. Each tab shows the turn-based transcript — speaker, timestamp, text — and, for an audio/video source, a player bar with play/pause and a seek bar whose displayed time accounts for the source recording's own timecode offset, so what's on screen matches what a production's own timecode-based notes say.

07Exporting a transcript

File > Export Transcript... (or the export control on a finished job) offers three formats:

Plain text (.txt) Word (.docx) PDF

Running as the desktop app, export opens a native Save dialog so the file goes exactly where you choose; running in a browser against a shared instance, it downloads through the browser's own download flow instead.

08Combining several jobs into one script

Export All (Combined) on the Dashboard merges every finished job's transcript into a single chronological script — useful when a scene was captured as several separate takes/imports but should read as one continuous document. If a particular mic ended up included by accident (e.g. a boom/mix track that duplicates what isolated mics already captured), click its ✓ button on the job card to exclude that one stream from the combined export; this is reversible at any time and doesn't touch that job's own individual transcript or re-run any processing.

09Live microphone transcription

The Live tab (when enabled) records straight from a connected microphone instead of importing a file. Pick a model and language, press REC, and speak — text appears as pauses are detected and each segment finishes transcribing, with a running chunk count and a live input-level meter. Press the record button again to stop; the session behaves like any other finished job afterward (viewable, exportable).

Running this for a team on a shared machine? See Administration for accounts, sessions, and deployment notes.