FAQ & glossary
Quick answers, and plain-language definitions for terms used throughout this documentation.
01Frequently asked
- Does transcription require an internet connection?
- No, once the models in use are already downloaded and cached locally. The first use of a given model requires a one-time download — see Security — the one network dependency.
- Where does my audio/video go?
- Nowhere but the local working folder on the machine (or local server) running BananaScribe. See Security — processing is local.
- How accurate is the transcript?
- Depends on the model chosen (heavier models are generally more accurate but slower), audio quality, and whether diarization/bleed suppression apply cleanly to the source. The optional text clean-up pass improves punctuation/capitalization/homophones, not word-level accuracy.
- Can I pause a long job and resume it later?
- Yes — progress is checkpointed at the chunk level, so pausing is safe and resuming picks up exactly where it left off, even across an app restart. See How It Works — chunking, checkpoints & pause/resume.
- What's the difference between importing files and importing MXF?
- Plain file import is for any individual audio/video file. MXF import is purpose-built for production takes with one track per microphone: it scans every take/track first, then lets you pick exactly which mics to actually transcribe. See User Guide — importing an MXF take.
- Why did a mic's dialogue show up twice, on two different tracks?
- Likely cross-track "bleed" — one person's mic picking up another nearby. Bleed suppression (see the glossary) addresses this for isolated per-character MXF tracks; it doesn't apply to a single standalone file with nothing to compare it against.
- Can several people use one installation at the same time?
- Yes, in shared mode — see Administration. Note that GPU transcription itself is still processed one job at a time in queue order, even with multiple users submitting work.
- What export formats are available?
- Plain text, Word (.docx), and PDF — per job, or merged across several finished jobs into one combined script. See User Guide — exporting.
02Glossary
- MXF
- Material eXchange Format — a professional audio/video container format commonly used for production sound and camera recordings, often carrying several separate tracks (e.g. one per lavalier mic) in a single take.
- Take
- One recorded unit of production audio/video (e.g. one take of a scene), potentially containing multiple tracks.
- Track / stream
- One channel of audio within a file or take — e.g. a single person's lavalier mic.
- ISO track
- An isolated per-character/per-mic track, as opposed to a pre-mixed one. See MXF import in the User Guide.
- MIX track
- A pre-mixed track combining multiple sources (e.g. a boom mix), as opposed to an isolated single-mic track.
- Diarization
- Determining who is speaking and when, as distinct from transcription, which determines what is said. See How It Works — diarization.
- Bleed / bleed suppression
- "Bleed" is one person's microphone also faintly picking up what a nearby person says. Bleed suppression is the processing step that attenuates a track at moments when another track in the same take clearly dominates, so a line is attributed once rather than duplicated. See How It Works.
- Turn
- One continuous block of speech by one speaker in the exported transcript — a new turn starts on a speaker change or after a long enough pause, so a transcript reads as a series of turns rather than one undifferentiated block of text.
- Chunk
- An internal processing unit — a stream is cut into chunks on silence boundaries, each transcribed and checkpointed individually. See How It Works.
- Sound report CSV
- A spreadsheet export from a production's sound department mapping file/track names to character names, used as a fallback when that mapping isn't already embedded in a file's own metadata.
- Job
- One unit of work in the Dashboard's queue — an imported file, MXF track, or live recording, with its own status and progress.
- Combined export
- Merging several finished jobs' transcripts into a single chronological script. See User Guide — combining several jobs.