How to Transcribe an Audio or Video File
File transcription is ScribeForge's core workflow, and it's available on every plan, free included. Here's what actually happens when you drop in a file.
Supported formats
The Transcribe screen accepts audio (MP3, WAV, FLAC, OGG, M4A and more) and video (MP4, MKV, AVI, WebM and more) — over 40 formats in total. For video files, ScribeForge extracts the audio track automatically; you never need to convert anything yourself first. See the dedicated audio to text and video to text pages for format-specific detail.
What happens after you click Transcribe
- ScribeForge extracts the audio track (for video) and prepares it for the model.
- Your chosen Whisper model processes the audio locally, in chunks, generating timestamped segments as it goes.
- If speaker diarization is enabled (Pro), each segment is also labelled with a speaker.
- The finished transcript is saved to your local library automatically — nothing is uploaded anywhere during this process.
Editing after transcription
Whisper is accurate but not perfect — click any segment in the transcript panel to correct a mishead name or technical term directly. If a term comes up often (a product name, a person's name), add it to the custom dictionary so future transcripts get it right automatically.
Free plan limits
Free-plan file transcription is capped at 15 minutes per file and 5 hours per month, using the tiny/base/small Whisper models. Upgrading removes both caps and unlocks every model. Full detail on the free transcription software page.