Draft preview: not yet available for purchase

Glossary

Glossary

The technical words on this site, explained once. Every dotted link in the technical sections lands here.

Terms, A to Z

9:16, vertical, Fill and Fit
The portrait shape of TikTok, Reels and Shorts. Fill crops the widescreen clip to fill the frame, Fit shrinks it with bars, and auto track crops towards the motion.
AAC, Opus
Audio compression formats. AAC plays everywhere; Opus is newer and smaller. Substrike offers whichever your PC can encode.
Answer key, reference transcript
The human checked transcript a model’s output is scored against. Ours were corrected word by word by the developer, starting from engine drafts.
ASR (automatic speech recognition)
Software that turns spoken audio into text. Substrike’s ASR is whisper.cpp running on your processor.
AudioSet tagger
A small sound model trained on Google’s AudioSet collection of labelled sounds. Substrike uses one to score laughter for the experimental Find laughter option.
Beam search
A way of decoding speech that keeps several likely readings in play (Substrike uses five) instead of committing to one word at a time. Slower, and more accurate.
Bitrate, VBV, size target
Bitrate is how much data per second a video uses. To fit a Discord limit, Substrike picks a bitrate, caps it with a VBV buffer so it never spikes, then measures the file and re-encodes lower if needed.
Burn in
Drawing the captions into the video frames so they show on any player. The opposite of a separate subtitle file.
Checksum, SHA-256
A fingerprint of a file. Every model Substrike downloads is checked against a fixed SHA-256 fingerprint, so a corrupted or tampered download is refused.
Code signing, SmartScreen
A publisher certificate that tells Windows who made an installer. Until Substrike’s is in place, Windows SmartScreen may warn about an unrecognised publisher.
Confidence, unsure word
The speech model’s own estimate of how likely a word is right. Words below a threshold are treated as unsure and shown in Review words.
CPU, GPU
The processor and the graphics card. Speech recognition uses the CPU; only export can use the GPU.
Diarisation, speaker detection
Working out who is speaking when, and grouping the same voice together. Substrike uses pyannote segmentation 3.0 (where a voice changes) and NeMo TitaNet-small (whose voice it is).
DTW word timestamps
Dynamic time warping, the method whisper.cpp uses to estimate when each word starts and ends. Align timing then refines those times.
Ed25519, activation token
The signature scheme our licence server uses. Your PC keeps a signed token and checks it offline, so the licence keeps working without the server.
FFmpeg, ffprobe
The open source video toolkit that reads your clip’s details, pulls out the audio, burns captions in and encodes the export. Substrike runs it as a separate program.
Forced alignment, Align timing
Given the audio and the words, work out exactly when each word starts and ends. It is why captions land on the beat.
GGUF, ONNX, ONNX Runtime
File formats and a runtime for running AI models on your PC without any cloud. GGUF is used by llama.cpp for the cleanup model; the timing and speaker models are ONNX files, run by ONNX Runtime.
H.264, HEVC, AV1
Video compression formats. H.264 plays everywhere and is the default. HEVC and AV1 are newer and smaller; support depends on your graphics card.
Hallucination, invented words
When a speech model writes words nobody said, usually over music, silence or noise. It is why Listen closer only runs on a stretch you select.
Hardware encoder, NVENC, AMF, Quick Sync, Media Foundation
Video encoding chips on NVIDIA (NVENC), AMD (AMF) and Intel (Quick Sync) graphics. Media Foundation is the encoder built into Windows, used when there is no graphics card encoder.
Karaoke captions, word by word captions
Captions where the current word lights up as it is spoken. They need accurate word timing.
LGPL, GPL, copyleft
Open source licences that let anyone use, study and change the software, and require its source to be available. Substrike’s FFmpeg build is LGPL-2.1 with its source published; Substrike ships no GPL code.
libx264, software encoder
A program that encodes H.264 video on the processor. Our FFmpeg build leaves it out on purpose, for licence and patent reasons, and uses hardware encoders or Windows’ own encoder instead.
llama.cpp, llama-server
An open program that runs language models on an ordinary PC. Substrike starts it on your own computer (address 127.0.0.1), so the cleanup model never talks to the internet.
LLM, cleanup model, Qwen2.5-1.5B-Instruct
A small language model (1.1 GB) that runs on your PC. Substrike shows it only the unsure words and their context and asks for a fix or a flag; it also translates and writes post text.
Merchant of record, Paddle
The company that legally sells you the licence, takes payment, works out tax for your country and issues the receipt. For Substrike that is Paddle.
Multi track audio
OBS can record your mic, the party and the game as separate audio tracks in one file. Substrike can caption only the voice tracks and keep the game track in the export.
Pooled
Errors added up across all clips and divided by all words, rather than averaging each clip’s percentage. It stops a short clip counting as much as a long one.
Prompt, Names and jargon
Text given to the speech model before it listens, which nudges it towards those spellings. It is how Substrike learns your gamertags.
Quantisation, precision, Q4_K_M, int8, bf16
How many bits each of a model’s numbers is stored with. Fewer bits (int8, Q4_K_M) means a smaller download and less memory for a small loss in quality; bf16 is a common higher precision format.
Replay buffer (OBS)
An OBS feature that keeps the last minute or so in memory and saves it when you press a hotkey. Where most Substrike clips come from.
RTF (real time factor)
Processing time divided by audio length. 0.5 means a 2 minute clip takes 1 minute.
Safe zone
The parts of a vertical video that TikTok, Instagram and YouTube cover with buttons and text. Substrike draws them so you can keep captions clear.
sherpa-onnx
An open toolkit that runs speech and sound models on your PC. Substrike uses it for speaker detection and laughter tagging.
SRT, ASS
Subtitle file formats. SRT is the plain one Premiere, Resolve, CapCut and YouTube import. ASS carries styling and animation; Substrike compiles your caption style to ASS and FFmpeg burns it in.
Tauri, WebView2
The framework the desktop app is built with: a Rust program showing its interface through WebView2, the web engine already in Windows.
VAD (voice activity detection)
Deciding whether anyone is talking at all. We tested using it to gate transcription and dropped it because it cut real words on busy clips.
Voiceprint, embedding
A short list of numbers that sums up how a voice sounds. It is not audio and cannot be played back. Substrike keeps them on your PC and treats them as sensitive information.
Watch folder
A folder Substrike keeps an eye on. New files saved into it are captioned automatically.
wav2vec2
A speech model from Meta that Substrike uses only for timing, not for recognising words. The download is 378 MB.
Webhook (Discord)
A private link you create in a Discord channel’s settings that lets a program post there. Anyone with it can post, so treat it like a password; Substrike keeps yours on your PC.
WER (word error rate)
The standard accuracy score for speech recognition. Every wrong, missing or extra word counts as one error, divided by the number of words; 19.3% WER means about 81 words in 100 were right.
Whisper, whisper.cpp, large-v3-turbo
Whisper is OpenAI’s open speech recognition model family, and whisper.cpp runs it on ordinary PCs. large-v3-turbo is the size Substrike recommends: 1.6 GB, and it made the fewest errors of the local models we tested on gaming chat.
Windows installation identifier
A value Windows creates when it is installed (the MachineGuid). At activation Substrike sends only a one way hash of it plus an app secret, never the value itself, so your licence can count your PCs.

Missing a word? Ask support and we will add it.