Good to know
The facts to check before you buy: what Substrike needs, how accurate it is today, and what is being finished for 1.0.
Describes the current pre-release build as of 23 September 2026. Re-checked before every release.
At a glance
- Windows
- 64-bit Windows 10 or 11, with the Microsoft Edge WebView2 Runtime (already on current Windows). No macOS or Linux version.
- Language
- English today. Spanish, Portuguese, French, German and Italian are in beta in the app, and more languages are coming soon. Beta accuracy has not been measured: those languages keep the speech model's own word timing, and their cleanup fixes wait for you instead of applying themselves. Captions can also be translated into other languages after transcription, using the same local model.
- Processing
- Everything runs on your PC. No footage is uploaded to Substrike or a speech service for processing; files leave your PC only for Discord posting, support attachments, downloads or a future licence check. Speech recognition uses the CPU at low priority, and a working NVIDIA, AMD or Intel GPU encoder speeds up export.
- Downloads
- About 190 MB installer, then about 3.1 GB of one-time model downloads for the recommended setup: 1,625 MB speech model, 378 MB word-timing model, 1,117 MB AI cleanup model. Each is checked against a fixed checksum and resumes if interrupted.
- Internet
- For the model downloads, update checks and optional Discord posting. Captioning and export work offline after that.
- Posting
- Substrike posts to Discord through your own webhook: a session digest and a compiled reel. For TikTok, YouTube and Instagram it makes the file and you upload it.
More detail, including smaller speech models, is on the download page.
Accuracy, honestly
About 70% of words right on three-player proximity chat over game audio. Measured on 2 hand-checked clips, 207 words, about 70 seconds of audio, using today’s recommended speech model, large-v3-turbo. The other 30% needed a fix: a wrong word, a missing word or an extra one.
This is deliberately the hard case, and it is a small sample. More clips are being checked now, and this page will be updated with what they show.
How it was measured
- Clip one: about 45 seconds of three players talking over each other with game music and yelling, 156 reference words, transcribed from scratch by hand. 66.7% word accuracy. Scored as a simple match rate instead, the same clip gives the 68.4% figure this site quoted before: an illustrative internal stress test on this one hard clip, not a general accuracy rate. Same clip, different scoring, not a change.
- Clip two: 25 seconds of three-player proximity chat, 51 words, corrected in the app and marked as checked. 78.4% word accuracy.
- Who checked them: the Substrike developer, on his own recordings, with one microphone setup. No independent reviewer yet.
- Word timing: after re-timing, the typical word starts about 64 ms from where it was spoken, measured on clip one (124 words). About a quarter of words were still more than 0.25 seconds off.
- Model choice (same one hard clip, not a rate): on clip one, the recommended model scored a 68.4% match rate against 63.8% for the largest model and 52% to 66% for three cloud speech services. On a single clip, gaps that small are not settled.
- Not yet measured: clean single-speaker speech, other microphones, accents and games, the smaller speech models on game audio, and processing time or memory use across PCs. The bar before quoting a headline figure is 30 clips across 3 or more games.
What the safety net catches, and what it does not
- The AI cleanup flags words the speech model was unsure of. A clip with more than three unsure words, or one the cleanup could not check, is held for you. One to three still exports, with the count shown on its row.
- Confident mistakes and missing words are not flagged, so a clip that exports on its own has not been proven correct. Watch it through before posting anything important.
- Cleanup is an optional 1.1 GB download. Without it, a batch holds every clip for review by default rather than exporting it as if it were checked. You can switch that off with “Export even if AI cleanup can’t run”.
- On clip one, turning cleanup on did not change the match rate. Its job is flagging, not guaranteed fixes.
- Names, slang and game jargon are the most common misses. The names and jargon field helps the model spell the names you give it.
- Speakers are grouped automatically, and game audio or overlapping speech can produce wrong or extra speakers. Speaker detection was tuned on a handful of the developer’s proximity-chat clips.
- Without the 378 MB timing model, captions still work but word timing is less precise. The app offers a one-click download.
Being worked on before launch
- Code signing. The test build is not signed with a Windows publisher certificate, so SmartScreen may warn about an unrecognised publisher. Updates are already checked against Substrike’s own update key.
- Clean-PC testing. The installer is being tested on several fresh Windows PCs, including an overnight watch-folder run. Minimum RAM and typical processing times will be published once measured.
- Visual C++ runtime. Included. Substrike ships its own copy of the Microsoft Visual C++ runtime, so a fresh Windows install needs nothing extra. We are still confirming this on clean PCs before launch.
- Licensing and the trial. A one-time online key check, then offline use, and a 7-day trial from your first successful export.
- FFmpeg source page. Substrike includes FFmpeg (LGPL-2.1). Licences for it and the other components are on the open source page and in the app (Home, About & licences). The source offer for our own FFmpeg build is published next to each release.
- First-run setup. The guided downloads and the 15-second preview on your own clip are new and are getting more testing across PCs.
Experimental features
These work, but sit outside the core workflow and have had little validation.
- Highlights finder. Ranks moments in a long recording by how loud and excited the audio gets. It does not know what a kill, a win or a joke is.
- Death detector. Looks for phrases in the transcript. Not validated across games or audio setups.
- Caption translation. Uses the local cleanup model to translate captions into other languages. Quality not measured, so check translated captions before posting.
