Models and your data

Models and your data

Substrike uses a few machine learning models. This page says what each one is, where it runs, where it came from, what we know and do not know about how it was trained, and what is stored or sent. Current as of 6 October 2026, checked against the app's code.

The short version

  • Substrike is not free of machine learning. The speech recogniser is a machine learning model, and the optional AI cleanup is a small language model.
  • All of them run on your PC. Your clips, audio, transcripts and voice fingerprints are not uploaded for processing.
  • We did not train any of these models. Each comes from someone else, under the licence listed below. Where the makers have not published their training data, we cannot tell you what is in it.
  • A speech model trained on better documented, openly licensed data is something we are exploring. There is no promise and no date.

Speech recogniser: Whisper

This is the model that turns the speech in your clip into words.

What it is
Whisper, an encoder decoder model from OpenAI. It listens to audio and writes out text one piece at a time, so it can occasionally mishear, and it can very occasionally write words that were not said. Substrike flags the words it was unsure of for you to check.
What it does in Substrike
Produces the transcript and word timings that the captions are built from. It is run through whisper.cpp, a separate program on your PC.
Where it runs
On your PC. No audio is sent to a speech service.
Optional?
You need one speech model. A small one, tiny.en (78 MB), comes inside the installer so you can try Substrike straight away. The recommended one, large-v3-turbo (about 1,625 MB), is a one time download you start yourself. Other sizes are listed in the app.
Where it comes from
OpenAI made Whisper. The files we download are the ggml conversions published by the whisper.cpp project on Hugging Face (repository ggerganov/whisper.cpp), pinned to one exact version and checked against a fixed checksum after download.
Licence
MIT, for the model and for whisper.cpp.
Training data: known
OpenAI has said Whisper was trained on a very large amount of audio and matching text collected from the internet.
Training data: not known
The data is not published or itemised, so we cannot tell you which creators, languages or sources are in it, or whether everyone in it agreed. We have not audited it and we do not claim it was ethically sourced.

Word timing: wav2vec2

A second, smaller model that lines each word up with the sound more precisely, once the words are known.

What it is
wav2vec2-base-960h, a speech model from Meta. It is used here to place words in time, not to decide what was said.
Where it runs
On your PC.
Optional?
Yes. Without it you still get captions, with looser word timing. It is a one time download of about 378 MB.
Where it comes from
Meta (Facebook AI) trained the base model. The file we download is an ONNX conversion made by the Xenova project (Transformers.js) and hosted on Hugging Face, pinned and checksum checked.
Licence
The base model is Apache-2.0. The conversion repository does not state a licence of its own.
Training data
Known: the base model was trained on LibriSpeech, about 960 hours of read English audiobooks. Not known: we have not checked the rights position of those recordings beyond what the publishers state.

AI cleanup: Qwen2.5 1.5B Instruct

This is the one optional language model in Substrike. If you do not want a language model on your PC, do not download it.

What it is
Qwen2.5-1.5B-Instruct, a small language model that generates text, in a compressed form (Q4_K_M).
What it does in Substrike
It looks at words the speech model was unsure of and suggests what the word most likely was, such as a misheard name. It is shown one flagged word in its sentence at a time and never rewrites the whole transcript. On English clips, a suggestion it rates at least 90 percent sure is applied automatically and recorded as an AI cleanup edit, and in the editor all of those automatic fixes can be undone in one step. Every other suggestion waits for your review. On clips in other languages nothing is applied automatically. It can also help with caption translation and suggested titles, which you review the same way. We have not shown that it improves accuracy.
Where it runs
On your PC. Substrike starts a small local program, llama-server, which talks to Substrike over your own computer's loopback address (127.0.0.1) and is shut down when the run ends. Nothing is sent to an AI service.
Optional?
Yes. It is not installed until you choose to download it (about 1,117 MB). Everything else works without it. If you install it, the editor and the batch queue use it on unsure words, as described above.
Where it comes from
Made by the Qwen team at Alibaba Cloud. We download the GGUF file from Hugging Face (repository Qwen/Qwen2.5-1.5B-Instruct-GGUF), pinned to one exact version and checked against a fixed checksum. It runs through llama.cpp.
Licence
Apache-2.0, for the model. llama.cpp is MIT.
Training data: known
The Qwen team describes it as pretrained on a very large amount of text and then tuned to follow instructions.
Training data: not known
The training text is not published or itemised. We cannot tell you what it contains or whether the authors agreed to its use, and we have not audited it.

Speaker detection: pyannote and TitaNet

Two small models that work out when the speaker changes and whose voice is whose, so each person can have their own caption colour and name.

What they are
A speaker change detector (pyannote segmentation 3.0) and a speaker voice recogniser (NVIDIA NeMo TitaNet-small), run through sherpa-onnx. They do not produce words.
Where they run
On your PC.
Optional?
They come inside the installer, so there is nothing to download. Speaker colours and names depend on them, and captions work without anyone naming a speaker.
Where they come from
pyannote segmentation 3.0 is from the pyannote project and TitaNet-small is from NVIDIA. We bundle ONNX exports published by the k2-fsa (sherpa-onnx) project on GitHub.
Licence
pyannote segmentation 3.0 is MIT. TitaNet-small is CC-BY-4.0, credited to NVIDIA in our third-party notices. sherpa-onnx is Apache-2.0.
Training data
Each is trained on published speech datasets named on its own model card, and some of those have their own terms. We have not audited them, and we do not say they are ethically sourced.

Two small helpers

These are not speech or language models, but they are machine learning models, so they are listed here.

Face finder (YuNet)
Finds faces in your video frames so captions can stay clear of them. It is built into the app, runs on your PC, and is MIT licensed, from the OpenCV model zoo. Its training data is not described in our files, so we cannot say more.
Laughter finder (experimental)
A small audio tagging model that helps find moments where people laugh. It is bundled, runs on your PC, and is Apache-2.0. Its training data is not described in our files.

What leaves your PC, and when

Substrike has no cloud account and does not upload your clips, audio, transcripts or voice fingerprints for processing. The only connections are these, and each one needs something from you or happens at start up. The privacy notice has the full table.

Model downloads
Only when you press Download. The request goes to huggingface.co, which can see your IP address. Nothing about you or your clips is sent.
Update check
A short check about 2.5 seconds after the app opens, and when you press Check for updates. It asks GitHub for a small file saying whether a newer version exists. You can block it in your firewall.
Licence check
Once when you enter a licence key, and when you move it to another PC. It sends the key and a one way hash of your Windows installation identifier to our licence service.
Discord posting
Only if you paste a Discord webhook address and choose to post. The video or message you chose goes to Discord.
Twitch clip import
Only if you connect a Twitch account.
Crash report
Only if you press Send report and then Send, after reading the preview. It never includes video, audio, transcripts, voiceprints, clip names, paths or your licence key.

Once the models are on your PC, if you never check for updates and do not use Discord or Twitch, the only connection Substrike makes is its start up update check.

What is stored on your PC, and how to delete it

Models
In Substrike's data folder, normally %APPDATA%\com.substrike.app\models. Delete a model file there to remove it. The small tiny.en model is copied back from the installer if it goes missing.
Transcripts and edits
Autosaves of your work on each video (transcript, edits, style choices), any project file you save, and the batch queue list. They stay on your PC. Delete the autosave or project files to remove them.
Voice fingerprints during processing
To tell speakers apart, Substrike makes a numeric voice fingerprint for each speaker in a clip. It is kept with that clip's transcript, in autosaves, project files and the batch queue, and is not uploaded. It is not audio.
Saved voices (voiceprints you choose to keep)
Off until you say yes. The first time you name a speaker, Substrike asks whether to remember that voice. If you agree, it stores the name you typed and up to four fingerprints per person in one small file in Substrike's data folder. It never learns a voice on its own. In Australia a voiceprint used to identify someone can count as sensitive biometric information, so only save people who are happy for you to.
Deleting saved voices
In the Transcript tab, open Saved voices, then delete one person or choose Delete all voices. The same list is in Settings under captions. Copies already inside an autosave or a project file you saved are separate, so delete those files too. Uninstalling Substrike and deleting its data folder removes everything.
Settings and history
Recent clips, chosen model, theme, export choices and style presets are kept on your PC. If you enter a Discord webhook address it is stored there in plain form, so treat it like a password.

What we are looking at next

We are exploring a speech model trained on better documented, openly licensed data. This is an exploration only. We are not promising it, and we have no date for it. If it happens, this page will say what changed.

If something here looks wrong or out of date, please tell us through support.