Overview
VoiceScribe turns speech into text entirely on the local machine. No network calls, no API keys, no per-request cost, and nothing leaves the device.
Feed it audio — a recording, a microphone stream, or an imported file — and it returns transcribed text, either in one pass or segment by segment as the speaker talks.
- Fully offline — the model ships inside your packaged build
- Streaming or one-shot recognition
- Multi-language, with optional translation to English
- Selectable model size, from Tiny to Large, including quantised variants
- Editor tool to download and manage models
- Tunable decoding: threads, beam size, temperature, prompts, token limits
Requirements
| Requirement | Details | |
|---|---|---|
| Unreal Engine | 4.27 and 5.8 (verified clean on both) | |
| Modules | VoiceScribe (Runtime), VoiceScribeEditor (Editor) | |
| CPU | x64 with AVX. Recognition is CPU-bound. | |
| Disk | Model dependent — roughly 75 MB (Tiny) to 3 GB (Large) | |
| Language | Blueprints and/or C++ | |
| Platforms | Win64 — built and verified. The code carries no Win64-only restriction, but other targets have not been verified by us. |
Installation
- Copy the
VoiceScribefolder into your project'sPlugins/directory. - C++ projects: regenerate project files and build. Blueprint-only projects: launch the editor and accept the build prompt.
- Open Edit → Plugins and confirm VoiceScribe is enabled.
- Go to Project Settings → VoiceScribe, choose a model size and language, then download the model. This is a one-off editor step.
- From C++, add
"VoiceScribe"toPublicDependencyModuleNames.
Quick start
Transcribe a buffer of audio
- Call Create Speech Recognizer and keep the result.
- Bind On Recognized Text Segment and On Recognition Finished.
- Call Start Speech Recognition.
- Push PCM in with Process Audio Data as it becomes available.
- Call Stop Speech Recognition when the speaker is done.
USpeechRecognizer* Recognizer = USpeechRecognizer::CreateSpeechRecognizer();
Recognizer->OnRecognizedTextSegment.AddDynamic(this, &AMyActor::HandleSegment);
Recognizer->OnRecognitionFinished.AddDynamic(this, &AMyActor::HandleFinished);
Recognizer->SetStreamingDefaults();
Recognizer->SetLanguage(ESpeechRecognizerLanguage::En);
Recognizer->StartSpeechRecognition();
// then, as audio arrives:
Recognizer->ProcessAudioData(PcmSamples, SampleRate, NumChannels, /*bLast=*/false);
Recognition API
| Node | Type | Description | |
|---|---|---|---|
| Create Speech Recognizer | Callable | Create a recognizer instance. | |
| Start Speech Recognition | Callable | Begin a recognition session. | |
| Process Audio Data | Callable | Push PCM samples in. | |
| Stop Speech Recognition | Callable | End the session and flush. | |
| Force Process Pending Audio Data | Callable | Process buffered audio immediately. | |
| Clear Audio Data | Callable | Discard buffered audio. | |
| Get Is Finished / Get Is Stopped / Get Is Stopping | Pure | Session state. | |
| On Recognized Text Segment | Event | Fires per transcribed segment. | |
| On Recognition Finished | Event | Fires when the session completes. | |
| On Recognition Progress | Event | Percentage progress. | |
| On Recognition Stopped | Event | Fires after an explicit stop. | |
| On Recognition Error | Event | Fires with a message on failure. |
Tuning recognition
| Node | Type | Description | |
|---|---|---|---|
| Set Streaming Defaults | Callable | Preset tuned for low latency. | |
| Set Non Streaming Defaults | Callable | Preset tuned for accuracy. | |
| Set Language | Callable | Spoken language of the audio. | |
| Set Translate To English | Callable | Translate the transcript into English. | |
| Set Num Of Threads | Callable | CPU threads used for decoding. | |
| Set Beam Size | Callable | Beam search width. Higher is slower and more accurate. | |
| Set Step Size | Callable | Audio chunk size for streaming. | |
| Set Audio Context Size | Callable | How much past audio to keep as context. | |
| Set Max Tokens | Callable | Cap tokens per segment. | |
| Set Initial Prompt | Callable | Bias decoding with expected vocabulary. | |
| Set Temperature To Increase | Callable | Fallback temperature on low confidence. | |
| Set Entropy Threshold | Callable | Threshold that triggers the temperature fallback. | |
| Set Single Segment | Callable | Force one segment for the whole input. | |
| Set No Context | Callable | Do not carry context between segments. | |
| Set Suppress Blank | Callable | Suppress blank outputs. | |
| Set Suppress Non Speech Tokens | Callable | Suppress non-speech markers. | |
| Set Speed Up | Callable | Trade accuracy for speed. | |
| Get / Set Recognition Parameters | Callable | Read or apply the whole parameter struct. | |
| Get Streaming / Non Streaming Defaults | Pure | Inspect the presets. | |
| Compute Levenshtein Similarity | Pure | Compare two strings — useful for voice commands. |
Choosing a model
Set the model under Project Settings → VoiceScribe. Quantised variants
(Q5_1, Q8_0) are considerably smaller and faster with a modest
accuracy cost.
| Model | Rough size | Best for | |
|---|---|---|---|
| Tiny / Tiny Q5_1 / Tiny Q8_0 | ~75 MB | Prototyping, short commands, mobile | |
| Base / Base Q5_1 | ~140 MB | Good general default | |
| Small / Small Q5_1 / Distil Small | ~460 MB | Noticeably better accuracy | |
| Medium | ~1.5 GB | High accuracy, desktop only | |
| Large | ~3 GB | Best accuracy, slowest |
Configuration
Recognition runs on the CPU. No optional acceleration backends are bundled, so there is nothing to enable — the settings below are the only build-time knobs.
| Switch | Default | Effect | |
|---|---|---|---|
| AVX baseline | on | CPU instruction baseline. Set via MinCpuArchX64 on UE 5.3+ and bUseAVX on older engines. Lower it in the Build.cs if you must support pre-AVX hardware. | |
| PCHUsage | NoSharedPCHs | The module raises the CPU baseline, so it builds its own PCH rather than sharing the engine's. |
VoiceScribeEditor module.Troubleshooting
| Symptom | Cause and fix | |
|---|---|---|
| No text comes back | Confirm a model is downloaded in Project Settings, that the audio is 16-bit PCM at a sane sample rate, and that you called Start Speech Recognition before pushing data. | |
| Recognition is very slow | Use a smaller or quantised model, raise Set Num Of Threads, and lower Set Beam Size. | |
| Transcript is garbled | Usually the wrong language. Set it explicitly with Set Language rather than relying on detection, and give an Initial Prompt for domain vocabulary. | |
| Packaged build is enormous | The selected model is included in the package. Choose a quantised or smaller model. | |
| Crash or illegal instruction on old CPUs | The module targets an AVX baseline. On pre-AVX hardware, lower it in the Build.cs and rebuild. |
Credits
© 2026 Alpha XP — alphaxp.net. Built and verified for Unreal Engine 4.27 and 5.8.
All licence notices, including third-party components, ship at
Source/ThirdParty/LICENSE-ThirdParty.txt.