Plugin Documentation

VoiceScribe
Offline speech recognition

Real-time, on-device speech-to-text with selectable models and languages. No network calls, no API keys, nothing leaves the machine.

v1.0.0
UE 4.27 & 5.8
Win64 (verified)
Category Audio
Offline
Alpha XP

Overview

VoiceScribe turns speech into text entirely on the local machine. No network calls, no API keys, no per-request cost, and nothing leaves the device.

Feed it audio — a recording, a microphone stream, or an imported file — and it returns transcribed text, either in one pass or segment by segment as the speaker talks.

  • Fully offline — the model ships inside your packaged build
  • Streaming or one-shot recognition
  • Multi-language, with optional translation to English
  • Selectable model size, from Tiny to Large, including quantised variants
  • Editor tool to download and manage models
  • Tunable decoding: threads, beam size, temperature, prompts, token limits
Privacy Because recognition runs locally, VoiceScribe suits projects with privacy or compliance constraints, and it keeps working with no internet connection.

Requirements

RequirementDetails
Unreal Engine4.27 and 5.8 (verified clean on both)
ModulesVoiceScribe (Runtime), VoiceScribeEditor (Editor)
CPUx64 with AVX. Recognition is CPU-bound.
DiskModel dependent — roughly 75 MB (Tiny) to 3 GB (Large)
LanguageBlueprints and/or C++
PlatformsWin64 — built and verified. The code carries no Win64-only restriction, but other targets have not been verified by us.

Installation

  1. Copy the VoiceScribe folder into your project's Plugins/ directory.
  2. C++ projects: regenerate project files and build. Blueprint-only projects: launch the editor and accept the build prompt.
  3. Open Edit → Plugins and confirm VoiceScribe is enabled.
  4. Go to Project Settings → VoiceScribe, choose a model size and language, then download the model. This is a one-off editor step.
  5. From C++, add "VoiceScribe" to PublicDependencyModuleNames.
Pick the model before packaging The selected model is baked into the build. Start with Tiny or Base while prototyping — the larger models are far more accurate but much slower and much bigger.

Quick start

Transcribe a buffer of audio

  1. Call Create Speech Recognizer and keep the result.
  2. Bind On Recognized Text Segment and On Recognition Finished.
  3. Call Start Speech Recognition.
  4. Push PCM in with Process Audio Data as it becomes available.
  5. Call Stop Speech Recognition when the speaker is done.
USpeechRecognizer* Recognizer = USpeechRecognizer::CreateSpeechRecognizer();

Recognizer->OnRecognizedTextSegment.AddDynamic(this, &AMyActor::HandleSegment);
Recognizer->OnRecognitionFinished.AddDynamic(this, &AMyActor::HandleFinished);

Recognizer->SetStreamingDefaults();
Recognizer->SetLanguage(ESpeechRecognizerLanguage::En);
Recognizer->StartSpeechRecognition();

// then, as audio arrives:
Recognizer->ProcessAudioData(PcmSamples, SampleRate, NumChannels, /*bLast=*/false);
Streaming versus one-shot Set Streaming Defaults favours latency and emits partial segments as the speaker talks. Set Non Streaming Defaults favours accuracy on a complete recording.

Recognition API

NodeTypeDescription
Create Speech RecognizerCallableCreate a recognizer instance.
Start Speech RecognitionCallableBegin a recognition session.
Process Audio DataCallablePush PCM samples in.
Stop Speech RecognitionCallableEnd the session and flush.
Force Process Pending Audio DataCallableProcess buffered audio immediately.
Clear Audio DataCallableDiscard buffered audio.
Get Is Finished / Get Is Stopped / Get Is StoppingPureSession state.
On Recognized Text SegmentEventFires per transcribed segment.
On Recognition FinishedEventFires when the session completes.
On Recognition ProgressEventPercentage progress.
On Recognition StoppedEventFires after an explicit stop.
On Recognition ErrorEventFires with a message on failure.

Tuning recognition

NodeTypeDescription
Set Streaming DefaultsCallablePreset tuned for low latency.
Set Non Streaming DefaultsCallablePreset tuned for accuracy.
Set LanguageCallableSpoken language of the audio.
Set Translate To EnglishCallableTranslate the transcript into English.
Set Num Of ThreadsCallableCPU threads used for decoding.
Set Beam SizeCallableBeam search width. Higher is slower and more accurate.
Set Step SizeCallableAudio chunk size for streaming.
Set Audio Context SizeCallableHow much past audio to keep as context.
Set Max TokensCallableCap tokens per segment.
Set Initial PromptCallableBias decoding with expected vocabulary.
Set Temperature To IncreaseCallableFallback temperature on low confidence.
Set Entropy ThresholdCallableThreshold that triggers the temperature fallback.
Set Single SegmentCallableForce one segment for the whole input.
Set No ContextCallableDo not carry context between segments.
Set Suppress BlankCallableSuppress blank outputs.
Set Suppress Non Speech TokensCallableSuppress non-speech markers.
Set Speed UpCallableTrade accuracy for speed.
Get / Set Recognition ParametersCallableRead or apply the whole parameter struct.
Get Streaming / Non Streaming DefaultsPureInspect the presets.
Compute Levenshtein SimilarityPureCompare two strings — useful for voice commands.

Choosing a model

Set the model under Project Settings → VoiceScribe. Quantised variants (Q5_1, Q8_0) are considerably smaller and faster with a modest accuracy cost.

ModelRough sizeBest for
Tiny / Tiny Q5_1 / Tiny Q8_0~75 MBPrototyping, short commands, mobile
Base / Base Q5_1~140 MBGood general default
Small / Small Q5_1 / Distil Small~460 MBNoticeably better accuracy
Medium~1.5 GBHigh accuracy, desktop only
Large~3 GBBest accuracy, slowest
Packaged size The model ships with your build. A Large model adds roughly 3 GB to the package — budget for it, or pick a quantised variant.

Configuration

Recognition runs on the CPU. No optional acceleration backends are bundled, so there is nothing to enable — the settings below are the only build-time knobs.

SwitchDefaultEffect
AVX baselineonCPU instruction baseline. Set via MinCpuArchX64 on UE 5.3+ and bUseAVX on older engines. Lower it in the Build.cs if you must support pre-AVX hardware.
PCHUsageNoSharedPCHsThe module raises the CPU baseline, so it builds its own PCH rather than sharing the engine's.
Editor settings Model size, language and download management are under Project Settings → VoiceScribe, provided by the VoiceScribeEditor module.

Troubleshooting

SymptomCause and fix
No text comes backConfirm a model is downloaded in Project Settings, that the audio is 16-bit PCM at a sane sample rate, and that you called Start Speech Recognition before pushing data.
Recognition is very slowUse a smaller or quantised model, raise Set Num Of Threads, and lower Set Beam Size.
Transcript is garbledUsually the wrong language. Set it explicitly with Set Language rather than relying on detection, and give an Initial Prompt for domain vocabulary.
Packaged build is enormousThe selected model is included in the package. Choose a quantised or smaller model.
Crash or illegal instruction on old CPUsThe module targets an AVX baseline. On pre-AVX hardware, lower it in the Build.cs and rebuild.

Credits

© 2026 Alpha XP — alphaxp.net. Built and verified for Unreal Engine 4.27 and 5.8.

All licence notices, including third-party components, ship at Source/ThirdParty/LICENSE-ThirdParty.txt.