In-process Piper text-to-speech for Node.js. No Python required.
pipertts runs Piper neural voices directly inside Node via onnxruntime-node: the ONNX session stays warm, so synthesis runs ~10x faster than spawning the Piper CLI per call (~0.22s/line vs ~2.2s/line on en_US-lessac-medium).
PiperNativeTTS): persistent ONNX session, streaming per-sentence chunksphoneme_type values: espeak, text, hebrew, lithuanian, pinyin, thai, japanesespeakerId, incl. multi-style models)Buffer or direct file write: wav, raw, mp3, ogg, opuslistPiperModels())| OS | Architecture | Requirement |
|---|---|---|
| Linux | x64 | None beyond npm install (espeak prebuild bundled) |
| Linux | arm64 | None beyond npm install (espeak prebuild bundled) |
| Windows | x64 | espeak-ng CLI (falls back automatically, e.g. via choco) |
onnxruntime-node, lindera-wasm, pyodide, and the pure-JS audio encoders ship as npm dependencies. ffmpeg is optional: when present it handles mp3/ogg/opus transcoding, otherwise pure-JS fallbacks cover mp3 (lamejs) and opus (opusscript); ogg needs ffmpeg.
npm install pipertts
pipertts does not ship Piper voice models. Get a .onnx model (and its .onnx.json) from the Piper project, or let the catalog helpers download one for you:
import { PiperNativeTTS, resolveModelPathFromOptions } from "pipertts";
const modelPath = await resolveModelPathFromOptions({
model: "en_US-lessac-medium",
modelsDir: "./models",
});
const tts = await PiperNativeTTS.load({ modelPath });
// ...or straight from a URL (any host, including Hugging Face):
const ttsUrl = await PiperNativeTTS.load({
modelUrl:
"https://huggingface.co/Artyom9800/glados_fr_piper_medium/resolve/main/fr_FR-glados_fr-medium.onnx",
modelsDir: "./models",
});
// config defaults to `<modelUrl>.json`; override with `configUrl` if needed.
// ...or from the PiperHub community catalog (~500 voices):
const ttsHub = await PiperNativeTTS.load({
model: "hub:glados-french-medium-voice",
modelsDir: "./models",
});
Phonemizer data (g2pw ~280MB, UniDic ~200MB, TLTK ~20MB, smaller tables for Hebrew/Arabic/Lithuanian) downloads once into ./piper-data/ on first use. Override with nativeDataDir or per-bundle paths (g2pwModelDir, unidicDir, tltkDataDir, nakdimonModelPath, tashkeelModelDir, lithuanianDataDir).
Piper releases: https://github.com/rhasspy/piper/releases
import { PiperNativeTTS } from "pipertts";
import * as fs from "node:fs";
async function main() {
const tts = await PiperNativeTTS.load({
modelPath: "./models/en_US-lessac-medium.onnx",
});
const result = await tts.synthesize("Hello from Piper.", {
outputFormat: "wav",
lengthScale: 0.9,
});
fs.writeFileSync("./hello.wav", result.audio);
// Multi-speaker / multi-style voices:
const excited = await tts.synthesize("What wonderful news!", {
speakerId: 1,
});
fs.writeFileSync("./excited.wav", excited.audio);
}
main().catch(console.error);
See examples/example-native.ts (run with npx tsx examples/example-native.ts).
import {
getPiperModelMetadata,
getPiperModelsByLanguage,
listPiperModels,
PiperNativeTTS,
} from "pipertts";
const models = await listPiperModels();
console.log(models.slice(0, 5)); // first catalog ids + "custom"
const metadata = await getPiperModelMetadata("en_US-lessac-medium");
console.log(metadata?.languageCode, metadata?.quality, metadata?.numSpeakers);
const englishModels = await getPiperModelsByLanguage("en");
const frenchModels = await getPiperModelsByLanguage("fr_FR");
console.log(englishModels.length, frenchModels.length);
getPiperModelMetadata("custom") returns null.
getPiperModelsByLanguage("en") matches all variants like en_US, en_GB, etc.
PiperHub indexes ~500 community voices (GLaDOS, game characters, finetunes) that the rhasspy manifest above does not cover. Each entry points at a Hugging Face repo; downloads go straight to HF:
import {
getPiperHubModelMetadata,
getPiperHubModelsByLanguage,
listPiperHubModels,
PiperNativeTTS,
} from "pipertts";
const slugs = await listPiperHubModels(); // 500+ slugs
const meta = await getPiperHubModelMetadata("glados-french-medium-voice");
console.log(meta.repo, meta.onnxUrl, meta.pageUrl);
const french = await getPiperHubModelsByLanguage("fr");
// Any of these selects the voice — bare slugs also work as `model`:
const tts = await PiperNativeTTS.load({
model: "hub:glados-french-medium-voice",
modelsDir: "./models",
});
Accepted model forms: rhasspy id/alias, hub:<slug>, full
piperhub.org/model/<slug>/ page URL, or bare slug. PiperTTS.create and
resolveModelPathFromOptions accept the same values, plus modelUrl /
configUrl for direct downloads (cached under
modelsDir/url-<hash>/, configUrl defaults to <modelUrl>.json).
Robustness: PiperHub downloads first try the voice's own repo, then fall
back automatically to the community mirror
(fox3000foxy/piper-voices-collection) when the author repo no longer
serves the files. Voices the catalog flags usable: false (broken files,
or piperPlus format needing a forked runtime) are refused with an
explanatory error — metadata exposes usable / piperPlus so you can
filter them via getPiperHubModelMetadata().
Every Piper phoneme_type is supported, verified against the Python reference:
| Type | Engine | Notes |
|---|---|---|
espeak |
N-API espeak-ng bridge (prebuilds) or CLI fallback | Byte-identical clauses |
text |
Codepoint passthrough | Trivial |
hebrew |
Nakdimon ONNX + rule-based IPA | Identical outputs |
lithuanian |
espeak + 189k-entry stress dictionary | Identical outputs |
pinyin |
g2pW BERT ONNX + WordPiece | Polyphonic disambiguation works (银行→yínháng) |
thai |
Real TLTK unmodified under Pyodide/WASM | 7/7 identical; first call takes seconds, later calls are fast |
japanese |
lindera-wasm + UniDic + TS morae engine | Segments exact; no lexical pitch accent — UniDic ships no accent data |
Arabic ar voices are diacritized with tashkeel automatically (disable with useTashkeel: false).
Inspect phonemization without synthesizing:
const phonemes = await tts.phonemize("Hello world.");
wav (default), raw PCM, mp3, ogg, opus — set via outputFormat
per call, or write straight to disk:
await tts.synthesize("Saved directly.", {
outputFormat: "mp3",
outputFile: "./direct.mp3",
});
src/native/espeak-bridge/ is an N-API port of espeakbridge.c
(byte-identical phonemes). Prebuilds ship with the package
(npm run build:prebuilds, CI in prebuilds.yml); local build with
npm run build:espeak-bridge (needs libespeak-ng-dev). Without an addon
it falls back to the espeak-ng CLI (~15ms/sentence). Set
PIPER_ESPEAK_BRIDGE=0 to force the CLI.
License note:
src/native/ports piper1-gpl (GPL-3.0-or-later);src/native/chinese.tsports g2pw (Apache-2.0, Yi-Chang Chen); Thai runs TLTK (BSD-3-Clause, Chulalongkorn University) unmodified; Japanese segments with lindera-wasm + UniDic (MIT, no accent data); the phonemizer data bundles keep their own licenses (tashkeel/hebrew: GPL/MIT, Lithuanian TSVs: CC-BY-4.0, g2pw model: Apache-2.0).
ESM:
import { PiperNativeTTS } from "pipertts";
const tts = await PiperNativeTTS.load({
modelPath: "./models/en_US-lessac-medium.onnx",
});
CommonJS:
const { PiperNativeTTS } = require("pipertts");
async function boot() {
const tts = await PiperNativeTTS.load({
modelPath: "./models/en_US-lessac-medium.onnx",
});
await tts.synthesize("Hello from CommonJS", {
outputFile: "./cjs.wav",
});
}
boot().catch(console.error);
PiperNativeTTS.load(options)Loads the ONNX model and resolves phonemizer resources (downloading data on first use).
| Option | Type | Required | Default |
|---|---|---|---|
modelPath |
string |
one of modelPath / modelUrl / model |
- |
modelUrl |
string |
one of modelPath / modelUrl / model |
- |
model |
string |
one of modelPath / modelUrl / model |
- |
configPath |
string |
no | auto (<model>.json) |
configUrl |
string |
no | auto (<modelUrl>.json) |
modelsDir |
string |
no | "models" (download cache) |
numThreads |
number |
no | ONNX default |
espeakBinary |
string |
no | espeak-ng |
session |
NativeSession |
no | created from modelPath |
nativeDataDir |
string |
no | "./piper-data" |
tashkeelModelDir / nakdimonModelPath / g2pwModelDir / lithuanianDataDir / tltkDataDir / unidicDir |
string |
no | auto-downloaded |
useTashkeel |
boolean |
no | true |
taskeenThreshold |
number |
no | 0.8 |
tts.synthesize(text, options?)Synthesizes speech and returns { audio: Buffer, sampleRate, text }.
Audio arrives sentence by sentence from synthesizeChunks(); chunks are
concatenated with sentenceSilence gaps, then optionally transcoded.
| Option | Type | Default | Notes |
|---|---|---|---|
speakerId |
number |
config default | Multi-speaker/style index |
noiseScale |
number |
0.667 |
Voice variability (output is stochastic) |
noiseWScale |
number |
0.8 |
Timing variability |
lengthScale |
number |
1.0 |
>1 slower, <1 faster |
sentenceSilence |
number |
0 |
Gap between sentences, in seconds |
outputFormat |
"wav" | "raw" | "mp3" | "ogg" | "opus" |
"wav" |
Transcoded when needed |
outputFile |
string |
- | Also writes the final audio there |
normalizeAudio |
boolean |
true |
Scale to peak |
volume |
number |
1.0 |
Gain multiplier |
tts.phonemize(text) / tts.synthesizeChunks(text, options?)Lower-level access: phonemize returns phonemes grouped by sentence;
synthesizeChunks is an async generator yielding { sampleRate, phonemes, phonemeIds, pcm16 } per sentence.
tts.getConfig() / tts.getModelPath()Introspect the resolved voice config and model path.
PiperTTS)PiperTTS shells out to python3 -m piper (one process per synthesis).
It stays available for environments where the native engine cannot run,
but it is ~10x slower and needs Python plus the Piper module installed:
import { PiperTTS } from "pipertts";
const tts = await PiperTTS.create({
modelPath: "./models/en_US-lessac-medium.onnx",
});
// or: piperBinaryPath, warmUpText/skipWarmup, synthesisTimeoutMs,
// defaultOptions (PiperInferenceOptions: modelPath, configPath, outputFile,
// outputFormat, speakerId, noiseScale, noiseWScale, lengthScale,
// sentenceSilence, jsonInput, numThreads, useCuda, logLevel, timeoutMs)
PiperTTS.create performs a warm-up inference unless skipWarmup is set.
See examples/example.ts.
model file not found: verify modelPath is correct.unknown model "...": call listPiperModels() and use one of returned ids, or set model: "custom".onnxruntime-node could not be loaded: reinstall dependencies (npm install).No module named piper / Python not found (wrapper only): install Python 3 and the Piper module, verify python3 -m piper --help.| Component | Supported |
|---|---|
| Node.js | 18+ |
| TypeScript | 5+ |
| Runtime | Linux x64/arm64, Windows x64 |
Notes:
npm install
npm run build
npm run typecheck
npm test
modelUrl / configUrl on PiperNativeTTS.load and PiperTTS.create:
direct .onnx downloads (any host, including Hugging Face), cached once
under modelsDir/url-<hash>/, config defaults to <modelUrl>.json.listPiperHubModels,
getPiperHubModelMetadata, getPiperHubModelsByLanguage,
ensurePiperHubModelDownloaded. model accepts rhasspy ids,
hub:<slug>, piperhub.org page URLs, and bare slugs.piper-voices-collection mirror
when an author repo stops serving files; voices flagged usable: false
(broken files, piper-plus format) fail fast with a clear error.resolveModelFiles used by both engines; 15 hermetic tests
(local HTTP server, no external network).generate-doc / publish-doc
scripts, CI docs job with --treatWarningsAsErrors).PiperNativeTTS: the CLI wrapper moves to a
clearly-labeled legacy section; no Python in setup, platform, or API docs.examples/example-native.ts: same flow as example.ts, fully in-process.resolveModelPathFromOptions now public (catalog ids for native users).matchAll Thai runs, zero-copy WAV parse. Inference was already 99%+.npm i failed: the install script was absent from the tarball.
files now ships scripts/install-bridge.js, binding.gyp, and the
bridge C source (verified with a blank-dir end-user install)._GNU_SOURCE for RTLD_DEFAULT).lint/check/typecheck battery green: root biome.json
(useNamingConvention off — public API and upstream snake_case must
stay), all other violations fixed, no logic changes.bun.lockb → bun.lock, real tsc --noEmit step,
bun.lock synced (missing opusscript/lindera/pyodide broke CI tests).*.test.js no longer ships in the npm tarball
(tsconfig.build.json; typecheck still covers tests).espeak_TextToPhonemesWithTerminator at runtime
(dlsym/GetProcAddress): old headers compile, pre-1.52 libraries
degrade to single-clause output.prebuild script renamed to build:prebuilds (npm treated it as a
pre-hook of build, running prebuildify on every build).node-gyp rebuild for packages with a root binding.gyp;
the new tolerant scripts/install-bridge.js prefers prebuilds, tries a
source build, and warns instead of failing (CLI fallback).Measured on en_US-lessac-medium: ~2.2s/line (CLI wrapper) vs ~0.22s/line
(native) — roughly 10x, from killing the per-call Python spawn + model reload.
Added:
PiperNativeTTS — persistent in-process ONNX session (onnxruntime-node),
one inference per sentence, streaming chunks.phoneme_type values ported and verified byte-identical (or
equivalent) against the Python reference:
espeak/text — N-API bridge (byte-identical clauses) with CLI fallback.hebrew — Nakdimon ONNX + 494 lines of IPA rules, identical outputs.lithuanian — espeak + 189k-entry stress dictionary, identical outputs.pinyin — full g2pW BERT port (WordPiece, features, 159MB graph);
polyphonic disambiguation works (银行→yínháng).thai — real TLTK unmodified under Pyodide/WASM, 7/7 identical.japanese — lindera-wasm + UniDic segmentation with a TS morae engine,
4/4 segment-identical to pyopenjtalk (no lexical pitch accent — UniDic
ships no accent data; documented limitation).ar voices get tashkeel diacritization automatically.wav/raw/mp3/ogg/opus in both engines: ffmpeg when
present, otherwise pure-JS fallbacks (lamejs for mp3, opusscript + Ogg
muxer for opus — 0.993 correlation vs source)../piper-data/ (g2pw ~280MB,
UniDic ~200MB, TLTK ~20MB, others small).prebuildify, CI matrix linux/win/macos).bun test suite: 71 tests (RBNF tables, transformer vectors, Ogg CRC,
gated real-model e2e). npm test was broken (biome test does not exist).outputFormat, numeric range validation,
synthesis timeouts, streaming catalog downloads with size/md5 checks,
executable checks with PATHEXT support, skippable warm-up.Not ported (documented): marine accent model, Pyodide-independent Thai,
lexical Japanese pitch accent.
Thin TypeScript wrapper around python3 -m piper: process spawn per call,
WAV-or-tempfile output, Hugging Face catalog helpers.