Speech synthesis & scripts

Voice QA Studio

Turn scripts into voices, and voices into usable data.

Synthesize text and XLSX scripts with Supertonic TTS. Set voice, speed and pauses per speaker, then export segments, combined audio and timestamp reports.

Voice QA Studio actual UI with fictional sample data

Actual program UI with original fictional sample data. Click to view full size. UI rendered with demo data; activity collection and speech synthesis were not run.

OVERVIEW

Overview

A Windows application for managing Supertonic TTS production. Create narration or multi-speaker conversations and connect audio with script timing for STT evaluation datasets and audiobook workflows.

WHAT YOU CAN DO

Features

Text synthesis

Split text into sentence segments and highlight the active segment. Choose playback, saving or both, including alternating male/female conversation mode.

XLSX script editing

Import scripts, add or remove rows, and edit speakers, text and voice IDs. Header-based mapping supports the original seven-column template and optional Style/Pitch columns.

10 voices & speed controls

Choose male M1–M5 and female F1–F5 voices with per-row speed. Style values and text/remark tags assist recommended speed and pauses.

Conversation pauses

Insert actual PCM silence after each row. Empty, AUTO or -1 pause values use the default; explicit values take priority.

Audio & XLSX reports

Save segment audio, a combined track and timestamp reports together. Use WAV, or MP3/FLAC when FFmpeg is available.

Streaming combined output

Stream long combined audio to disk rather than accumulating all PCM in memory. Save edited scripts or blank XLSX templates.

Who it is for

  • Creators making audiobooks, narration and multi-speaker audio
  • People preparing STT evaluation audio with reference scripts and timestamps
YOUR FIRST STEPS

How to use

  1. Prepare the runtime and TTS model, then import text or an XLSX script.
  2. Set voice, speed and pauses for each speaker and preview playback.
  3. Save all to generate segments, combined audio and an XLSX report.
  4. Save your edited script to XLSX for future sessions.

Supported environment

  • Windows 10/11 64-bit · C# WPF / .NET 8
  • Python 3.10 or later · internet for initial environment, packages and model setup
  • MP3/FLAC output: FFmpeg on PATH · assets folder required when moving to another PC

Availability

AVAILABILITY

Not publicly released

This program is currently private. Source code and installation files are not publicly available. Download information will be provided here when a public release is ready.

Things to know

Pitch values are preserved in import, editing, exports and reports, but audio pitch post-processing is disabled by default. This is not direct engine pitch control.

Parenthesized and standalone Chinese characters are removed before synthesis. Some failures fall back to Windows TTS, which can change the voice; review the audio.

Initial startup prepares a virtual environment, packages and model downloads. Include the assets folder when moving the executable to another PC.

Information checked on October 8, 2026, against repository documentation. Availability and requirements may change.