ZCRM

Voice conversion for Vietnamese & English

Explore the features ofVONIA

From text-to-speech to transcription and API integration, here are the main feature groups in VONIA.

Text-to-speech (TTS)

30 ready-made preset voices, no sample recording needed.

Enter text, pick one of 30 built-in preset voices (named after stars), and get back an audio file. Non-verbal cues can be inserted into the text (laughter, sighs...) along with pronunciation fixes for technical terms, plus text normalization so phone numbers, prices, and symbols are read correctly instead of misread. Output supports multiple formats: WAV, MP3, PCM.

See pricing
Text input and voice selection screen in VONIA
30ready-made preset voices

Voice cloning

Cloned voices are saved for use in later projects.

Given a 3–10 second reference audio sample, VONIA clones the voice to read any new text; it can automatically transcribe the sample if the original script isn't provided. Cloned voices are saved to the Voice Library for review or deletion, reusable across future projects instead of resubmitting an audio sample every time.

See pricing
Cloning a voice from a 3-10 second audio sample in VONIA
3–10 saudio sample to clone a voice

Voice design by attributes

No audio sample on hand? Choose gender, age, pitch, and regional accent directly to create a new voice for test content without needing to find someone to record a sample.

See pricing

Multi-voice dialogue

Many characters, one finished audio file.

Build a dialogue with multiple characters, each assigned its own voice (preset, cloned, or designed), then merge everything into a single finished audio file. Suited for podcast scripts or ads with dialogue, without recording multiple people.

See pricing
Building multi-character dialogue with a distinct voice per character in VONIA

Speech-to-text (STT)

Upload audio or video, get text with timestamps.

Upload audio or video to transcribe it into timestamped text segments, running on the Whisper model locally on VONIA's server with no external paid API calls. Export a standard SRT subtitle file directly for video, no separate captioning tool needed.

See pricing
Speech-to-text (STT) transcription interface in VONIA

Flexible input/output formats

VONIA accepts many audio/video input formats for cloning or transcription (mp3, wav, m4a, aac, ogg, flac, mp4, mov, webm...), so users don't need to convert formats before uploading. TTS output supports WAV, MP3, or PCM depending on integration needs.

See pricing

API for developers

Integrate VONIA with internal systems via API and webhooks.

VONIA offers an OpenAI-compatible endpoint (`/v1/audio/speech`), letting businesses already integrated with OpenAI TTS switch over with almost no code changes: just update the server address and API key. A dedicated API also supports `voice_id` and `seed` to keep a voice's identity consistent across calls, important when generating bulk content with the same voice. Authentication uses an API key header for all server-to-server integrations.

See pricing
Integration webhook configuration in VONIA

Self-hosted infrastructure

Dedicated GPU infrastructure, predictable running costs.

The voice conversion model runs on VONIA's own GPU, without pay-per-call API calls to OpenAI/Google, which helps control operating costs at higher usage volumes.

See pricing
Runtime environment setup for the VONIA system

Login & account management

Manage your account and license right in the app.

Login via Google SSO, with sessions kept for 7 days. The web interface automatically switches to a mobile layout (Clone, TTS, STT, Settings in a bottom tab bar), usable directly on a phone without installing a separate app.

See pricing
Account settings screen in VONIA

See VONIA pricing

Contact us for a quote tailored to your content volume.