Voice conversion for Vietnamese & English
Explore the features ofVONIA
From text-to-speech to transcription and API integration, here are the main feature groups in VONIA.
Text-to-speech (TTS)
30 ready-made preset voices, no sample recording needed.
Enter text, pick one of 30 built-in preset voices (named after stars), and get back an audio file. Non-verbal cues can be inserted into the text (laughter, sighs...) along with pronunciation fixes for technical terms, plus text normalization so phone numbers, prices, and symbols are read correctly instead of misread. Output supports multiple formats: WAV, MP3, PCM.

Voice cloning
Cloned voices are saved for use in later projects.
Given a 3–10 second reference audio sample, VONIA clones the voice to read any new text; it can automatically transcribe the sample if the original script isn't provided. Cloned voices are saved to the Voice Library for review or deletion, reusable across future projects instead of resubmitting an audio sample every time.

Voice design by attributes
No audio sample on hand? Choose gender, age, pitch, and regional accent directly to create a new voice for test content without needing to find someone to record a sample.
Multi-voice dialogue
Many characters, one finished audio file.
Build a dialogue with multiple characters, each assigned its own voice (preset, cloned, or designed), then merge everything into a single finished audio file. Suited for podcast scripts or ads with dialogue, without recording multiple people.

Speech-to-text (STT)
Upload audio or video, get text with timestamps.
Upload audio or video to transcribe it into timestamped text segments, running on the Whisper model locally on VONIA's server with no external paid API calls. Export a standard SRT subtitle file directly for video, no separate captioning tool needed.

Flexible input/output formats
VONIA accepts many audio/video input formats for cloning or transcription (mp3, wav, m4a, aac, ogg, flac, mp4, mov, webm...), so users don't need to convert formats before uploading. TTS output supports WAV, MP3, or PCM depending on integration needs.
API for developers
Integrate VONIA with internal systems via API and webhooks.
VONIA offers an OpenAI-compatible endpoint (`/v1/audio/speech`), letting businesses already integrated with OpenAI TTS switch over with almost no code changes: just update the server address and API key. A dedicated API also supports `voice_id` and `seed` to keep a voice's identity consistent across calls, important when generating bulk content with the same voice. Authentication uses an API key header for all server-to-server integrations.

Self-hosted infrastructure
Dedicated GPU infrastructure, predictable running costs.
The voice conversion model runs on VONIA's own GPU, without pay-per-call API calls to OpenAI/Google, which helps control operating costs at higher usage volumes.

Login & account management
Manage your account and license right in the app.
Login via Google SSO, with sessions kept for 7 days. The web interface automatically switches to a mobile layout (Clone, TTS, STT, Settings in a bottom tab bar), usable directly on a phone without installing a separate app.
