Research & Innovation

Pushing the boundaries of Voice AI

Vodex is an AI-native voice automation company building speech intelligence for real-world, high-stakes conversations in mortgage, insurance, and collections.

Debt Collection

Where this work began

The team's background is foundational AI/ML and speech synthesis. Vodex is the product of years spent inside these models, not a wrapper around someone else's.

First experiments in speech synthesis — photo

2020

First experiments in speech synthesis

The team's background is foundational AI/ML and speech synthesis. Vodex is the product of years spent inside these models, not a wrapper around someone else's.

Retrieval and language modeling — photo

2021

Retrieval and language modeling

Work with FAISS, Haystack-RAG, and BLOOM built the retrieval and language foundations that conversational agents depend on.

An AI-native voice automation company — photo

Today

An AI-native voice automation company

Vodex builds speech intelligence for real-world, high-stakes conversations in the mortgage, insurance, and collections industries.

Running a collections operation? See Vodex for Debt Collection

Speech synthesis

Why we built our own TTS

Existing text-to-speech fell short on naturalness, expressiveness, and control. Production conversation needs behaviors off-the-shelf models often lack: laughing, sighing, pausing, expressing uncertainty. So we trained our own.

Orpheus core architecture

Our pretrained TTS model is built on the Orpheus core architecture, tuned for production conversation rather than demos.

21,000+ hours of training data

Trained on more than 21,000 hours of diverse, expressive speech, so the model has heard how people actually talk.

Zen-Tokenizer inside

Our proprietary neural audio codec is integrated at the audio tokenization layer, not bolted on afterward.

Narrowband and wideband

Supports both 8kHz telephony pipelines and 16kHz wideband audio from the same stack.

Running a collections operation? See Vodex for Debt Collection

Named Voices

Voices that feel human

Each production voice is a distinct speaker with its own character, trained for a specific kind of conversation.

Shreya

The most soft-spoken and emotionally expressive of our voices. Ideal for empathetic agentic use cases where tone carries the conversation.

Shweta

Trained to sound like a professional audiobook narrator. Smooth, clear storytelling with steady pacing.

Aastha

A latency-optimized voice that preserves expressiveness while minimizing response time. Built for real-time back-and-forth.

Running a collections operation? See Vodex for Debt Collection

Latency

Fast enough to interrupt

A voice agent that pauses too long breaks the illusion. We keep pushing time-to-first-byte down through codebook structure refinement and distributed deployment techniques.

Learn More About Integrations

189ms

Average time-to-first-byte (TTFB) today

<80ms

Target response time in production

Roadmap

What comes next

We draw inspiration from Kyutai (Moshi's full-duplex and MIMI tokenizer work), Snac's neural codecs, Canaophy Labs, Sesame Labs and their CSM stack, and Carson et al. (2025). Agentic voice AI is a long-term mission, and we want company on the road.

Buy Now Pay Later (BNPL) — photo

Buy Now Pay Later (BNPL)

Short-term installment collections, payment reminders, failed payment follow-ups and more.

Medical & Healthcare — photo

Medical & Healthcare

HIPAA-compliant patient payment reminders, EOB clarification, and deductible notifications.

Buy Here Pay Here (BHPH) — photo

Buy Here Pay Here (BHPH)

High-frequency outreach for late payments, reminders, and CPI or confusion clarification.

Enterprise-grade security & compliance

Built for enterprises

Exceptional performance, scalability, and dedicated support. Deploy compliant AI agents that accelerate recovery while cutting cost-per-contact.

SOC 2 Type II · ISO 27001 · 24/7 support

Ready to supercharge
your engagement?

Join the enterprises already using Vodex to turn reminders, collections and follow-ups into recovered revenue.