CAMLIN SPEECH

Streaming STT, Australian voices, and frame-accurate lip-sync

Words land on screen as you speak. The turn ends on a natural pause, not a timer. Camlin Voice answers with Australian voices recorded in Sydney, lip-sync included. Our engine by default; the big cloud engines when a tender insists.
Multi
Speech engines
Camlin default; swap per tender
8
Recognition modes
Spoken to structured
16
Business field types
Dates, numbers, IDs
Try the live demo

Opens the speech portal — mic on, your words land as you speak.

Camlin Speech

Recognition, prompting, and validation

LIVE EXAMPLE
Recognition result

Recognition + prompting
Capture, validate, and respond in one service
Structured output
Business fields and confidence checks
THE SPEECH STACK

Own the speech layer, keep provider choice

Recognition, text-to-speech, prompts and structured capture in one layer shared by every channel.

Streaming STT, ended on a natural pause

Words commit as the caller speaks. The turn closes on a natural pause, not a fixed timer, and the pause length is authorable per Interaction.

Word-by-word live captions
Authorable natural-pause endpointing
Sydney or private deployment

Tenant vocabularies and field capture

Register the names, brands and reference formats your callers actually say. Recognition leans toward them, and the output lands in business fields.

Per-tenant lexicons
Names, IDs, dates, and references
Structured output for Interactions

Camlin Voice: TTS with lip-sync built in

Every synthesis response carries the audio and the lip-sync track that drives the avatar on a live call. One engine, no third party in the path.

52 ARKit blendshapes at 60 fps
Audio + lip-sync in one response
AU voices recorded in Sydney

Provider choice when you need it

The Camlin engine is the default. Cloud speech engines stay available per journey for tenders, languages or fallback, authored in Architect.

Major cloud speech engines supported
Provider per journey
8 recognition modes
SPEECH PIPELINE

How speech flows through the platform

Spoken or typed input through recognition, validation and response.
Spoken words become checked business fieldsA man on the phone says “Reference 2448 0193, the 14th of October, $120.” The words stream in as they are said and the turn ends on the natural pause; the avatar on a mobile, listening, hands the journey a reference, a date and an amount, each with a confidence check.“Reference 2448 0193, the14th of October, $120.”Streamed as it is spokenTO THE JOURNEYREFERENCE2448 0193DATE14 OctAMOUNT$120.00Eight recognition modes, sixteen field types; every field carries a confidence check.

Caller speaks or types

Audio or text input arrives from the channel — phone, web chat, avatar, or SMS.

The channel determines which speech provider and mode to use based on the Interaction design.

CAMLIN VOICE

Australian voices, frame-accurate lip-sync

Australian voices, recorded in Sydney. Audio and lip-sync track come back in one JSON response, so the voice on your IVR can drive your avatar too.
HEAR IT LIVE

Hear the difference

Side-by-side clips where Camlin Speech nails Australian phrases the cloud STTs drop. Then try the live engine on your own voice.

More live demos on the speech portal: speech.camlinconnect.com.au. Prefer to stay on this site? See the embedded caption preview.

Recent Speech releases are listed on What’s new.