Recognize spoken input.
Speech-to-text supplies the transcript. Your application interprets the request and decides what work follows.
Turn a spoken request into text and your application's reply into audio.
Recognition and synthesis run on Origon's speech engine and GPU infrastructure.
Speech-to-text supplies the transcript. Your application interprets the request and decides what work follows.
Text-to-speech reads the response your application produces. Combine the two for a spoken exchange.
Speech in Origon’s native voice stack
Sub-300ms
Voice round trip on Origon's native voice stack.
In Origon's native voice stack, an agent can pause and respond when a person interrupts. Speech works with the media infrastructure to support that exchange.
This figure describes the native voice stack on Origon infrastructure. It is not a standalone STT or TTS API latency guarantee. Evaluate latency for your application and configuration.
Capacity planning
The number of simultaneous recognition and synthesis tasks depends on the GPU hardware allocated to the workload.
Origon runs the reserved Speech service; your team runs the application. Capacity and operating terms follow the agreed workload.
Voice Network
Add Voice Network when the application needs phone numbers, PSTN or SIP connectivity.
Speech and Voice Network are separate services. Use either on its own or combine them for a voice application.
Evaluation
Confirm the languages, voice options and interfaces for your configuration. Use representative terminology and audio conditions to assess recognition and generated speech.
For a complete voice application, include response time and interruption handling. The supported configuration, capacity and delivery timing are agreed before reservation.
Language and terminology
Audio conditions
Interaction behavior
Measurement boundary