Model picker guide
Every Rymi agent runs a language model (the brain) and a voice (the sound). On a custom agent you can pick any combination. The models you choose only change the per-minute call cost (component cost + a flat $0.02/min platform fee).
Quick recommendations
I want the cheapest agent that still feels good
Pick GPT-4o Mini or Claude Haiku 4.5 for the LLM, OpenAI TTS for voice. Latency low, cost low, quality fine for most short flows.
I want the most natural-sounding agent
Pick Claude Sonnet 4.6 for the LLM (or Opus if budget allows), ElevenLabs for voice. Best for high-stakes calls: concierge, executive support, premium sales.
I need the lowest latency possible
Pick a realtime path: GPT Realtime with its native voice, or Gemini 2.5 Flash with native audio. Avoid stacking separate TTS providers. Each hop adds 100–200 ms. Realtime voices aren't native-sounding in Hindi, so use them for English.
My users speak Hindi or other Indic languages
Pick Saaras v4 for speech recognition and Sarvam Bulbul v3 for voice, with GPT-4.1 or Sarvam 105B Conversations as the LLM. Bulbul gives a native Hindi voice, and Sarvam runs from Indian regions for lower latency.
Language models
Anthropic (Claude)
Strong reasoning, careful tone. Good default for most production agents.
| Model | Best for | Notes |
|---|---|---|
claude-haiku-4-5 | Fast, friendly tone, handles 80% of support / qualification flows | Cheapest Claude |
claude-sonnet-4-6 | Balanced quality + speed. Solid for sales discovery and multi-step playbooks | |
claude-opus-4-8 | Highest reasoning. Use when nuance matters most: complex objection handling, escalations | Flagship · most capable |
claude-opus-4-7 | Prior flagship. Strong reasoning, still selectable for pinned deployments | |
claude-opus-4-6 | Prior Opus generation. Still selectable for pinned deployments |
All three Opus models bill at the same per-minute rate; pick on capability, not cost. Opus is the most expensive Claude tier; Haiku the cheapest.
OpenAI (GPT)
Wide tool support, strong realtime variant for low-latency calls.
| Model | Best for | Notes |
|---|---|---|
gpt-4o-mini | Short verification or routing flows | Cheapest |
gpt-4o | General-purpose flagship. Reliable for most agent shapes | |
gpt-realtime-mini | Low-latency native-audio voice on a budget | Realtime · retiring 20 Jan 2027 |
gpt-realtime-1.5 | High-fidelity realtime voice. Pick for premium concierge experiences | Realtime · premium |
gpt-realtime-2 | OpenAI's newest realtime model | Realtime |
Google (Gemini)
Native multimodal audio path. Strong default for voice-first agents.
| Model | Best for | Notes |
|---|---|---|
gemini-2.5-flash-lite | High-volume top-of-funnel | Cheapest |
gemini-2.5-flash | Balanced quality and speed. Pair with native audio for low end-to-end latency | Native audio |
gemini-2.5-pro | Highest Gemini quality for nuanced, multi-step flows |
Sarvam (India-optimized)
Tuned for Indian English, Hindi, and other Indic languages. Lower latency in India.
| Model | Best for |
|---|---|
sarvam-105b-conversations | Built for voice agents: first words in about a third of a second, in Hindi and 10 other Indian languages. Experimental while Rymi verifies it on live calls |
sarvam-105b | Sarvam's reasoning model. Slower to start speaking; prefer the conversations model for calls |
Speech recognition
The speech-to-text model hears the caller. Pick it for the caller's language.
| Model | Best for |
|---|---|
flux-general-en (Deepgram Flux) | English calls. Detects the end of the caller's turn itself, so replies start sooner |
flux-general-multi (Deepgram Flux) | Callers who switch languages mid-call |
nova-3-general, nova-3-multilingual (Deepgram Nova-3) | General English or multilingual transcription |
nova-2-phonecall (Deepgram Nova-2) | Low-bitrate phone audio |
saaras:v4 (Sarvam Saaras v4) | The 22 scheduled Indian languages, Hindi included, plus global English accents. Same price as Saaras v3 |
saaras:v3 (Sarvam Saaras v3) | The previous Saaras model, still available |
scribe_v2_realtime (ElevenLabs Scribe) | Multilingual realtime transcription |
Realtime models (GPT Realtime, Gemini native audio) bundle their own speech recognition, so you don't pick one separately.
Voices
| Provider | What it is | Best for |
|---|---|---|
| Gemini native audio | Built into the Gemini stack with no extra hop | Lowest end-to-end latency. Default if you pick a Gemini model. |
| OpenAI TTS | 13 voices, mostly gendered (alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, cedar) | Cheap, fast, consistent. Limited expressive range. |
| ElevenLabs | 22 curated voices across 18 languages, accent control; connect your own key to use your full ElevenLabs library | Highest perceived quality and variety. Best for brand-sensitive deployments. |
| Deepgram Aura 2 | 12 curated English voices | Low-latency, natural English TTS. Strong fallback option. |
| Sarvam Bulbul v3 | Indic-optimized TTS | Hindi and other Indic languages with natural prosody. |
| Cartesia Sonic | Low-latency premium TTS with high naturalness, BYO key supported | Newest premium voice tier. |
When a model is retired
Vendors shut models down. When an agent uses a model, voice or managed stack within 30 days of its shutdown date, Studio shows a banner on the agent with the date and a suggested replacement, and marks the agent Retiring in the sidebar. Switch before the date, because the vendor stops serving the model after it. Retired voices and stacks drop out of the pickers.
Bring your own keys
Connect your OpenAI / Anthropic / ElevenLabs / Cartesia / Groq / and other provider keys under Settings → BYO Providers to route through your own accounts. If no key is connected, Rymi falls back to the platform default. See Voice Providers API for the full BYOK list.
What's next
- Get your number on a Rymi agent
- Custom Personas: tune voice and identity in detail
- API: Voice Providers

