ElevenLabs
What is HardGraph? HardGraph publishes curated, provenance-backed agent skills grounded in reproducible vendor documentation.
ElevenLabs is a voice AI platform, not a single API: text-to-speech is the core primitive, but the same account also reaches speech-to-text, voice cloning and design, a conversational Agents platform, dubbing, sound effects, and music. Most mistakes come from treating these as one product instead of picking deliberately.
The decision that shapes everything else
Model choice trades latency against fidelity, per request, not per account. Flash-tier models generate audio in under a second and are the only realistic choice when a human is waiting live — an Agent, a voice assistant. Higher-fidelity models sound better and more expressive but cost more time and per character, suiting pre-generated content nobody waits on — narration, IVR prompts, audiobooks. There is no single "best" model.
Voice cloning versus voice design
These get conflated constantly. Cloning reproduces a specific existing voice from an audio sample and needs the speaker's consent. Design generates a wholly synthetic voice from a text description ("a calm, older narrator") with no source recording and no consent question. Cloning where a description would suffice adds a rights-management obligation design avoids entirely.
REST versus WebSocket, and what Agents actually is
A one-shot REST call returns a complete audio file and suits text known upfront. WebSocket streaming is for text arriving token by token from an LLM, letting audio start before the full text exists — what makes sub-second latency possible. REST for streaming text adds a full round trip waiting for the LLM to finish before speech starts.
Conversational Agents are not "TTS plus a loop you write." The platform is a managed product with its own turn-taking, interruption handling, telephony integration, and tool-calling lifecycle — hand-building an equivalent means reimplementing barge-in detection it already solves. Reach for raw TTS/STT outside a live conversation, Agents when it is one.
What surprises people
Billing is metered in characters consumed, not calls or duration; different models consume characters at different rates, so identical scripts cost different amounts by model. Dubbing preserves timing and emotional delivery across languages by design, unlike hand-chaining translation into TTS.
What to verify rather than recall
Exact model names and their latency/quality tiers, character-to-credit ratios, supported
languages, rate limits, and endpoint paths change as ElevenLabs ships new models. Confirm these
against the mirrored corpus under references/vendor/ or the live docs rather than asserting a
remembered model or limit — a stale identifier fails as an API error, but a stale latency
assumption fails silently as a bad experience.