SKILL PROCEDURE

ElevenLabs

Use when generating speech from text, cloning or designing a synthetic voice, transcribing audio, building a real-time conversational voice agent, dubbing or translating audio/video, or generating sound effects or music with the ElevenLabs API — including choosing a model, deciding between REST and WebSocket streaming, and understanding what separates the raw audio APIs from the Agents platform. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.

text-to-speechvoice-aispeech-to-textconversational-aiaudio
BEGINNER GUIDE

Understand ElevenLabs before using it

CATEGORY

ElevenLabs is catalogued under AI and machine learning.

START HERE WHEN

Your work repeatedly involves the concepts tagged above. Open the full procedure below when the current task matches them.

Compare related skills

SKILLCATEGORYSHARED CONCEPTSEXPLANATION
ElevenLabsAI and machine learningCurrent skillUse when generating speech from text, cloning or designing a synthetic voice, transcribing audio, building a real-time conversational voice agent, dubbing or translating audio/video, or generating sound effects or music with the ElevenLabs API — including choosing a model, deciding between REST and WebSocket streaming, and understanding what separates the raw audio APIs from the Agents platform. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.
VapiAI and voice
conversational-aispeech-to-texttext-to-speech
Use when building a phone or web voice AI agent with Vapi — assembling a voice pipeline from transcription, LLM, and voice providers, wiring Twilio/SIP telephony or WebRTC web calls, giving an assistant function-calling tools or a custom-LLM webhook, coordinating multiple assistants with squads, or tuning latency and cost across provider choices. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.
Retell AIAI and voice
conversational-ai
Realtime conversational voice AI platform. Use when building an AI voice agent that speaks in real time — configuring a Retell agent (LLM + transcription + TTS), wiring Twilio/Vonage/SIP telephony or WebRTC web calls, handling function-calling tools and interruption/barge-in, using the realtime WebSocket protocol or SDKs, embedding web calls, or reading call analytics — across Retell Studio and the Retell API.

ElevenLabs

What is HardGraph? HardGraph publishes curated, provenance-backed agent skills grounded in reproducible vendor documentation.

ElevenLabs is a voice AI platform, not a single API: text-to-speech is the core primitive, but the same account also reaches speech-to-text, voice cloning and design, a conversational Agents platform, dubbing, sound effects, and music. Most mistakes come from treating these as one product instead of picking deliberately.

The decision that shapes everything else

Model choice trades latency against fidelity, per request, not per account. Flash-tier models generate audio in under a second and are the only realistic choice when a human is waiting live — an Agent, a voice assistant. Higher-fidelity models sound better and more expressive but cost more time and per character, suiting pre-generated content nobody waits on — narration, IVR prompts, audiobooks. There is no single "best" model.

Voice cloning versus voice design

These get conflated constantly. Cloning reproduces a specific existing voice from an audio sample and needs the speaker's consent. Design generates a wholly synthetic voice from a text description ("a calm, older narrator") with no source recording and no consent question. Cloning where a description would suffice adds a rights-management obligation design avoids entirely.

REST versus WebSocket, and what Agents actually is

A one-shot REST call returns a complete audio file and suits text known upfront. WebSocket streaming is for text arriving token by token from an LLM, letting audio start before the full text exists — what makes sub-second latency possible. REST for streaming text adds a full round trip waiting for the LLM to finish before speech starts.

Conversational Agents are not "TTS plus a loop you write." The platform is a managed product with its own turn-taking, interruption handling, telephony integration, and tool-calling lifecycle — hand-building an equivalent means reimplementing barge-in detection it already solves. Reach for raw TTS/STT outside a live conversation, Agents when it is one.

What surprises people

Billing is metered in characters consumed, not calls or duration; different models consume characters at different rates, so identical scripts cost different amounts by model. Dubbing preserves timing and emotional delivery across languages by design, unlike hand-chaining translation into TTS.

What to verify rather than recall

Exact model names and their latency/quality tiers, character-to-credit ratios, supported languages, rate limits, and endpoint paths change as ElevenLabs ships new models. Confirm these against the mirrored corpus under references/vendor/ or the live docs rather than asserting a remembered model or limit — a stale identifier fails as an API error, but a stale latency assumption fails silently as a bad experience.

References

Hardgraph / curated knowledge for agents.

STATIC EXPORT · CANONICAL SOURCE