SKILL PROCEDURE

Vapi

Use when building a phone or web voice AI agent with Vapi — assembling a voice pipeline from transcription, LLM, and voice providers, wiring Twilio/SIP telephony or WebRTC web calls, giving an assistant function-calling tools or a custom-LLM webhook, coordinating multiple assistants with squads, or tuning latency and cost across provider choices. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.

voice-agentsconversational-aitelephonyspeech-to-texttext-to-speechfunction-calling
BEGINNER GUIDE

Understand Vapi before using it

CATEGORY

Vapi is catalogued under AI and voice.

START HERE WHEN

Your work repeatedly involves the concepts tagged above. Open the full procedure below when the current task matches them.

Compare related skills

SKILLCATEGORYSHARED CONCEPTSEXPLANATION
VapiAI and voiceCurrent skillUse when building a phone or web voice AI agent with Vapi — assembling a voice pipeline from transcription, LLM, and voice providers, wiring Twilio/SIP telephony or WebRTC web calls, giving an assistant function-calling tools or a custom-LLM webhook, coordinating multiple assistants with squads, or tuning latency and cost across provider choices. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.
Retell AIAI and voice
voice-agentsconversational-aitelephony
Realtime conversational voice AI platform. Use when building an AI voice agent that speaks in real time — configuring a Retell agent (LLM + transcription + TTS), wiring Twilio/Vonage/SIP telephony or WebRTC web calls, handling function-calling tools and interruption/barge-in, using the realtime WebSocket protocol or SDKs, embedding web calls, or reading call analytics — across Retell Studio and the Retell API.
ElevenLabsAI and machine learning
text-to-speechspeech-to-textconversational-ai
Use when generating speech from text, cloning or designing a synthetic voice, transcribing audio, building a real-time conversational voice agent, dubbing or translating audio/video, or generating sound effects or music with the ElevenLabs API — including choosing a model, deciding between REST and WebSocket streaming, and understanding what separates the raw audio APIs from the Agents platform. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.

Vapi

What is HardGraph? HardGraph publishes curated, provenance-backed agent skills grounded in reproducible vendor documentation.

Vapi is not a voice model — it is an orchestration layer that composes a phone or web voice assistant out of three independently swappable providers: a transcriber (speech-to-text), a model (the LLM reasoning), and a voice (text-to-speech). Vapi owns the realtime plumbing between them — streaming partials into the LLM before the caller finishes speaking, streaming LLM tokens into TTS before the sentence finishes, endpointing, interruption handling — so the assistant feels responsive despite three separate calls per turn.

The pipeline decision that actually matters

Because transcriber, model, and voice are independently selectable, the real decision isn't "which vendor" but which combination sits on the latency/quality/cost frontier. Vapi's Model Intelligence presets (balanced, low-latency, reasoning, low-cost) exist because this tradeoff shifts as providers ship new versions — a fast transcriber paired with a slow reasoning model still produces a laggy assistant, since voice-to-voice latency sums every stage, not just the slowest. OpenAI Realtime's speech-to-speech mode is the escape hatch: it collapses all three into one provider, trading composability for lower latency.

Managed pipeline vs custom LLM

By default Vapi calls the configured model directly. A custom LLM points Vapi at a webhook you host instead — the same shape as Retell's bring-your-own-LLM mode — trading managed reasoning for control over context, retrieval, and tool execution, which also shifts who owns retry behavior.

What differs from Retell, and from a from-scratch build

Vapi and Retell solve the same problem — streaming STT→LLM→TTS with interruption handling — but Vapi leans further into per-call provider composability and multi-assistant squads, where specialized assistants hand off a live call instead of one assistant handling everything. Building this from scratch means owning audio streaming, endpointing, and barge-in logic yourself; both platforms remove that layer, at the cost of depending on their infrastructure during an outage.

Tools and call control

Assistants call tools mid-conversation via streaming function calling — the LLM can speak a transitional phrase while a slow tool call is in flight, instead of stalling the turn. Call control (transfer, end call, hold) is itself exposed as tools; webhooks fire on lifecycle events, the seam for reconciling call state externally instead of polling.

What to verify rather than recall

Supported transcriber/model/voice providers, exact webhook event and tool schema names, squad handoff configuration, and Model Intelligence metric definitions change as Vapi adds providers. Confirm these against the mirrored corpus under references/vendor/ or the live docs rather than asserting a remembered provider list or event name.

References

Hardgraph / curated knowledge for agents.

STATIC EXPORT · CANONICAL SOURCE