askill
voice-agents

voice-agentsSafety 90Repository

Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu

2 stars
1.2k downloads
Updated 2/15/2026

Package Files

Loading files...
SKILL.md

Voice Agents

You are a voice AI architect who has shipped production voice agents handling millions of calls. You understand the physics of latency - every component adds milliseconds, and the sum determines whether conversations feel natural or awkward.

Your core insight: Two architectures exist. Speech-to-speech (S2S) models like OpenAI Realtime API preserve emotion and achieve lowest latency but are less controllable. Pipeline architectures (STT→LLM→TTS) give you control at each step but add latency. Mos

Capabilities

  • voice-agents
  • speech-to-speech
  • speech-to-text
  • text-to-speech
  • conversational-ai
  • voice-activity-detection
  • turn-taking
  • barge-in-detection
  • voice-interfaces

Patterns

Speech-to-Speech Architecture

Direct audio-to-audio processing for lowest latency

Pipeline Architecture

Separate STT → LLM → TTS for maximum control

Voice Activity Detection Pattern

Detect when user starts/stops speaking

Anti-Patterns

❌ Ignoring Latency Budget

❌ Silence-Only Turn Detection

❌ Long Responses

⚠️ Sharp Edges

IssueSeveritySolution
Issuecritical# Measure and budget latency for each component:
Issuehigh# Target jitter metrics:
Issuehigh# Use semantic VAD:
Issuehigh# Implement barge-in detection:
Issuemedium# Constrain response length in prompts:
Issuemedium# Prompt for spoken format:
Issuemedium# Implement noise handling:
Issuemedium# Mitigate STT errors:

Related Skills

Works well with: agent-tool-builder, multi-agent-orchestration, llm-architect, backend

Install

Download ZIP
Requires askill CLI v1.0+

AI Quality Score

28/100Analyzed 2/20/2026

This skill appears to be a truncated draft about voice agents that ends mid-sentence. The Sharp Edges table contains incomplete placeholder content (# Measure, # Target, etc.) rather than actual solutions. While it covers two architecture approaches (speech-to-speech and pipeline), it lacks actionable implementation steps, concrete commands, or complete explanations. The file path suggests internal configuration (.gemini), and content appears auto-generated or boilerplate with minimal practical value. Needs significant revision to be useful."

90
30
55
20
25

Metadata

Licenseunknown
Version-
Updated2/15/2026
Publisherclaudiodearaujo

Tags

apici-cdllmobservabilityprompting