azure-speech
Expert knowledge for Azure Speech in Foundry Tools development including troubleshooting, best practices, decision making, limits & quotas, security, configuration, integrations & coding patterns, and deployment. Use when using STT/TTS, custom voice or avatars, Speech containers, Voice Live, or telephony/WebRTC streaming, and other Azure Speech in Foundry Tools related development tasks. Not for Azure Content Understanding in Foundry Tools (use azure-content-understanding), Azure AI Vision (use azure-ai-vision), Azure AI Video Indexer (use azure-video-indexer), Azure Translator (use azure-translator).
How do I install this agent skill?
npx skills add https://github.com/microsoftdocs/agent-skills --skill azure-speechIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill is a documentation reference for Azure Speech services. It provides a structured index of official Microsoft documentation and instructions for the agent to fetch updated content from trusted Microsoft domains.
- Socketpass
No alerts
- Snykwarn
Risk: MEDIUM · 1 issue
- Runlayerpass
1/1 file flagged
- ZeroLeakspass
Score: 93/100 · 2 sections analyzed
What does this agent skill do?
Azure Speech in Foundry Tools Skill
This skill provides expert guidance for Azure Speech in Foundry Tools. Covers troubleshooting, best practices, decision making, limits & quotas, security, configuration, integrations & coding patterns, and deployment. It combines local quick-reference content with remote documentation fetching capabilities.
How to Use This Skill
IMPORTANT for Agent: Use the Category Index below to locate relevant sections. For categories with line ranges (e.g.,
L35-L120), useread_filewith the specified lines. For categories with file links (e.g.,[security.md](security.md)), useread_fileon the linked reference file
IMPORTANT for Agent: If
metadata.generated_atis more than 3 months old, suggest the user pull the latest version from the repository. Ifmcp_microsoftdocstools are not available, suggest the user install it: Installation Guide
This skill requires network access to fetch documentation content:
- Preferred: Use
mcp_microsoftdocs:microsoft_docs_fetchwith query stringfrom=learn-agent-skill. Returns Markdown. - Fallback: Use
fetch_webpagewith query stringfrom=learn-agent-skill&accept=text/markdown. Returns Markdown.
Category Index
| Category | Lines | Description |
|---|---|---|
| Troubleshooting | L36-L43 | Diagnosing and fixing common Azure Speech issues across TTS, STT, SDK, containers, CRL compatibility, and retrieving session/transcription IDs for support. |
| Best Practices | L44-L61 | Best practices for audio/video prep, custom voice/avatar training, latency and memory tuning, accuracy boosts (phrases/keywords), reliability (CRL, backups), and Voice Live handling/evaluation |
| Decision Making | L62-L75 | Guidance on choosing access methods and devices, and step-by-step migration paths between Speech/Custom Voice/STT/TTS APIs, Long Audio, and retired intent recognition features. |
| Limits & Quotas | L76-L82 | Managing custom speech/voice models and endpoints, plus quotas, capacity limits, and scaling constraints for Azure Speech workloads. |
| Security | L83-L94 | Securing Azure AI Speech: auth (Entra, RBAC), network isolation (VNet, Private Link, sovereign clouds), encryption/BYOK, BYOS storage, and consent/compliance for personal/professional voice. |
| Configuration | L95-L128 | Configuring Azure Speech behavior: recognition, TTS, avatars, containers, logging, storage, SSML, audio inputs, language/diarization, and Voice Live/SDK runtime and tracing options. |
| Integrations & Coding Patterns | L129-L164 | Patterns and APIs for integrating Azure Speech and Voice Live with apps, agents, telephony, and WebRTC, including STT, TTS, avatars, translation, function calling, and streaming events. |
| Deployment | L165-L175 | Deploying and running Azure Speech services (STT, TTS, language ID) via containers, Kubernetes/Helm, and batch APIs, including custom models and on-premises setups. |
Troubleshooting
| Topic | URL |
|---|---|
| Retrieve Speech to text session and transcription IDs for support | https://learn.microsoft.com/en-us/azure/ai-services/speech-service/how-to-get-speech-session-id |
| Resolve common Azure Speech in Foundry issues | https://learn.microsoft.com/en-us/azure/ai-services/speech-service/known-issues |
| Troubleshoot Azure Speech containers deployment issues | https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-container-faq |
| Diagnose and fix common Azure Speech SDK issues | https://learn.microsoft.com/en-us/azure/ai-services/speech-service/troubleshooting |
Best Practices
Decision Making
Limits & Quotas
| Topic | URL |
|---|---|
| Manage custom speech model and endpoint lifecycle | https://learn.microsoft.com/en-us/azure/ai-services/speech-service/how-to-custom-speech-model-and-endpoint-lifecycle |
| Custom voice endpoint limits for Speech service | https://learn.microsoft.com/en-us/azure/ai-services/speech-service/professional-voice-deploy-endpoint |
| Review quotas and limits for Azure Speech workloads | https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-services-quotas-and-limits |
Security
Configuration
Integrations & Coding Patterns
Deployment
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/microsoftdocs/agent-skills/azure-speech">View azure-speech on skillZs</a>