# assemblyai.com > AI-optimized mirror of assemblyai.com containing 50 pages totalling 55,476 words of clean markdown content, structured data, and semantic HTML. Original source: https://assemblyai.com/. Last updated: 2026-04-30T22:30:27.161Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [The best way to build Voice AI apps](/site-root.html): With AssemblyAI's industry-leading Speech AI models, transcribe speech to text and extract insights from your voice data. (1,075 words) ## Articles & Blog Posts - [blog/llm-use-cases/index.html](/blog/llm-use-cases/index.html) (1 words) - [blog/top-tools-for-live-transcription/index.html](/blog/top-tools-for-live-transcription/index.html) (1 words) - [pricing/index.html](/pricing/index.html) (1 words) - [blog/best-audio-file-formats-for-speech-to-text/index.html](/blog/best-audio-file-formats-for-speech-to-text/index.html) (1 words) - [docs/api-reference/streaming-api/universal-streaming/universal-streaming-mdx.html](/docs/api-reference/streaming-api/universal-streaming/universal-streaming-mdx.html) (1 words) - [blog/introducing-universal-streaming/index.html](/blog/introducing-universal-streaming/index.html) (1 words) - [docs/streaming/whisper-streaming/index.html](/docs/streaming/whisper-streaming/index.html) (1 words) - [docs/streaming/universal-streaming/index.html](/docs/streaming/universal-streaming/index.html) (1 words) - [docs/streaming/universal-3-pro/prompting/index.html](/docs/streaming/universal-3-pro/prompting/index.html) (1 words) - [docs/streaming/universal-streaming/message-sequence-md.html](/docs/streaming/universal-streaming/message-sequence-md.html) (1 words) - [Overview](/docs/integrations/n-8-n/index.html): Integrate AssemblyAI with 1000+ apps and services using n8n's automation platform. (2,290 words) - [legal/terms-of-service/index.html](/legal/terms-of-service/index.html) (1 words) - [blog/best-real-time-speech-to-text-apps/index.html](/blog/best-real-time-speech-to-text-apps/index.html) (1 words) - [docs/streaming/label-speakers-and-separate-channels/index.html](/docs/streaming/label-speakers-and-separate-channels/index.html) (1 words) - [Are there language-specific models for better accuracy?](/blog/speech-to-text-accuracy/index.html): Speech to text accuracy depends on audio quality, accents, models, and setup. Learn how WER works and improve transcription results for real teams today. (1,818 words) - [How DALL-E 2 Actually Works](/blog/how-dall-e-2-actually-works/index.html): How does OpenAI's groundbreaking DALL-E 2 model actually work? Check out this detailed guide to learn the ins and outs of DALL-E 2. (3,335 words) - [Overview](/docs/pre-recorded-audio/getting-started/transcribe-an-audio-file/index.html): Learn how to transcribe and analyze an audio file. (847 words) - [How to evaluate speech recognition models](/blog/how-to-evaluate-speech-recognition-models/index.html): Learn how to evaluate speech-to-text models beyond Word Error Rate. This guide covers Semantic WER, Missed Entity Rate, ground truth correction, and practical benchmarking frameworks for 2026. (3,227 words) - [Reinforcement Learning With (Deep) Q-Learning Explained](/blog/reinforcement-learning-with-deep-q-learning-explained.html): In this video, we learn about Reinforcement Learning and (Deep) Q-Learning. (1,269 words) - [products/streaming-speech-to-text/index.html](/products/streaming-speech-to-text/index.html) (1 words) - [2025 INSIGHTS REPORT](/2025-research-report/index.html): Industry survey reveals conversation intelligence adoption rates, accuracy challenges, and generative AI integration strategies. Essential insights for speech AI decision-makers. (1,461 words) - [blog/word-error-rate-is-broken/index.html](/blog/word-error-rate-is-broken/index.html) (1 words) - [blog/top-speaker-diarization-libraries-and-apis/index.html](/blog/top-speaker-diarization-libraries-and-apis/index.html) (1 words) - [legal/business-associate-agreement/index.html](/legal/business-associate-agreement/index.html) (1 words) - [legal/data-processing-addendum/index.html](/legal/data-processing-addendum/index.html) (1 words) - [blog/what-is-speaker-diarization-and-how-does-it-work/index.html](/blog/what-is-speaker-diarization-and-how-does-it-work/index.html) (1 words) - [Best medical speech recognition software and APIs in 2026](/blog/best-medical-speech-recognition-software-and-apis/index.html): Compare 8 leading medical speech recognition solutions, APIs vs software (2,758 words) - [What kinds of businesses use automatic transcription?](/blog/corporate-transcription-services/index.html): Corporate transcription services deliver accurate, secure transcripts for meetings, calls, and legal needs with enterprise-level compliance and fast turnaround. (1,610 words) - [Authentication](/docs/api-reference/transcripts/submit/index.html): > For the complete documentation index, see llms.txt(https://www.assemblyai.com/docs/llms.txt) (2,565 words) - [Best AI playgrounds in 2026](/blog/6-best-ai-playgrounds/index.html): If you’re looking for ways to play around with AI in 2026, these are the best AI playgrounds to try. (2,149 words) - [Gemini 3 Pro vs GPT-5 vs Claude 4.5: Which model wins for audio workflows?](/blog/gemini-3-pro-vs-gpt-5-vs-claude-4-5/index.html): Gemini 3 Pro brings smarter summaries and actionable insights to audio workflows. Compare it to GPT-5, Claude 4.5, and other leading LLMs. (925 words) - [Decoding Strategies: How LLMs Choose The Next Word](/blog/decoding-strategies-how-llms-choose-the-next-word/index.html): Large Language Models are trained to guess the next word. But when generating text, the combination of their probability estimates with algorithms known as decoding strategies is what determines how they actually choose words. Learn how decoding strategies work in this article. (4,019 words) - [7 best conversation intelligence software in 2026](/blog/conversation-intelligence-software/index.html): Learn about the best conversation intelligence platforms for 2026. You'll learn what each does best, where they fall short, and how to pick the right solution. (3,009 words) - [Authenticate with a temporary token](/docs/streaming/authenticate-with-a-temporary-token/index.html) (284 words) - [Inside AssemblyAI's NYC voice agents January 2026 meetup: Production insights from the front lines](/blog/nyc-voice-agents-meetup-january-2026/index.html): 100+ voice AI builders gathered at AssemblyAI's NYC office to discuss production insights, where voice agents break, and 2026 predictions. (1,151 words) - [Word accuracy and error rate](/docs/pre-recorded-audio/benchmarks/index.html): Review the latest benchmarks and performance metrics for AssemblyAI's pre-recorded Speech-to-text models. (419 words) - [Advancing and democratizing Voice AI technology for the world](/about/index.html): Our vision is to create new, superhuman Speech AI models that will unlock entirely new classes of applications and products to be built leveraging voice data. (541 words) - [Voice Agent Orchestrators Compared: Vapi vs Pipecat vs LiveKit with AssemblyAI](/blog/vapi-vs-pipecat-vs-livekit/index.html): Vapi vs Pipecat vs LiveKit: Compare architecture, control, transport, pricing, and setup tradeoffs to choose the right platform for your voice AI use case. (2,762 words) - [Top 8 open source STT options for voice applications in 2026](/blog/top-open-source-stt-options-for-voice-applications/index.html): This comprehensive comparison examines eight open source STT solutions, analyzing their technical capabilities, implementation requirements, and ideal use cases to help you build voice applications from scratch. (3,164 words) - [Edge Routing](/docs/streaming/endpoints-and-data-zones/index.html) (332 words) - [Project instructions](/docs/coding-agent-prompts/index.html): Give your AI coding agent accurate, up-to-date knowledge of AssemblyAI's APIs and SDKs. (593 words) - [Introducing our Voice Agent API: The fastest path to a working voice agent](/blog/slam-1-public-beta/index.html): Improve transcription accuracy for your industry terminology and specific use cases through prompting, without complex custom model development. (2,389 words) - [Quickstart](/docs/integrations/recall/index.html) (935 words) - [Self-Hosted Voice AI](/deployments/self-hosted-voice-ai/index.html): AssemblyAI's industry-leading speech AI is available to deploy on your own infrastructure, in the cloud or on-premises. (438 words) - [Service Level Agreement](/legal/service-level-agreement/index.html) (581 words) - [Speech Understanding Models](/docs/speech-understanding/getting-started/index.html): Extract structured insights from audio with Speech Understanding models. (158 words) - [Blog](/blog/index.html): Explore the latest in AI-powered speech-to-text technology with AssemblyAI. 'Speech & Text' delivers expert insights, industry trends, and deep dives into automatic speech recognition (ASR), transcription, and audio intelligence. (9,222 words) - [How are word/transcript level confidence scores calculated? | AssemblyAI | Documentation](/docs/faq/how-are-word-transcript-level-confidence-scores-calculated.html) (122 words) - [Page Not Found](/docs/streaming/diarization-and-multichannel-md.html) (9 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/content/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/content/robots.txt): Crawler directives