Introducing our Voice Agent API: The fastest path to a working voice agent
Slam-1 is the most powerful prompt-based Speech Language Model that’s customizable to your transcription needs. Improve transcription accuracy for your industry terminology and specific use cases through prompting, without complex custom model development.
JD Prater, Head of Product Marketing
Ryan O'Connor, Senior Developer Educator
%20Blog%20-%20Slam-1.png)
Slam-1 represents a fundamental shift in speech recognition technology
Slam-1 is designed for speech-to-text tasks, combining large language models with specialized audio processing. It aims to improve transcription for specific use cases without the complexity of custom models.
Superior accuracy that humans prefer
Slam-1 outperforms our industry-leading Universal model, with a 72% human preference rating in blind tests, enhancing user satisfaction and engagement.
Why users prefer Slam-1
Key benefits include:
- Accuracy for key entities: Reduces error rates for alphanumerics and specific terminology significantly.
- Formatting accuracy: Enhances readability with proper capitalization, punctuation, and spacing.
- General accuracy: Maintains a low word error rate (WER) across diverse data sets.
Customization and Fine-Tuning
Slam-1 offers simple customization for various industries, requiring minimal effort to achieve high accuracy.
Provide key terms for Slam-1
Users can provide domain-specific terms that enhance recognition and understanding in transcripts.
Applications
Medical
Slam-1 significantly reduces missed entity rates in medical transcription, providing comprehensive records necessary for patient care.
import requests
import time
base_url = "https://api.assemblyai.com"
headers = {"authorization": "<YOUR_API_KEY>"}
# select Slam-1 and specify your audio file
data = {
"audio_url": "https://assembly.ai/sports_injuries.mp3",
"speech_model": "slam-1",
"keyterms_prompt": ["differential diagnosis", "myocardial infarction", "hypertension", "Wellbutrin XL 150mg", "lumbar radiculopathy", "bilateral paresthesia", "metastatic adenocarcinoma", "idiopathic thrombocytopenic purpura"]
}
# submit the transcription request
response = requests.post(base_url + "/v2/transcript", headers=headers, json=data)
Legal
For legal contexts, Slam-1 maintains high accuracy and efficiently handles related terminology.
import requests
import time
base_url = "https://api.assemblyai.com"
headers = {"authorization": "<YOUR_API_KEY>"}
# select Slam-1 and specify your audio file
data = {
"audio_url": "https://assembly.ai/sports_injuries.mp3",
"speech_model": "slam-1",
"keyterms_prompt": ["motion for summary judgment", "voir dire", "amicus curiae", "Duran v. Peabody Coal Company"]
}
# submit the transcription request
response = requests.post(base_url + "/v2/transcript", headers=headers, json=data)
Slam-1’s multi-modal advantage
Slam-1 processes audio and language together, enhancing transcription accuracy compared to traditional methods.
Looking forward
New capabilities are in development, including contextual information incorporation and capturing disfluencies.
Choosing the right model for your needs
- Slam-1: For accuracy and customization with specialized vocabularies.
- Universal: For speed and broader language support.
Get started with Slam-1 today
The public beta is accessible via the API endpoint. Simply include the speech_model parameter with a value of "slam-1".