Introducing our Voice Agent API: The fastest path to a working voice agent

Slam-1 now in public beta: the most powerful prompt-based Speech Language Model to unlock real world outcomes

Slam-1 is the most powerful prompt-based Speech Language Model that’s customizable to your transcription needs. Improve transcription accuracy for your industry terminology and specific use cases through prompting, without complex custom model development.

JD Prater, Head of Product Marketing
Ryan O'Connor, Senior Developer Educator

%20Blog%20-%20Slam-1.png)

Slam-1 represents a fundamental shift in speech recognition technology

Slam-1 is designed for speech-to-text tasks, combining large language models with specialized audio processing. It aims to improve transcription for specific use cases without the complexity of custom models.

Superior accuracy that humans prefer

Slam-1 outperforms our industry-leading Universal model, with a 72% human preference rating in blind tests, enhancing user satisfaction and engagement.

Why users prefer Slam-1

Key benefits include:

  • Accuracy for key entities: Reduces error rates for alphanumerics and specific terminology significantly.
  • Formatting accuracy: Enhances readability with proper capitalization, punctuation, and spacing.
  • General accuracy: Maintains a low word error rate (WER) across diverse data sets.

Customization and Fine-Tuning

Slam-1 offers simple customization for various industries, requiring minimal effort to achieve high accuracy.

Provide key terms for Slam-1

Users can provide domain-specific terms that enhance recognition and understanding in transcripts.

Applications

Medical

Slam-1 significantly reduces missed entity rates in medical transcription, providing comprehensive records necessary for patient care.

import requests
import time

base_url = "https://api.assemblyai.com"
headers = {"authorization": "<YOUR_API_KEY>"}

# select Slam-1 and specify your audio file
data = {
    "audio_url": "https://assembly.ai/sports_injuries.mp3",
    "speech_model": "slam-1",
    "keyterms_prompt": ["differential diagnosis", "myocardial infarction", "hypertension", "Wellbutrin XL 150mg", "lumbar radiculopathy", "bilateral paresthesia", "metastatic adenocarcinoma", "idiopathic thrombocytopenic purpura"]
}
# submit the transcription request
response = requests.post(base_url + "/v2/transcript", headers=headers, json=data)

Legal

For legal contexts, Slam-1 maintains high accuracy and efficiently handles related terminology.

import requests
import time

base_url = "https://api.assemblyai.com"
headers = {"authorization": "<YOUR_API_KEY>"}

# select Slam-1 and specify your audio file
data = {
    "audio_url": "https://assembly.ai/sports_injuries.mp3",
    "speech_model": "slam-1",
    "keyterms_prompt": ["motion for summary judgment", "voir dire", "amicus curiae", "Duran v. Peabody Coal Company"]
}
# submit the transcription request
response = requests.post(base_url + "/v2/transcript", headers=headers, json=data)

Slam-1’s multi-modal advantage

Slam-1 processes audio and language together, enhancing transcription accuracy compared to traditional methods.

Looking forward

New capabilities are in development, including contextual information incorporation and capturing disfluencies.

Choosing the right model for your needs

  • Slam-1: For accuracy and customization with specialized vocabularies.
  • Universal: For speed and broader language support.

Get started with Slam-1 today

The public beta is accessible via the API endpoint. Simply include the speech_model parameter with a value of "slam-1".