Skill detail
modelslab-audio-generation
Generates speech, music, and sound effects using ModelsLab's v7 Voice API, supporting TTS, STT, voice conversion, and dubbing.
Inspect before use
Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.
SKILL.md
The saved excerpt is a snapshot from review. The external source remains the complete and most current version.
---
name: modelslab-audio-generation
description: Generate speech, music, and sound effects using ModelsLab's v7 Voice API. Supports text-to-speech, speech-to-text, speech-to-speech, music generation, sound effects, dubbing, song extension, and song inpainting via ElevenLabs and Inworld models.
---
# ModelsLab Audio Generation
Generate high-quality audio including speech, music, voice conversion, sound effects, and dubbing using AI.
## When to Use This Skill
- Convert text to natural-sounding speech (TTS)
- Transcribe speech to text
- Transform voice characteristics (speech-to-speech)
- Generate music from text prompts
- Create sound effects
- Dub audio into different languages
- Extend or inpaint songs
- Build voice assistants or audiobooks
## Available APIs (v7)
### Voice Endpoints
- **Text to Speech**: `POST https://modelslab.com/api/v7/voice/text-to-speech`
- **Speech to Text**: `POST https://modelslab.com/api/v7/voice/speech-to-text`
- **Speech to Speech**: `POST https://modelslab.com/api/v7/voice/speech-to-speech`
- **Music Generation**: `POST https://modelslab.com/api/v7/voice/music-gen`
- **Sound Generation**: `POST https://modelslab.com/api/v7/voice/sound-generation`
- **Create Dubbing**: `POST https://modelslab.com/api/v7/voice/create-dubbing`
- **Song Extender**: `POST https://modelslab.com/api/v7/voice/song-extender`
- **Song Inpaint**: `POST https://modelslab.com/api/v7/voice/song-inpaint`
- **Fetch Result**: `POST https://modelslab.com/api/v7/voice/fetch/{id}`
> **Note**: v6 endpoints (`/api/v6/voice/text_to_speech`, etc.) still work but v7 is the current version. Parameter names have changed in v7 (e.g., `text` is now `prompt`, `audio` is now `init_audio`).
## Discovering Audio Models
```bash
# Search audio/voice models
modelslab models search --feature audio_gen
# Search by provider
modelslab models search --search "eleven"
# Get model details
modelslab models detail --id eleven_multilingual_v2
```
## Audio Model IDs
| model_id | Name | Use With |
|----------|------|----------|
| `eleven_multilingual_v2` | ElevenLabs Multilingual v2 | text-to-speech |
| `eleven_english_sts_v2` | ElevenLabs Voice Changer | speech-to-speech |
| `scribe_v1` | ElevenLabs Scribe | speech-to-text |
| `eleven_sound_effect` | ElevenLabs Sound Effects | sound-generation |
| `music_v1` | ElevenLabs Music | music-gen |
| `inworld-tts-1` | Inworld TTS | text-to-speech |
## Text to Speech
```python
import requests
import time
def text_to_speech(text, api_key, voice_id="21m00Tcm4TlvDq8ikWAM", model_id="eleven_multilingual_v2"):
"""Convert text to speech.
Args:
text: The text to convert to speech
api_key: Your ModelsLab API key
voice_id: ElevenLabs voice ID (see Available Voices below)
model_id: TTS model to use
"""
response = requests.post(
"https://modelslab.com/api/v7/voice/text-to-speech",
json={
"key": api_key,
"prompt": text, # v7 uses "prompt" not "text"
"voice_id": voice_id,
"model_id": model_id
}
)
data = response.json()
if data["status"] == "success":
return data["output"][0]
elif data["status"] == "processing":
return poll_audio_result(data["id"], api_key)
else:
raise Exception(f"Error: {data.get('message', 'Unknown error')}")
# Usage
audio_url = text_to_speech(
"Hello! Welcome to ModelsLab. This is a test of our text-to-speech API.",
"your_api_key"
)
print(f"Audio URL: {audio_url}")
```
## Speech to Text (Transcription)
```python
def speech_to_text(audio_url, api_key, model_id="scribe_v1"):
"""Transcribe speech from audio to text.
Args:
audio_url: URL of audio file (must be publicly accessible)
model_id: STT model to use
"""
response = requests.post(
"https://modelslab.com/api/v7/voice/speech-to-text",
json={
"key": api_key,
"init_audio": audio_url, # v7Read the full source on GitHub (opens external page)