Skill detail

modelslab-audio-generation

Generates speech, music, and sound effects using ModelsLab's v7 Voice API, supporting TTS, STT, voice conversion, and dubbing.

MatchDirectReviewed for Audio and Voice
Sourcemodelslab/​skillsExternal source
Reported installs63Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: modelslab-audio-generation
description: Generate speech, music, and sound effects using ModelsLab's v7 Voice API. Supports text-to-speech, speech-to-text, speech-to-speech, music generation, sound effects, dubbing, song extension, and song inpainting via ElevenLabs and Inworld models.
---

# ModelsLab Audio Generation

Generate high-quality audio including speech, music, voice conversion, sound effects, and dubbing using AI.

## When to Use This Skill

- Convert text to natural-sounding speech (TTS)
- Transcribe speech to text
- Transform voice characteristics (speech-to-speech)
- Generate music from text prompts
- Create sound effects
- Dub audio into different languages
- Extend or inpaint songs
- Build voice assistants or audiobooks

## Available APIs (v7)

### Voice Endpoints
- **Text to Speech**: `POST https://modelslab.com/api/v7/voice/text-to-speech`
- **Speech to Text**: `POST https://modelslab.com/api/v7/voice/speech-to-text`
- **Speech to Speech**: `POST https://modelslab.com/api/v7/voice/speech-to-speech`
- **Music Generation**: `POST https://modelslab.com/api/v7/voice/music-gen`
- **Sound Generation**: `POST https://modelslab.com/api/v7/voice/sound-generation`
- **Create Dubbing**: `POST https://modelslab.com/api/v7/voice/create-dubbing`
- **Song Extender**: `POST https://modelslab.com/api/v7/voice/song-extender`
- **Song Inpaint**: `POST https://modelslab.com/api/v7/voice/song-inpaint`
- **Fetch Result**: `POST https://modelslab.com/api/v7/voice/fetch/{id}`

> **Note**: v6 endpoints (`/api/v6/voice/text_to_speech`, etc.) still work but v7 is the current version. Parameter names have changed in v7 (e.g., `text` is now `prompt`, `audio` is now `init_audio`).

## Discovering Audio Models

```bash
# Search audio/voice models
modelslab models search --feature audio_gen

# Search by provider
modelslab models search --search "eleven"

# Get model details
modelslab models detail --id eleven_multilingual_v2
```

## Audio Model IDs

| model_id | Name | Use With |
|----------|------|----------|
| `eleven_multilingual_v2` | ElevenLabs Multilingual v2 | text-to-speech |
| `eleven_english_sts_v2` | ElevenLabs Voice Changer | speech-to-speech |
| `scribe_v1` | ElevenLabs Scribe | speech-to-text |
| `eleven_sound_effect` | ElevenLabs Sound Effects | sound-generation |
| `music_v1` | ElevenLabs Music | music-gen |
| `inworld-tts-1` | Inworld TTS | text-to-speech |

## Text to Speech

```python
import requests
import time

def text_to_speech(text, api_key, voice_id="21m00Tcm4TlvDq8ikWAM", model_id="eleven_multilingual_v2"):
    """Convert text to speech.

    Args:
        text: The text to convert to speech
        api_key: Your ModelsLab API key
        voice_id: ElevenLabs voice ID (see Available Voices below)
        model_id: TTS model to use
    """
    response = requests.post(
        "https://modelslab.com/api/v7/voice/text-to-speech",
        json={
            "key": api_key,
            "prompt": text,             # v7 uses "prompt" not "text"
            "voice_id": voice_id,
            "model_id": model_id
        }
    )

    data = response.json()

    if data["status"] == "success":
        return data["output"][0]
    elif data["status"] == "processing":
        return poll_audio_result(data["id"], api_key)
    else:
        raise Exception(f"Error: {data.get('message', 'Unknown error')}")

# Usage
audio_url = text_to_speech(
    "Hello! Welcome to ModelsLab. This is a test of our text-to-speech API.",
    "your_api_key"
)
print(f"Audio URL: {audio_url}")
```

## Speech to Text (Transcription)

```python
def speech_to_text(audio_url, api_key, model_id="scribe_v1"):
    """Transcribe speech from audio to text.

    Args:
        audio_url: URL of audio file (must be publicly accessible)
        model_id: STT model to use
    """
    response = requests.post(
        "https://modelslab.com/api/v7/voice/speech-to-text",
        json={
            "key": api_key,
            "init_audio": audio_url,    # v7
Read the full source on GitHub (opens external page)
Context

Related work