Skill-Details

moss-tts-nano-speech

Experten-Skill für MOSS-TTS-Nano, ein 0,1B mehrsprachiges Echtzeit-TTS-Modell, das auf der CPU läuft, mit Voice-Cloning und Streaming-Unterstützung.

ÜbereinstimmungDirektGeprüft für Audio und Stimme
Quellereason-machines/​trending-skillsExterne Quelle
Gemeldete Installationen465Nur Popularitätssignal

Vor Nutzung prüfen

Die automatische Prüfung bewertet Relevanz, nicht Sicherheit oder Empfehlung. Lies vor der Nutzung die Quellanweisungen.

Gespeicherte Quellvorschau

SKILL.md

Dieser Auszug wurde bei der Prüfung gespeichert. Die externe Quelle enthält die vollständige und aktuelle Version.

---
name: moss-tts-nano-speech
description: Expert skill for using MOSS-TTS-Nano, a 0.1B parameter multilingual real-time TTS model that runs on CPU with voice cloning support.
triggers:
  - generate speech with MOSS TTS
  - text to speech with voice cloning
  - moss tts nano inference
  - run MOSS-TTS-Nano locally
  - multilingual TTS CPU inference
  - clone voice with MOSS
  - streaming audio generation python
  - tiny TTS model deployment
---

# MOSS-TTS-Nano Speech Generation Skill

> Skill by [ara.so](https://ara.so) — Daily 2026 Skills collection.

MOSS-TTS-Nano is an open-source multilingual tiny TTS model (0.1B parameters) from MOSI.AI and the OpenMOSS team. It uses an Audio Tokenizer + LLM autoregressive pipeline to generate 48 kHz stereo speech in real time, supports 20 languages, voice cloning, streaming inference, and runs on CPU without a GPU.

## Installation

### Conda (recommended)

```bash
conda create -n moss-tts-nano python=3.12 -y
conda activate moss-tts-nano

git clone https://github.com/OpenMOSS/MOSS-TTS-Nano.git
cd MOSS-TTS-Nano

pip install -r requirements.txt
pip install -e .
```

### Fix WeTextProcessing if it fails

```bash
conda install -c conda-forge pynini=2.1.6.post1 -y
pip install git+https://github.com/WhizZest/WeTextProcessing.git
```

After `pip install -e .` the `moss-tts-nano` CLI command is available in the active environment.

## Model Weights

Models are auto-downloaded from Hugging Face on first run:
- TTS model: `OpenMOSS-Team/MOSS-TTS-Nano`
- Audio tokenizer: `OpenMOSS-Team/MOSS-Audio-Tokenizer-Nano`

ModelScope mirrors are available at `openmoss/MOSS-TTS-Nano` and `openmoss/MOSS-Audio-Tokenizer-Nano`.

## CLI Commands

### Generate speech (voice clone mode)

```bash
moss-tts-nano generate \
  --prompt-speech assets/audio/zh_1.wav \
  --text "欢迎关注模思智能、上海创智学院与复旦大学自然语言处理实验室。"
```

Output defaults to `generated_audio/moss_tts_nano_output.wav`.

### Generate from a text file (long-form)

```bash
moss-tts-nano generate \
  --prompt-speech assets/audio/zh_1.wav \
  --text-file my_script.txt \
  --output output.wav
```

### Launch local web demo

```bash
moss-tts-nano serve
# or directly:
python app.py
```

Opens at `http://127.0.0.1:18083` — model stays loaded in memory for fast repeated requests.

### Direct Python entrypoint

```bash
python infer.py \
  --prompt-audio-path assets/audio/zh_1.wav \
  --text "Hello, this is a test of MOSS-TTS-Nano."
```

Output: `generated_audio/infer_output.wav`

## Python API Usage

### Basic voice clone inference

```python
from infer import MossTTSNanoInference

# Initialize once (downloads weights on first run)
tts = MossTTSNanoInference()

# Voice clone: synthesize text in the style of the reference audio
audio = tts.infer(
    text="欢迎使用MOSS语音合成系统。",
    prompt_audio_path="assets/audio/zh_1.wav",
)

# Save output
import soundfile as sf
sf.write("output.wav", audio, samplerate=48000)
```

### English voice clone

```python
from infer import MossTTSNanoInference

tts = MossTTSNanoInference()

audio = tts.infer(
    text="Welcome to MOSS TTS Nano, a tiny but capable text to speech model.",
    prompt_audio_path="assets/audio/en_sample.wav",
)

import soundfile as sf
sf.write("english_output.wav", audio, samplerate=48000)
```

### Streaming inference (low latency)

```python
from infer import MossTTSNanoInference
import soundfile as sf
import numpy as np

tts = MossTTSNanoInference()

chunks = []
for audio_chunk in tts.infer_stream(
    text="This sentence is generated chunk by chunk for low latency playback.",
    prompt_audio_path="assets/audio/en_sample.wav",
):
    chunks.append(audio_chunk)
    # process or play chunk in real time here

full_audio = np.concatenate(chunks)
sf.write("streamed_output.wav", full_audio, samplerate=48000)
```

### Long-text synthesis with chunked voice cloning

```python
from infer import MossTTSNanoInference

tts = MossTTSNanoInference()

long_text = """
MOSS-TTS-Nano supports long-form synthesis through automatic chunkin
Vollständige Quelle auf GitHub lesen (öffnet externe Seite)
Kontext

Verwandte Arbeit