Detalle del Skill
gemini-tts
Genera voz a partir de texto usando modelos TTS de Google Gemini, con soporte para múltiples voces, conversaciones multihablante y streaming.
Revisar antes de usar
La revisión automática comprueba relevancia, no seguridad ni respaldo. Lee las instrucciones de la fuente antes de usar este Skill.
SKILL.md
Este extracto es una copia guardada durante la revisión. La fuente externa contiene la versión completa y actual.
--- name: gemini-tts description: Generate speech from text using Google Gemini TTS models via scripts/. Use for text-to-speech, audio generation, voice synthesis, multi-speaker conversations, and creating audio content. Supports multiple voices and streaming. Triggers on "text to speech", "TTS", "generate audio", "voice synthesis", "speak this text". license: MIT version: 1.0.0 keywords: text-to-speech, TTS, audio generation, voice synthesis, multi-speaker, streaming, Kore, Puck, Charon, Fenrir, Aoede, Zephyr --- # Gemini Text-to-Speech Generate natural-sounding speech from text using Gemini's TTS models through executable scripts with support for multiple voices and multi-speaker conversations. ## When to Use This Skill Use this skill when you need to: - Convert text to natural speech - Create audio for podcasts, audiobooks, or videos - Generate multi-speaker conversations - Stream audio for long content - Choose from multiple voice options - Create accessible audio content - Generate voiceovers for presentations - Batch convert text to audio files ## Available Scripts ### scripts/tts.js **Purpose**: Convert text to speech using Gemini TTS models **When to use**: - Any text-to-speech conversion - Multi-speaker conversation generation - Streaming audio for long texts - Voiceovers for content creation - Accessible audio generation **Key parameters**: | Parameter | Description | Example | |-----------|-------------|---------| | `text` | Text to convert (required) | `"Hello, world!"` | | `--voice`, `-v` | Voice name | `Kore` | | `--output`, `-o` | Base name for output file | `welcome` | | `--output-dir` | Output directory for audio | `audio/` | | `--no-timestamp` | Disable auto timestamp | Flag | | `--model`, `-m` | TTS model | `gemini-2.5-flash-preview-tts` | | `--stream`, `-s` | Enable streaming | Flag | | `--speakers` | Multi-speaker mapping | `"Joe:Kore,Jane:Puck"` | **Output**: WAV audio file path ## Workflows ### Workflow 1: Basic Text-to-Speech ```bash node scripts/tts.js "Hello, world! Have a wonderful day." ``` - Best for: Quick audio generation, simple messages - Voice: `Kore` (default, clear and professional) - Output: `audio/tts_output_YYYYMMDD_HHMMSS.wav` (auto timestamp) ### Workflow 2: Choose Different Voice ```bash node scripts/tts.js "Welcome to our podcast about technology trends" --voice Puck --output welcome ``` - Best for: Friendly, conversational content - Voice options: Kore, Puck, Charon, Fenrir, Aoede, Zephyr, Sulafat - Output: `audio/welcome_YYYYMMDD_HHMMSS.wav` ### Workflow 3: Multi-Speaker Conversation ```bash node scripts/tts.js "TTS the following conversation: Joe: How's it going today? Jane: Not too bad, how about you? Joe: I'm working on a new project. Jane: Sounds exciting, tell me more!" --speakers "Joe:Kore,Jane:Puck" --output conversation ``` - Best for: Dialogues, interviews, role-playing content - Format: Marked conversation with speaker names - Script automatically routes text to appropriate voices - Output: `audio/conversation_YYYYMMDD_HHMMSS.wav` ### Workflow 4: Long Content with Streaming ```bash node scripts/tts.js "This is a very long text that would benefit from streaming..." --stream --output long-form ``` - Best for: Podcasts, audiobooks, long articles - Streaming: Processes audio in chunks for long texts - Output: `audio/long-form_YYYYMMDD_HHMMSS.wav` ### Workflow 5: Professional Voiceover ```bash node scripts/tts.js "Welcome to our quarterly earnings presentation. Today we'll discuss our growth metrics and future plans." --voice Charon --output voiceover ``` - Best for: Corporate content, presentations, formal announcements - Voice: `Charon` (deep, authoritative) - Use when: Professional, serious tone required ### Workflow 6: Custom Output Directory ```bash node scripts/tts.js "Save to specific folder." --output-dir ./my-projects/podcasts/ --output episode1 ``` - Best for: Organized project structures - Directory created automatically if it doesn't exist - Output:Leer la fuente completa en GitHub (abre una página externa)