Skill 详情

audiogram

将音频片段转换为带字幕的短视频,包含振幅驱动波形、频率条和社交框架。

匹配类型可能匹配已针对 YouTube 视频 审核
来源iart-ai/​youtube-video-skills外部来源
报告安装量236仅表示受欢迎程度

使用前先检查

自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。

已保存的来源预览

SKILL.md

这段内容是审核时保存的快照。外部来源才是完整且最新的版本。

---
name: audiogram
description: This skill should be used when the user asks to "make an audiogram", "turn a podcast into a video", "create a waveform video", "make a podcast clip video", "convert audio to video", "animate a waveform/audio bars", "add a moving waveform to my voiceover", or "make a square/vertical clip for a podcast quote". Covers amplitude-driven waveforms, frequency bars, synced captions, cover-art layout, a progress bar, and 1:1 / 9:16 social framing.
version: 0.1.0
---

# Audiogram

Turn an audio clip — a podcast highlight, a voiceover, a quote — into a short, captioned video that stops the scroll in a muted feed. The moving element (a waveform or equalizer bars) signals "there is sound here," the captions carry the message for the 80% who watch on mute, and cover art + title + a progress bar give it a finished, branded frame.

## When to use

- A podcast soundbite or quote clip for Reels / TikTok / Shorts / feed posts.
- A voiceover or narration that needs a visual so it can post as video.
- Music or any audio where a reactive waveform/bars is the hero.

## The one rule that prevents 90% of bugs

**Drive every bar height from the current frame, never from a real-time analyser loop.** Live `AnalyserNode` + `requestAnimationFrame` reads "what is playing right now" — but a video renderer paints frames out of order and faster/slower than real time, so the wave desyncs or freezes. Instead, decode the whole file to amplitude samples once, then compute the displayed value as a pure function of the frame. In Remotion this is `useAudioData()` + `visualizeAudio()`; in plain canvas it is `decodeAudioData()` into a sample array you index by frame.

```tsx
import { useAudioData, visualizeAudio } from "@remotion/media-utils";
import { useCurrentFrame, useVideoConfig } from "remotion";

const audioData = useAudioData(staticFile("episode.mp3"));
if (!audioData) return null; // still loading
const { fps } = useVideoConfig();
const bars = visualizeAudio({
  audioData,
  frame: useCurrentFrame(),
  fps,
  numberOfSamples: 32, // MUST be a power of two
}); // → Float array, length 32, each 0–1, low freq → high freq
```

## Two ways to render the wave

| Look | Data | API | Best for |
|---|---|---|---|
| Equalizer **bars** | frequency spectrum, values 0–1 | `visualizeAudio()` | music, energy, "feed the bars" |
| Smooth **oscilloscope wave** | time-domain amplitude, −1…1 | `visualizeAudioWaveform()` | voice, podcasts, a calm minimal look |

`visualizeAudio` returns lows on the left, highs on the right. For a centered equalizer, take the first N bars and **mirror** them around the middle so the bass sits in the center. `numberOfSamples` must be a power of two (16/32/64); use 16–32 for a chunky branded look, 64+ for a detailed spectrum. See `references/waveform-render.md` for full bars and oscilloscope components.

## Long files: don't load the whole episode

`useAudioData()` reads the entire file into memory — fine for a 30–90s clip, slow and memory-heavy for a full episode. For anything long, trim first or use `useWindowedAudioData()`, which fetches only the audio around the current frame via HTTP range requests.

```tsx
import { useWindowedAudioData } from "@remotion/media-utils";
const { audioData } = useWindowedAudioData(
  staticFile("full-episode.mp3"),
  fps,
  /* windowInSeconds */ 10,
);
```

Best practice: **trim to the soundbite** (≤90s) before rendering. A tight clip is a better social asset and a lighter render.

## The layout

A finished audiogram is five stacked layers. Keep them in fixed zones so one composition crops cleanly to every aspect.

| Layer | Job | Notes |
|---|---|---|
| Background | brand color / subtle gradient / blurred cover | never busy enough to fight captions |
| Cover art + title | who/what this is | small square cover + episode/show title, top zone |
| Waveform / bars | the motion that signals audio | center band, the hero element |
| Captions | the message, for muted viewers | high contrast
在 GitHub 阅读完整来源 (打开外部页面)
相关上下文

相关工作