Skill detail

audiogram

Turns an audio clip into a short captioned video with amplitude-driven waveforms, frequency bars, and social framing.

MatchPossibleReviewed for YouTube Videos
Sourceiart-ai/​youtube-video-skillsExternal source
Reported installs236Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: audiogram
description: This skill should be used when the user asks to "make an audiogram", "turn a podcast into a video", "create a waveform video", "make a podcast clip video", "convert audio to video", "animate a waveform/audio bars", "add a moving waveform to my voiceover", or "make a square/vertical clip for a podcast quote". Covers amplitude-driven waveforms, frequency bars, synced captions, cover-art layout, a progress bar, and 1:1 / 9:16 social framing.
version: 0.1.0
---

# Audiogram

Turn an audio clip — a podcast highlight, a voiceover, a quote — into a short, captioned video that stops the scroll in a muted feed. The moving element (a waveform or equalizer bars) signals "there is sound here," the captions carry the message for the 80% who watch on mute, and cover art + title + a progress bar give it a finished, branded frame.

## When to use

- A podcast soundbite or quote clip for Reels / TikTok / Shorts / feed posts.
- A voiceover or narration that needs a visual so it can post as video.
- Music or any audio where a reactive waveform/bars is the hero.

## The one rule that prevents 90% of bugs

**Drive every bar height from the current frame, never from a real-time analyser loop.** Live `AnalyserNode` + `requestAnimationFrame` reads "what is playing right now" — but a video renderer paints frames out of order and faster/slower than real time, so the wave desyncs or freezes. Instead, decode the whole file to amplitude samples once, then compute the displayed value as a pure function of the frame. In Remotion this is `useAudioData()` + `visualizeAudio()`; in plain canvas it is `decodeAudioData()` into a sample array you index by frame.

```tsx
import { useAudioData, visualizeAudio } from "@remotion/media-utils";
import { useCurrentFrame, useVideoConfig } from "remotion";

const audioData = useAudioData(staticFile("episode.mp3"));
if (!audioData) return null; // still loading
const { fps } = useVideoConfig();
const bars = visualizeAudio({
  audioData,
  frame: useCurrentFrame(),
  fps,
  numberOfSamples: 32, // MUST be a power of two
}); // → Float array, length 32, each 0–1, low freq → high freq
```

## Two ways to render the wave

| Look | Data | API | Best for |
|---|---|---|---|
| Equalizer **bars** | frequency spectrum, values 0–1 | `visualizeAudio()` | music, energy, "feed the bars" |
| Smooth **oscilloscope wave** | time-domain amplitude, −1…1 | `visualizeAudioWaveform()` | voice, podcasts, a calm minimal look |

`visualizeAudio` returns lows on the left, highs on the right. For a centered equalizer, take the first N bars and **mirror** them around the middle so the bass sits in the center. `numberOfSamples` must be a power of two (16/32/64); use 16–32 for a chunky branded look, 64+ for a detailed spectrum. See `references/waveform-render.md` for full bars and oscilloscope components.

## Long files: don't load the whole episode

`useAudioData()` reads the entire file into memory — fine for a 30–90s clip, slow and memory-heavy for a full episode. For anything long, trim first or use `useWindowedAudioData()`, which fetches only the audio around the current frame via HTTP range requests.

```tsx
import { useWindowedAudioData } from "@remotion/media-utils";
const { audioData } = useWindowedAudioData(
  staticFile("full-episode.mp3"),
  fps,
  /* windowInSeconds */ 10,
);
```

Best practice: **trim to the soundbite** (≤90s) before rendering. A tight clip is a better social asset and a lighter render.

## The layout

A finished audiogram is five stacked layers. Keep them in fixed zones so one composition crops cleanly to every aspect.

| Layer | Job | Notes |
|---|---|---|
| Background | brand color / subtle gradient / blurred cover | never busy enough to fight captions |
| Cover art + title | who/what this is | small square cover + episode/show title, top zone |
| Waveform / bars | the motion that signals audio | center band, the hero element |
| Captions | the message, for muted viewers | high contrast
Read the full source on GitHub (opens external page)
Context

Related work