Skill detail

deepgram-js-speech-to-text

Uses Deepgram Speech-to-Text v1 in JavaScript/TypeScript for prerecorded and live audio transcription via REST and WebSocket.

Needs (stated): Node.js

MatchDirectReviewed for Audio and Voice
Sourcedeepgram/​deepgram-js-sdkExternal source
Reported installs69Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: deepgram-js-speech-to-text
description: Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (`/v1/listen`) for prerecorded or live audio transcription. Covers `client.listen.v1.media.transcribeUrl` / `transcribeFile` (REST) plus `client.listen.v1.createConnection()` / `connect()` (WebSocket). Use `deepgram-js-audio-intelligence` for summarize/sentiment/topics/diarize overlays, `deepgram-js-conversational-stt` for Flux turn-taking on `/v2/listen`, and `deepgram-js-voice-agent` for full-duplex assistants. Triggers include "transcribe", "speech to text", "STT", "listen.v1", "nova-3", "live transcription", and "websocket transcription".
---

# Using Deepgram Speech-to-Text (JavaScript / TypeScript SDK)

Basic transcription for prerecorded audio (REST) or live audio (WebSocket) via `/v1/listen`.

## When to use this product

- **REST (`client.listen.v1.media.transcribeUrl` / `transcribeFile`)** — one-shot transcription of a finished URL or file. Good for batch jobs, caption generation, offline processing.
- **WebSocket (`client.listen.v1.createConnection()` / `connect()`)** — continuous streaming transcription. Good for live captions, microphone audio, telephony streams, browser or Node realtime apps.

**Use a different skill when:**
- You also want summaries, topics, intents, sentiment, language detection, or redaction guidance on the same `/v1/listen` call → `deepgram-js-audio-intelligence`.
- You need Flux turn-taking and end-of-turn events on `/v2/listen` → `deepgram-js-conversational-stt`.
- You need a full interactive assistant with STT + LLM + TTS over one socket → `deepgram-js-voice-agent`.

## Authentication

```js
require("dotenv").config();

const { DeepgramClient } = require("@deepgram/sdk");

const deepgramClient = new DeepgramClient({
  apiKey: process.env.DEEPGRAM_API_KEY,
});
```

Use the exported `DeepgramClient` from `src/CustomClient.ts`, not `DefaultDeepgramClient`. The wrapper adds the required `Token` auth prefix, session headers, and patched WebSocket behavior.

## Quick start — REST (prerecorded URL)

From `examples/04-transcription-prerecorded-url.ts`:

```js
const data = await deepgramClient.listen.v1.media.transcribeUrl({
  url: "https://dpgr.am/spacewalk.wav",
  model: "nova-3",
  language: "en",
  punctuate: true,
  paragraphs: true,
  utterances: true,
});

console.log(
  "Transcription:",
  data.results?.channels?.[0]?.alternatives?.[0]?.transcript,
);
```

## Quick start — REST (prerecorded file)

From `examples/05-transcription-prerecorded-file.ts`:

```js
const { createReadStream } = require("fs");

const data = await deepgramClient.listen.v1.media.transcribeFile(
  createReadStream("./examples/spacewalk.wav"),
  {
    model: "nova-3",
    language: "en",
    punctuate: true,
    paragraphs: true,
    utterances: true,
    smart_format: true,
  }
);
```

`transcribeFile(...)` accepts multiple upload shapes in this SDK: `fs.ReadStream`, `Buffer`, `ReadableStream`, `Blob`, `File`, `ArrayBuffer`, and `Uint8Array` (see `examples/23-file-upload-types.ts`).

## Quick start — WebSocket (live streaming)

From `examples/07-transcription-live-websocket.ts`:

```js
const deepgramConnection = await deepgramClient.listen.v1.createConnection({
  model: "nova-3",
  language: "en",
  punctuate: "true",
  interim_results: "true",
});

deepgramConnection.on("message", (data) => {
  if (data.type === "Results") {
    console.log("Transcript:", data);
  }
});

deepgramConnection.connect();
await deepgramConnection.waitForOpen();

// Swap this for a mic capture (e.g. `node-microphone` / `MediaRecorder`)
// in real apps; the repo examples use `createReadStream` over a sample WAV.
const { createReadStream } = require("node:fs");
const audioStream = createReadStream("samples/spacewalk.wav");

audioStream.on("data", (chunk) => {
  deepgramConnection.sendMedia(chunk);
});

audioStream.on("end", () => {
  deepgramConnection.sendFinalize({ type: "Finalize" });
});
`
Read the full source on GitHub (opens external page)
Context

Related work