Skill 详情

deepgram-js-speech-to-text

在 JavaScript/TypeScript 中使用 Deepgram Speech-to-Text v1,通过 REST 和 WebSocket 进行预录音频和实时音频转录。

声明的前提(自述): Node.js

匹配类型直接匹配已针对 音频与语音 审核
来源deepgram/​deepgram-js-sdk外部来源
报告安装量69仅表示受欢迎程度

使用前先检查

自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。

已保存的来源预览

SKILL.md

这段内容是审核时保存的快照。外部来源才是完整且最新的版本。

---
name: deepgram-js-speech-to-text
description: Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (`/v1/listen`) for prerecorded or live audio transcription. Covers `client.listen.v1.media.transcribeUrl` / `transcribeFile` (REST) plus `client.listen.v1.createConnection()` / `connect()` (WebSocket). Use `deepgram-js-audio-intelligence` for summarize/sentiment/topics/diarize overlays, `deepgram-js-conversational-stt` for Flux turn-taking on `/v2/listen`, and `deepgram-js-voice-agent` for full-duplex assistants. Triggers include "transcribe", "speech to text", "STT", "listen.v1", "nova-3", "live transcription", and "websocket transcription".
---

# Using Deepgram Speech-to-Text (JavaScript / TypeScript SDK)

Basic transcription for prerecorded audio (REST) or live audio (WebSocket) via `/v1/listen`.

## When to use this product

- **REST (`client.listen.v1.media.transcribeUrl` / `transcribeFile`)** — one-shot transcription of a finished URL or file. Good for batch jobs, caption generation, offline processing.
- **WebSocket (`client.listen.v1.createConnection()` / `connect()`)** — continuous streaming transcription. Good for live captions, microphone audio, telephony streams, browser or Node realtime apps.

**Use a different skill when:**
- You also want summaries, topics, intents, sentiment, language detection, or redaction guidance on the same `/v1/listen` call → `deepgram-js-audio-intelligence`.
- You need Flux turn-taking and end-of-turn events on `/v2/listen` → `deepgram-js-conversational-stt`.
- You need a full interactive assistant with STT + LLM + TTS over one socket → `deepgram-js-voice-agent`.

## Authentication

```js
require("dotenv").config();

const { DeepgramClient } = require("@deepgram/sdk");

const deepgramClient = new DeepgramClient({
  apiKey: process.env.DEEPGRAM_API_KEY,
});
```

Use the exported `DeepgramClient` from `src/CustomClient.ts`, not `DefaultDeepgramClient`. The wrapper adds the required `Token` auth prefix, session headers, and patched WebSocket behavior.

## Quick start — REST (prerecorded URL)

From `examples/04-transcription-prerecorded-url.ts`:

```js
const data = await deepgramClient.listen.v1.media.transcribeUrl({
  url: "https://dpgr.am/spacewalk.wav",
  model: "nova-3",
  language: "en",
  punctuate: true,
  paragraphs: true,
  utterances: true,
});

console.log(
  "Transcription:",
  data.results?.channels?.[0]?.alternatives?.[0]?.transcript,
);
```

## Quick start — REST (prerecorded file)

From `examples/05-transcription-prerecorded-file.ts`:

```js
const { createReadStream } = require("fs");

const data = await deepgramClient.listen.v1.media.transcribeFile(
  createReadStream("./examples/spacewalk.wav"),
  {
    model: "nova-3",
    language: "en",
    punctuate: true,
    paragraphs: true,
    utterances: true,
    smart_format: true,
  }
);
```

`transcribeFile(...)` accepts multiple upload shapes in this SDK: `fs.ReadStream`, `Buffer`, `ReadableStream`, `Blob`, `File`, `ArrayBuffer`, and `Uint8Array` (see `examples/23-file-upload-types.ts`).

## Quick start — WebSocket (live streaming)

From `examples/07-transcription-live-websocket.ts`:

```js
const deepgramConnection = await deepgramClient.listen.v1.createConnection({
  model: "nova-3",
  language: "en",
  punctuate: "true",
  interim_results: "true",
});

deepgramConnection.on("message", (data) => {
  if (data.type === "Results") {
    console.log("Transcript:", data);
  }
});

deepgramConnection.connect();
await deepgramConnection.waitForOpen();

// Swap this for a mic capture (e.g. `node-microphone` / `MediaRecorder`)
// in real apps; the repo examples use `createReadStream` over a sample WAV.
const { createReadStream } = require("node:fs");
const audioStream = createReadStream("samples/spacewalk.wav");

audioStream.on("data", (chunk) => {
  deepgramConnection.sendMedia(chunk);
});

audioStream.on("end", () => {
  deepgramConnection.sendFinalize({ type: "Finalize" });
});
`
在 GitHub 阅读完整来源 (打开外部页面)
相关上下文

相关工作