Skill detail

datalineage-bigquery-asset-impact-analysis

Performs a narrow BigQuery lineage impact analysis, not general data analysis.

MatchPossibleReviewed for data analysis
Sourcegoogle/skillsExternal source
Reported installs2,016Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: datalineage-bigquery-asset-impact-analysis
metadata:
  category: BigDataAndAnalytics
description: >-
  Analyzes the downstream impact (blast radius) when a BigQuery table or view is broken, stale, or modified.
  Identifies all downstream tables, dashboards, and processes that will be affected.
  Use when:
  - Performing a blast radius or impact analysis for a BigQuery table or view.
  - Assessing the consequences of modifying, deleting, or pausing updates to a BigQuery asset.
  - Identifying downstream dependencies (tables, dashboards, processes) of a BigQuery asset.
  Don't use for:
  - General BigQuery querying or data analysis (use BigQuery-related tools instead).
  - Non-BigQuery assets (e.g., Cloud Storage files) unless they are part of the BigQuery lineage.
  - Creating or modifying lineage links directly.
---

# BigQuery Asset Impact Analysis

This skill guides the agent in performing a downstream impact analysis (blast
radius assessment) when a BigQuery table or view is reported as broken, stale,
missing, or when a user is planning maintenance and wants to know the
consequences of modifying or pausing updates to an asset.

It relies primarily on the **Google Cloud Data Lineage (Knowledge Catalog) MCP Server**
to discover relationships between assets.

## Prerequisites

This skill requires access to the Google Cloud Data Lineage API and an active
client connection to the Data Lineage MCP Server. For detailed connection
configurations and tool schemas, refer to [MCP Usage](references/mcp-usage.md).

## Analysis Workflow

### 1. Resolve the Asset's Fully Qualified Name (FQN)

*   Ensure you have the correct FQN format for the BigQuery asset:
    *   *Format:* `bigquery:{project_id}.{dataset_id}.{table_or_view_id}`
    *   *Example:* `bigquery:my-prod-project.analytics.orders`


### 2. Determine Locations and Parent Path

Identify the locations to search and construct the Data Lineage API request:

*   **Discover Asset Location**: Run the command `bq show --format=json
    {project_id}:{dataset_id}` and extract the `location` field (e.g.,
    `us-central1` or `us`). If location discovery fails due to permissions or
    missing tools, prompt the user for the dataset's location.
*   **Set Parent Path**: Set the `parent` path using the project ID and the
    MCP server's location. Consult the `DataLineageServer` tool definition
    to find the configured region or location (e.g., `us`). The format is:
    `projects/{project_id}/locations/{mcp_server_location}`.
*   **Configure Search Scope**: Include the discovered asset location in the
    `locations` array of the payload (e.g., `["us-central1"]` or `["us",
    "us-central1"]`).

### 3. Retrieve the Downstream Lineage Graph

Call the `DataLineageServer:search_lineage` tool to fetch downstream
relationships.

*   **Direction**: Set to `DOWNSTREAM`.
*   **Search Parameters**: Use `max_depth = 10` and `max_process_per_link = 5`
    as robust defaults.

### 4. Identify the Blast Radius

Traverse the returned lineage links to build the impact graph:

*   **Affected Assets**: The `target` of each link represents a downstream asset
    that depends on your source asset.
*   **Transform Processes**: Inspect the `processes` field on each link. This
    identifies the ETL pipelines, BigQuery Views, or Scheduled Queries that
    propagate the data.
*   **Direct vs. Indirect Impact**:
    *   **Direct Impact (Depth 1)**: Assets directly consuming the source asset.
        If a link has `dependency_type: EXACT_COPY`, mark the target as
        "Directly Stale / Identical Copy".
    *   **Indirect Impact (Depth > 1)**: Assets further down the stream that
        will experience cascading stale data or failures.

### 5. Summarize and Format the Output

Present your findings clearly to the user using the following structure:

1.  **Executive Summary**: State the total number of downstream assets affected
    and the maximum depth of the impact.
2.  **Critical Path**: Highlight hig
Read the full source on GitHub (opens external page)
Context

Related work