Detalle del Skill

data-analysis

Decision-oriented dataset analysis with quality checks and business insights.

CoincidenciaDirectaRevisado para análisis de datos
Fuenteakillness/jeo-skillsFuente externa
Instalaciones reportadas547Solo señal de popularidad

Revisar antes de usar

La revisión automática comprueba relevancia, no seguridad ni respaldo. Lee las instrucciones de la fuente antes de usar este Skill.

Vista previa guardada

SKILL.md

Este extracto es una copia guardada durante la revisión. La fuente externa contiene la versión completa y actual.

---
name: data-analysis
description: >-
  Analyze datasets to extract insights, identify patterns, and generate reports.
allowed-tools: Read Grep Glob Bash
metadata:
  tags: data-analysis, csv, sql, notebooks, telemetry, experimentation, analytics
  platforms: Claude, ChatGPT, Gemini, Codex
  version: 2.0
---






# Data Analysis

## When to use this skill
- The user has a **dataset, export, report extract, query result, or shaped event / telemetry table** and wants evidence-backed conclusions.
- The task is to **understand what changed**, compare segments, summarize performance, or explain anomalies in business terms.
- The request mentions **CSV, JSON, SQL tables, retention, cohorts, funnels, conversion, spend, telemetry, event exports, or KPIs**.
- The work needs **data-quality checks before conclusions**.
- The user needs a **concise analysis narrative**, not just raw code snippets.

Do **not** use this skill as the main workflow when:
- The main goal is repeated anomaly or code-pattern scanning across code/data assets → use `pattern-detection`.
- The main goal is building or tuning a specific BI dashboard / Looker Studio + BigQuery workflow → use `looker-studio-bigquery`.
- The task is repository navigation or call-site tracing rather than dataset reasoning → use `codebase-search`.
- The problem is raw log triage / incident reconstruction rather than dataset analysis → use `log-analysis`.

## Core idea
Data analysis is a staged reasoning workflow:
1. clarify the decision question
2. profile the data and trust level
3. choose the cheapest analysis lane that can answer it
4. separate observation from interpretation
5. finish with evidence, caveats, and next actions

Do **not** jump straight into charts or code. The goal is decision-quality analysis.

## Instructions

### Step 1: Frame the analysis question
Before touching the data, define:
- **Decision to support** — what action or judgment depends on this analysis?
- **Primary metric(s)** — conversion, retention, revenue, latency, churn, balance, spend efficiency, etc.
- **Dimensions / segments** — time, channel, cohort, region, plan, device, feature flag, player segment
- **Comparison mode** — before/after, control/treatment, top vs bottom segments, expected vs actual
- **Time window** — day/week/month/release/experiment period

If the request is vague, restate it as:
> "We need to explain [metric/outcome] for [audience] over [time window] and identify the strongest drivers or caveats."

### Step 2: Run a trust check before analysis
Always start with data-quality triage.

#### Minimum trust checklist
- row count / extract size
- schema and types
- missing values / null-heavy columns
- duplicates or repeated IDs
- time range coverage and timezone assumptions
- segment completeness (channels, countries, devices, builds, player groups)
- obvious join / aggregation errors
- outliers or impossible values

Default check pattern:
```python
import pandas as pd

# df = pd.read_csv(...)
print(df.shape)
print(df.dtypes)
print(df.head())
print(df.isna().sum().sort_values(ascending=False).head(15))
print(df.duplicated().sum())
```

If trust is low, stop promising conclusions and explicitly switch the output to:
- what is trustworthy
- what is suspect
- what additional cleanup or data is needed

### Step 3: Choose the analysis lane

| Lane | Use when | Typical tools | What success looks like |
|---|---|---|---|
| Spreadsheet-scale triage | Small extracts, PM/ops handoff, quick KPI sanity checks | Sheets / Excel / quick table review | Fast overview, obvious errors and top movements surfaced |
| SQL slicing | Data already lives in a DB / warehouse or needs grouped filters fast | SQL / DuckDB / warehouse query | Clean aggregates, cohorts, funnels, comparisons |
| Notebook / statistical analysis | Multiple metrics, cohort logic, experiment reasoning, telemetry or richer transformations | pandas / notebooks / scripts | Reproducible calculations and richer interpretation |
| Stakeholder-rea
Leer la fuente completa en GitHub (abre una página externa)
Contexto

Trabajo relacionado