Skill 詳細
data-analysis
General Excel/CSV exploration, SQL analysis, summaries, and exports.
使用前に確認
自動レビューは関連性のみを確認し、安全性や推奨を保証しません。使用前に出典の説明を読んでください。
SKILL.md
これはレビュー時に保存された抜粋です。完全で最新の内容は外部ソースを確認してください。
--- name: data-analysis description: Use this skill when the user uploads Excel (.xlsx/.xls) or CSV files and wants to perform data analysis, generate statistics, create summaries, pivot tables, SQL queries, or any form of structured data exploration. Supports multi-sheet Excel workbooks, aggregation, filtering, joins, and exporting results to CSV/JSON/Markdown. --- # Data Analysis Skill ## Overview This skill analyzes user-uploaded Excel/CSV files using DuckDB — an in-process analytical SQL engine. It supports schema inspection, SQL-based querying, statistical summaries, and result export, all through a single Python script. ## Core Capabilities - Inspect Excel/CSV file structure (sheets, columns, types, row counts) - Execute arbitrary SQL queries against uploaded data - Generate statistical summaries (mean, median, stddev, percentiles, nulls) - Support multi-sheet Excel workbooks (each sheet becomes a table) - Export query results to CSV, JSON, or Markdown - Handle large files efficiently with DuckDB's columnar engine ## Workflow ### Step 1: Understand Requirements When a user uploads data files and requests analysis, identify: - **File location**: Path(s) to uploaded Excel/CSV files under `/mnt/user-data/uploads/` - **Analysis goal**: What insights the user wants (summary, filtering, aggregation, comparison, etc.) - **Output format**: How results should be presented (table, CSV export, JSON, etc.) - You don't need to check the folder under `/mnt/user-data` ### Step 2: Inspect File Structure First, inspect the uploaded file to understand its schema: ```bash python /mnt/skills/public/data-analysis/scripts/analyze.py \ --files /mnt/user-data/uploads/data.xlsx \ --action inspect ``` This returns: - Sheet names (for Excel) or filename (for CSV) - Column names, data types, and non-null counts - Row count per sheet/file - Sample data (first 5 rows) ### Step 3: Perform Analysis Based on the schema, construct SQL queries to answer the user's questions. #### Run SQL Query ```bash python /mnt/skills/public/data-analysis/scripts/analyze.py \ --files /mnt/user-data/uploads/data.xlsx \ --action query \ --sql "SELECT category, COUNT(*) as count, AVG(amount) as avg_amount FROM Sheet1 GROUP BY category ORDER BY count DESC" ``` #### Generate Statistical Summary ```bash python /mnt/skills/public/data-analysis/scripts/analyze.py \ --files /mnt/user-data/uploads/data.xlsx \ --action summary \ --table Sheet1 ``` This returns for each numeric column: count, mean, std, min, 25%, 50%, 75%, max, null_count. For string columns: count, unique, top value, frequency, null_count. #### Export Results ```bash python /mnt/skills/public/data-analysis/scripts/analyze.py \ --files /mnt/user-data/uploads/data.xlsx \ --action query \ --sql "SELECT * FROM Sheet1 WHERE amount > 1000" \ --output-file /mnt/user-data/outputs/filtered-results.csv ``` Supported output formats (auto-detected from extension): - `.csv` — Comma-separated values - `.json` — JSON array of records - `.md` — Markdown table ### Parameters | Parameter | Required | Description | |-----------|----------|-------------| | `--files` | Yes | Space-separated paths to Excel/CSV files | | `--action` | Yes | One of: `inspect`, `query`, `summary` | | `--sql` | For `query` | SQL query to execute | | `--table` | For `summary` | Table/sheet name to summarize | | `--output-file` | No | Path to export results (CSV/JSON/MD) | > [!NOTE] > Do NOT read the Python file, just call it with the parameters. ## Table Naming Rules - **Excel files**: Each sheet becomes a table named after the sheet (e.g., `Sheet1`, `Sales`, `Revenue`) - **CSV files**: Table name is the filename without extension (e.g., `data.csv` → `data`) - **Multiple files**: All tables from all files are available in the same query context, enabling cross-file joins - **Special characters**: Sheet/file names with spaces or special characters are auto-sanitized (spaces → underscores). Use double quotesGitHub で全文を読む (外部ページ)