Skill 详情

construction-sweep

Builds and enriches databases of construction builders and contractors.

匹配类型直接匹配已针对 建筑施工 审核
来源lotfb86/sharedskillsluke外部来源
报告安装量1仅表示受欢迎程度

使用前先检查

自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。

已保存的来源预览

SKILL.md

这段内容是审核时保存的快照。外部来源才是完整且最新的版本。

---
name: construction-sweep
description: County-by-county scrape of construction builders (GCs, home builders, remodelers). Dedups across counties, classifies vs trades/suppliers, enriches via Firecrawl websites + Facebook pages. Outputs one master CSV per state.
triggers:
  - "/construction-sweep"
  - "sweep construction"
  - "scrape builders"
---

# construction-sweep

**Purpose:** Build a state-level database of construction builders by running a repeatable per-county pipeline: Apify Google Maps scrape → cross-county dedup → rules classifier → Firecrawl + FB enrichment → append to master CSV.

**First time on this machine?** Read `SETUP.md` in this skill dir FIRST. Verify prereqs before invoking.

## Invocation

User types one of:
- `/construction-sweep <State>` — pick next undone county, run full pipeline
- `/construction-sweep <State> <County>` — run specific county
- `/construction-sweep <State> --status` — show progress, stats, next county
- `/construction-sweep <State> --bootstrap` — initialize state folder (idempotent)
- `/construction-sweep <State> --dry-run` — confirm next county + est cost, no API calls

State folder defaults to `./<State>/` in cwd. Override with `--dir <path>`.

## Pipeline (per county)

Each step writes a checkpoint file so a crash doesn't restart from scratch. Resume by re-running same county — skill detects completed stages and skips.

| # | Stage | Script | Input | Output | Wait |
|---|---|---|---|---|---|
| 1 | Bootstrap state | `scripts/bootstrap_state.py <state> <dir>` | state name | master.csv + seen.json + processed.txt + counties.json | — |
| 2 | Pick next county | `scripts/pick_next_county.py <dir>` | state dir | county name | — |
| 3 | Confirm w/ user | (inline) | — | yes/no | user |
| 4 | Scrape GMaps | Apify MCP `compass/crawler-google-places` async | county + state | runId | 5-15 min poll |
| 5 | Pull dataset | Apify API `/v2/datasets/<id>/items` | datasetId | `<county>_raw.json` | — |
| 6 | Dedup county | `scripts/dedup_county.py <county> <raw> <dir>` | raw.json | `<county>_new.json` + updated seen.json + master backfills | — |
| 7 | Classify | `scripts/classify_county.py <county> <dir>` | new.json | `<county>_classified.json` | — |
| 8 | Prep enrichment | `scripts/prep_enrichment.py <county> <dir>` | classified.json | `<county>_fc_urls.json` + `<county>_fb_input.json` | — |
| 9a | FC batch submit | `scripts/submit_firecrawl.py <county> <dir>` | fc_urls.json | FC batchId | 3-8 min poll |
| 9b | FB batch submit | Apify MCP `apify/facebook-pages-scraper` async | fb_input.json | Apify runId | 3-8 min poll |
| 10a | FC pull | `scripts/pull_firecrawl.py <county> <batchId> <dir>` | batchId | `<county>_fc.json` | — |
| 10b | FB pull | Apify API dataset pull | runId | `<county>_fb.json` | — |
| 11 | Append | `scripts/append_county.py <county> <dir>` | all above | master.csv row per KEEP + seen.json update | — |
| 12 | Log | append to `processed.txt` | — | — | — |

## Credentials required

- **Firecrawl API key**: `~/.firecrawl_key` (chmod 600). Get from firecrawl.dev.
- **Apify**: MCP server configured in `~/.claude.json`. Skill uses `mcp__Apify__call-actor` + public `/v2/actor-runs/<id>` endpoint (no token needed for status polls).

If either missing, abort + tell user what to install. SETUP.md has exact commands.

## Sleep/poll pattern

Use `ScheduleWakeup` tool for polling:
- After Apify scrape launch → sleep 240s, then poll `/v2/actor-runs/<runId>` for SUCCEEDED
- After FC + FB parallel launch → sleep 270s (cache-warm window), poll both
- Repeat with shorter sleeps (60-120s) if still RUNNING

**Do NOT poll in a tight loop.** Each poll burns context. Sleep ≥60s between checks.

## Error handling

| Symptom | Fix |
|---|---|
| Apify scrape fails or 0 items | Check input, retry once. If still fails, flag county + skip. |
| Firecrawl "we do not support this site" | URL in blocklist. Check `data/blocklists.json`, strip bad URL, resubmit. |
| FC pagination returns <expected | Bump `max_page
在 GitHub 阅读完整来源 (打开外部页面)
相关上下文

相关工作