← 全部项目

在 GitHub 查看 ↗Python · 2026-10-05

scraw-fd-open-data-mcp

Wire (柏讯) product line · the open-data supply line of FindData — the unified concept-driven crawler

The unified concept-driven crawler for fd-open-data-mcp. Replaces the per-source scraw-* projects. Two modes:

Overview

fd-open-data-mcp plan-crawl emits a CrawlPlan (concepts + entity scope + date range). scraw-fd-open-data-mcp crawl plan.json runs it: for each (concept x entity x date) the adapter registry builds params, run_upstream fetches, and the PG pipeline idempotently upserts into semantic_observations on the remote Postgres. scrapy-redis dedups by a synthesized (source, command, params) request key.

Architecture

See docs/ARCHITECTURE.md.

Quickstart

cd scraw-fd-open-data-mcp
uv venv --python 3.12
uv pip install --python .venv/bin/python -e .
cp .env.example .env  # set FD_OPEN_DATA_MCP_DATABASE_URL, REDIS_URL, SCRAPYD_URL

# crawl (forward): plan then run
fd-open-data-mcp plan-crawl --concept-id 234 --entity-type stock --start 2026-07-01 --end 2026-07-10 -o plan.json
scraw-fd-open-data-mcp crawl plan.json

# migrate (legacy -> semantic_observations)
scraw-fd-open-data-mcp migrate astock --symbol 000001   # sample

Configuration

See docs/CONFIG.md. Env: FD_OPEN_DATA_MCP_DATABASE_URL (canonical store), REDIS_URL, SCRAPYD_URL, JSONL_PATH.

Deploy & Schedule

See docs/DEPLOY.md. ./deploy.sh builds the egg and deploys to scrapyd; python schedule.py concept_crawl --plan plan.json schedules via the scrapyd API.

Tests

.venv/bin/python -m pytest tests/test_smoke.py -v   # import + settings + scrapy list

Project layout

scraw-fd-open-data-mcp/
├── pyproject.toml          # scrapy entry point -> settings; console script
├── scrapy.cfg              # deploy -> scrapyd
├── deploy.sh  schedule.py
├── scraw_fd_open_data_mcp/
│   ├── settings.py         # scrapy-redis + pipelines (PG@300, JSONL@400)
│   ├── items.py            # ObservationItem
│   ├── pipelines.py        # ObservationUpsertPipeline + JsonLinesPipeline
│   ├── db.py               # write_observations -> semantic_observations
│   ├── cli.py              # crawl / migrate
│   └── spiders/concept_crawl_spider.py
├── tests/test_smoke.py
└── docs/

Image publish

Images are built by GitHub Actions (.github/workflows/image.yml) and pushed to Tencent TCR personal edition: ccr.ccs.tencentyun.com/finddata/scraw-fd-open-data-mcp:sha-<short> (+ rolling main). Release ritual: dev pushes go to gitee; publishing = git push github main. The Jenkins→Harbor path is the fallback channel only.