← All repos

View on GitHub ↗Python · 2026-10-06

fd-open-data-mcp

Wire (柏讯) product line · the open-data supply line of FindData — 柏讯线开源核心:开放数据本体与语义供给。

English | 中文

An open-data ontology MCP: a semantic concept layer over multi-datasource financial/economic data. You ask for data by concept + entity (e.g. “price.close of Moutai”, “GDP of China”); the system resolves concepts to physical columns in each source, ranks candidate sources by quality and reachability, fetches from the best one (with failover), caches per concept, and refreshes on a per-concept schedule.

It consumes finddata’s fd-* datasource registry and fd-entities-indicators read-only, and adds the unified layer on top.

One-command install

A single self-contained block that bootstraps the whole finddata open-data stack (hub + all datasource packages + the ontology database). Re-runnable; stops at the first error.

# 1) Install the full stack from PyPI.
#    fd-open-data-protocol comes in transitively; fd-polygon and fd-cn-report
#    self-register via entry points. Drop "[data]" for a light install (MCP
#    server + CLI only, no akshare/yfinance SDKs). "[data]" excludes the
#    browser-rendering stack; add it via "[data,browser]" when
#    scrapling/playwright is needed.
pip install "fd-open-data-mcp[data]" fd-polygon fd-cn-report

# 2) Initialize the ontology DB and wire every layer: catalog -> concepts
#    -> column bindings -> per-source entity ids -> refresh schedules
#    -> manifest.
fd-open-data-mcp migrate \
  && fd-open-data-mcp import-catalog \
  && fd-open-data-mcp consume-concepts \
  && fd-open-data-mcp propose-bindings \
  && fd-open-data-mcp seed-entities \
  && fd-open-data-mcp generate-schedules \
  && fd-open-data-mcp register-discovered

# 3) Start the MCP server (stdio transport, for any MCP client).
fd-open-data-mcp serve

Real-data fetching needs datasource keys in the environment (never commit them): POLYGON_API_KEY, EDGAR_IDENTITY, plus the LLM_* / ES_* vars required by fd-cn-report. See each package’s configuration section.

Architecture

Consumed (read-only)                 Added by fd-open-data-mcp
 fd-akshare/yfinance/world/           concept_bindings     (column -> concept)
 cn-report/cn-gov/polygon/            entity_source_identifiers (per-source ids)
 datacommons                          source_rankings     (quality x reach x freshness)
 fd-entities-indicators               semantic_observations (read-through cache)
   indicator_defs (926 concepts)      fetch_log / schedules / executions / policies
   countries/cities/symbols/sw_industries   entities / relationships (graph)
        │
   Loaders: import_catalog, consume_concepts, propose_bindings,
            seed_entity_identifiers, generate_refresh_schedules
        │
   Runtime: read() -> cache hit? : dispatch (ranked + failover) -> cache -> log
   Search : semantic_search (concepts) + graph_search (entity graph) + ai_search

Eight capabilities (see openspec/changes/add-fd-open-data-mcp/specs/): open-data-catalog, semantic-layer, entity-identity, source-ranking, concept-fetch, scheduled-refresh, entity-graph, vector-search.

Install

cd fd-open-data-mcp
uv sync                  # base install

# full datasource support (akshare, yfinance, edgar, world bank, ...)
uv sync --extra data

The database path defaults to fd_open_data_mcp/metadata/daas.db, overridable via FD_OPEN_DATA_MCP_DATABASE_URL. FINDDATA_ROOT (default: the parent finddata/ directory) locates fd-* providers.

Note: before using SEC EDGAR data, set EDGAR_IDENTITY="your_email@example.com" in the environment.

Quick start

# 1. Create the ontology tables
fd-open-data-mcp migrate

# 2. Import catalogs (akshare 673, yfinance 12, cn-gov 11, cn-report 44, edgar 6, ...)
fd-open-data-mcp import-catalog
# or a single provider: fd-open-data-mcp import-catalog akshare

# 3. Consume the 926 indicator_defs as concepts and propose column->concept bindings
fd-open-data-mcp consume-concepts
fd-open-data-mcp propose-bindings

# 4. Seed per-source entity identifiers (stocks for akshare/yfinance, countries for worldbank)
fd-open-data-mcp seed-entities

# 5. Generate refresh schedules from indicator_defs.frequency
fd-open-data-mcp generate-schedules

# 6. Read by concept + entity (read-through cache + ranked dispatch + failover)
fd-open-data-mcp read --concept-id 234 --entity-type stock --entity-id 1 --date 2024-07-26

MCP server

fd-open-data-mcp serve          # FastMCP, stdio transport

71 tools, registered across six files (49 @mcp.tool in server.py + 22 from the five register_* modules; count verified 2026-10-06 against the live instance’s tools/list = 71):

Group Tools
Catalog/bindings (server.py) import_catalog, register_datasource, update_entity, update_concept, update_binding, consume_concepts, propose_bindings, list_concepts, list_concept_families, list_registry_entries, record_concept_mapping, list_concept_mappings, import_crosswalks, list_bindings, review_bindings, confirm_binding
Entity identity (server.py) seed_entity_identifiers, resolve_entity, add_entity_identifier, list_entities, get_entity, add_entity
Entity graph (server.py) list_relationships, add_relationship, graph_search
Vector/semantic search (server.py) semantic_search, semantic_search_entities, semantic_search_unified, re_embed_concept, ai_search
Ranking/read (server.py) rank_sources, read, read_series, fetch
Search scopes (server.py) scope_create, scope_list, scope_update, scope_delete, scope_bind_caller, scope_unbind_caller, scope_list_bindings, scope_stats
Refresh schedules/indexes (server.py) plan_crawl, generate_refresh_schedules, list_schedules, run_schedule, list_cnreport_rules, enumerate_wbgapi_indicators, register_discovered
Crawl policy control plane (policy_tools.py) policy_create, policy_list, policy_get, policy_update, policy_enable, policy_disable, policy_delete, policy_trigger_now, policy_runs, policy_estimate, data_stats
Crawl platform control (platform_tools.py) platform_sources, platform_runs, platform_trigger, platform_cancel_run, platform_cancel_pending
Crawl visibility (visibility_tools.py) crawl_status
Coverage gaps (coverage_tools.py) coverage_report
Authenticated-crawling identity pool (auth_tools.py) auth_status, auth_events, auth_request_login, auth_launch_login

Catalog expansion and the read boundary

The list_concepts family (MCP tools and the list-concepts CLI) also returns entries from the unified indicator registry (registry_entries table) that passed the naming-quality gate (verified), covering world_bank / gta_panel / china_city_panel / the fd_open_data mirror; yearbook’s numeric-placeholder entries stay unverified and out of the catalog.

Full-registry enumeration (list_registry_entries, v0.5.38)

The parallel tiered-browsing channel (ADR-0002): pages through ALL registered entries — verified and not — ordered by (source_db, native_code), each row carrying the authoritative verified flag, with an optional status filter (verified | registered; default all). The catalog/search/ read verified gate is untouched: list_concepts, ai_search and the read paths still exclude unverified entries. Consumers: the Wire public catalog mirror re-keyed its ids to source_db:native_code from this feed. An optional scope filters rows by the indicator-scope rules (source_db / domain / semantic_code dimensions); unscoped calls return a plain list (the Wire sync contract).

Data sources

The dispatcher run_upstream() routes by source name to adapters. The table below is ordered by verified reachability, not self-reported status.

Verified online ✅

Source Notes
akshare A-share stocks/funds/financials
yfinance Yahoo Finance global equities
cn-report Chinese financial reports (44 tools, see fd-cn-report)
edgar SEC EDGAR filings (needs EDGAR_IDENTITY)
wbgapi World Bank data API
cnstats National Bureau of Statistics data
ckan CKAN catalog data
nbs-gdp NBS GDP data
datacommons Google Data Commons (needs DC_API_KEY)
polygon Polygon.io US equities (external package fd-polygon)
edinet Japan EDINET filings
dartlab Korea DARTLab disclosures

The polygon and datacommons runners live in external fd-* packages, lazily loaded via the manifest’s fetch.module — fd-open-data-mcp does not depend on polygon-api-client / requests unless a fetch actually happens.

Stub adapters ⚠️ (sample data only, no network)

Source Status
cisa-industry stub — static sample rows
amac-fund stub
shfe-metal-futures stub
agriculture stub (DCE)
cme-agricultural-futures stub (CME)
chemicals stub
electronics stub
nonferrous stub
flowers-kifc stub
fin_platforms stub (Wind)
sac-securities stub

These adapters expose run_<source>() entries and manifest registrations, so list-sources self-reports “✅ Full support” — but they return synthetic sample rows, not real exchange or association data. Read the adapter file before depending on any of them.

Read-only catalogs

Source Status
cn-gov manifest registration only (11 ministry catalogs)
world CKAN + Chinese NBS statistical catalogs

Crawl control center (panel + reconciler)

A policy describes what to crawl: concept × entity scope × date range × frequency × mode. CrawlPolicy objects are created on the panel, compiled into CrawlPlans by the reconciler, and executed by scraw-fd-open-data-mcp writing into semantic_observations.

# Start the control panel (default http://0.0.0.0:8000)
FD_OPEN_DATA_MCP_DATABASE_URL=<db url> fd-open-data-mcp panel

# Run the reconciler once (due policies -> launch; close stale runs)
python -m fd_open_data_mcp.refresh.reconciler

Environment variables:

Tests

uv run --with pytest pytest -q

Design notes / v1 limitations

Full specifications: openspec/changes/add-fd-open-data-mcp/.

Proxy pool & circuit breaker

The fetch stack rotates IPs through a proxy pool to avoid source bans, with per-(source, real_source×proxy_ip) circuit breakers. This infrastructure is deployed in the scraw namespace of the remote k8s cluster.

# Manually trigger a proxy sync
kubectl create job --from=cronjob/proxy-pool-sync proxy-pool-sync-manual -n scraw

# Check proxy pool health
kubectl exec -n scraw fd-open-pg-789d56dbb5-fkdbl -- \
  psql -U postgres -d postgres -c "SELECT status, count(*) FROM proxies GROUP BY status;"

Key files: fd_open_data_mcp/proxy/ (selector, breaker, ban rules, injection); real-source failover in fd_open_data_mcp/fetch/dispatch.py.

Standard real-source names

Libraries (akshare, yfinance) call several real sources underneath; the breaker tracks health per real source, not per library, enabling smart failover.

LLM configuration (PDF report extraction)

fd-cn-report uses an LLM to extract financial indicators from annual-report PDFs. The current provider is DeepSeek on Ark (see the fd-cn-report README):

# fd-open-data-mcp/.env
LLM_BASE_URL=…/api/plan/v1     # Ark endpoint
LLM_API_KEY=…                  # Ark key
LLM_MODEL=deepseek-v4-flash

LLM_API_KEY and OPENAI_API_KEY (legacy) are both accepted; LLM_API_KEY wins when both are set.

Image publish

Images are built by GitHub Actions (.github/workflows/image.yml) and pushed to Tencent TCR personal edition: ccr.ccs.tencentyun.com/finddata/fd-open-data-mcp:sha-<short> (+ rolling main). Release ritual: dev pushes go to gitee; publishing = git push github main. The Jenkins→Harbor path is the fallback channel only.

License

MIT