sec-edgar-pipeline
SEC EDGAR extraction pipeline: setup, filing discovery by CIK, recipe-driven extraction, and report generation.
How do I install this agent skill?
npx skills add https://github.com/bobmatnyc/claude-mpm-skills --skill sec-edgar-pipelineIs this agent skill safe to install?
- Gen Agent Trust Hubwarn
This skill provides a pipeline for SEC EDGAR data extraction. It generates and executes code dynamically based on data patterns and processes external financial filings, which introduces a surface for indirect instructions and requires caution when running generated scripts.
- Socketwarn
1 alert: gptAnomaly
- Snykwarn
Risk: MEDIUM · 1 issue
- Runlayerwarn
2/2 files flagged
What does this agent skill do?
SEC EDGAR Pipeline
Overview
This pipeline is centered on edgar-analyzer and the EDGAR data sources. The core loop is: configure credentials, create a project with examples, analyze patterns, generate code, run extraction, and export reports.
Setup (Keys + User Agent)
Use the setup wizard to configure required keys:
python -m edgar_analyzer setup
# or
edgar-analyzer setup
Required entries:
OPENROUTER_API_KEY- (Optional)
JINA_API_KEY EDGARuser agent string ("Name email@example.com")
End-to-End CLI Workflow
# 1. Create project
edgar-analyzer project create my_project --template minimal
# 2. Add examples + project.yaml
# projects/my_project/examples/*.json
# 3. Analyze examples
edgar-analyzer analyze-project projects/my_project
# 4. Generate extraction code
edgar-analyzer generate-code projects/my_project
# 5. Run extraction
edgar-analyzer run-extraction projects/my_project --output-format csv
Outputs land in projects/<name>/output/.
EDGAR-Specific Conventions
- CIK values are 10-digit, zero-padded (e.g.,
0000320193). - Rate limit: SEC API allows 10 requests/sec. Scripts use ~0.11s delays.
- User agent is mandatory; include name + email.
Scripted Example (Apple DEF 14A)
edgar/scripts/fetch_apple_def14a.py shows the direct flow:
- Fetch latest DEF 14A metadata
- Download HTML
- Parse Summary Compensation Table (SCT)
- Save raw HTML + extracted JSON + ground truth
Recipe-Driven Extraction
edgar/recipes/sct_extraction/config.yaml defines a multi-step pipeline:
- Fetch DEF 14A filings by company list
- Extract SCT tables with
SCTAdapter - Validate with
sct_validator - Write results to
output/sct
Report Generation
edgar/scripts/create_csv_reports.py converts JSON results into:
executive_compensation_<timestamp>.csvtop_25_executives_<timestamp>.csvcompany_summary_<timestamp>.csv
Troubleshooting
- No filings found: confirm CIK formatting and filing type (DEF 14A vs DEF 14A/A).
- API errors: slow down requests and confirm user-agent is set.
- Extraction errors: regenerate code or use manual ground truth in POC scripts.
Related Skills
universal/data/reporting-pipelinestoolchains/python/testing/pytest
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/bobmatnyc/claude-mpm-skills/sec-edgar-pipeline">View sec-edgar-pipeline on skillZs</a>