CAPGov — Public Data Platform
High-performance ETL pipelines turning large public datasets into evidence for policy making.
- Role
- Data Engineering Researcher
- Year
- 2023 — 2026
- Stack
- Python, SQL, Docker, ETL
Nearly three years building the data infrastructure behind evidence-based policy research: pipelines that ingest large public datasets, validate them, and land them somewhere analysts can trust.
- Architected and deployed high-performance ETL pipelines in Python and SQL to automate the processing of large public datasets.
- Modernized legacy infrastructure into a fully containerized Docker microservices environment, guaranteeing dev/prod parity and cutting new-developer onboarding time by 40%.
- Wrote automated data validation that eliminated nearly 100% of manual entry errors feeding analytical dashboards.