Services
Services
We take on consulting and contracting engagements of any size — from a one-week data pull to multi-year platform builds. Here's what we do best.
Education data collection & web scraping
Large-scale, resilient data collection from the messy reality of the education web: district job boards, policy handbooks, state department portals, and public records.
- Distributed crawling and scraping platforms (hundreds of independent site scrapers, modern anti-bot tooling, tiered proxy management)
- Scheduled collection with automated health monitoring and failure alerting
- Validation against reference datasets such as NCES Common Core Data
- Document acquisition and extraction: PDFs, legacy Office formats, scanned policy documents
Data processing & engineering
Pipelines that turn raw collections into research-ready datasets.
- ETL with Python (Pandas, Polars), PostgreSQL, and async task queues (Celery/Redis)
- Text extraction and normalization at scale (Docling, Unstructured)
- Deduplication, entity resolution, and coverage analysis
- Secure APIs for data delivery and ingestion (Django REST Framework, FastAPI)
Data visualization & reporting
Findings only matter if people can see them.
- Interactive web dashboards with live monitoring and log streaming
- Tableau workbooks and ArcGIS geovisualization
- Publication-quality figures and policy-brief graphics
- Searchable public data portals with sort, filter, and export
Education research & policy analysis
Partner-level research experience across the full project life cycle.
- Study design, survey instruments, and data collection planning
- Statistical and econometric analysis (Stata, R, SAS, SPSS)
- Program evaluation and cost-benefit analysis
- Writing and editorial support for working papers, journal articles, and legislative testimony
AI integration workflows
Applied LLM systems built for research rigor — reproducible, auditable, and cost-controlled.
- Retrieval-augmented generation (RAG) over policy documents and research literature: embeddings, vector search, reranking, grounded chat
- Batch LLM pipelines: multi-model routing, prompt caching, response caching, and structured outputs at hundreds of thousands of queries
- Human-in-the-loop review tooling for verifying machine-collected data
- Speech-to-text, translation, and multilingual workflows
- MCP servers that expose your infrastructure to LLM agents safely
Infrastructure & operations
Everything we build, we can also run.
- Docker/Podman deployment, Nginx, TLS, VPS provisioning with Ansible
- Private networking (Tailscale) for sensitive research data
- Security hardening: dependency auditing, static analysis, CVE remediation
Get in touch to talk about your project.