Services

We take on consulting and contracting engagements of any size — from a one-week data pull to multi-year platform builds. Here's what we do best.

Education data collection & web scraping

Large-scale, resilient data collection from the messy reality of the education web: district job boards, policy handbooks, state department portals, and public records.

  • Distributed crawling and scraping platforms (hundreds of independent site scrapers, modern anti-bot tooling, tiered proxy management)
  • Scheduled collection with automated health monitoring and failure alerting
  • Validation against reference datasets such as NCES Common Core Data
  • Document acquisition and extraction: PDFs, legacy Office formats, scanned policy documents

Data processing & engineering

Pipelines that turn raw collections into research-ready datasets.

  • ETL with Python (Pandas, Polars), PostgreSQL, and async task queues (Celery/Redis)
  • Text extraction and normalization at scale (Docling, Unstructured)
  • Deduplication, entity resolution, and coverage analysis
  • Secure APIs for data delivery and ingestion (Django REST Framework, FastAPI)

Data visualization & reporting

Findings only matter if people can see them.

  • Interactive web dashboards with live monitoring and log streaming
  • Tableau workbooks and ArcGIS geovisualization
  • Publication-quality figures and policy-brief graphics
  • Searchable public data portals with sort, filter, and export

Education research & policy analysis

Partner-level research experience across the full project life cycle.

  • Study design, survey instruments, and data collection planning
  • Statistical and econometric analysis (Stata, R, SAS, SPSS)
  • Program evaluation and cost-benefit analysis
  • Writing and editorial support for working papers, journal articles, and legislative testimony

AI integration workflows

Applied LLM systems built for research rigor — reproducible, auditable, and cost-controlled.

  • Retrieval-augmented generation (RAG) over policy documents and research literature: embeddings, vector search, reranking, grounded chat
  • Batch LLM pipelines: multi-model routing, prompt caching, response caching, and structured outputs at hundreds of thousands of queries
  • Human-in-the-loop review tooling for verifying machine-collected data
  • Speech-to-text, translation, and multilingual workflows
  • MCP servers that expose your infrastructure to LLM agents safely

Infrastructure & operations

Everything we build, we can also run.

  • Docker/Podman deployment, Nginx, TLS, VPS provisioning with Ansible
  • Private networking (Tailscale) for sensitive research data
  • Security hardening: dependency auditing, static analysis, CVE remediation

Get in touch to talk about your project.