Staff Engineer, Data PlatformThe RoleWe're looking for a
Staff Engineer, Data Platform to own one of the most important technical problems at NationGraph:
turning the outside world's fragmented government information into a proprietary data advantage.This is not a traditional data engineering role focused on maintaining a warehouse or internal analytics.
You'll own the technical ecosystem that:
- Discovers external data
- Acquires it reliably
- Understands and extracts information from it
- Normalizes and connects it
- Validates its quality
- Makes it available to NationGraph's products and models
The scope starts with more than 110,000 independent state and local government agencies, but extends to federal data, Canada, and eventually public-sector information globally.
You'll work across:
- Data engineering
- Distributed systems
- Information retrieval
- Data modeling
- LLMs and agents
- Applied ML
- Entity resolution
- Knowledge graphs
You'll partner closely with Product, ML Research, and Infrastructure to determine both:
- How we acquire data
- What data NationGraph should have that nobody else does
What You'll Do- Own our external data platform end-to-end
- Design systems spanning discovery, acquisition, extraction, normalization, entity resolution, validation, storage, serving, and monitoring.
- Establish the architecture and abstractions other engineers build on.
- Map the world of government data
- Develop a deep understanding of where government information lives.
- Understand how it is published, how it changes, and how information across thousands of institutions can be connected.
- Build systems for messy, real-world data
- Work across government websites, APIs, procurement systems, PDFs, spreadsheets, meeting records, and public records.
- Build for changing schemas, broken sources, conflicting records, and edge cases.
- Use AI to rethink the traditional data stack
- Work with our ML Research team to use LLMs, agents, and emerging models to:
- Discover new sources
- Understand unfamiliar schemas
- Extract structured information
- Resolve entities
- Monitor data quality
- Detect when sources change
- Build proprietary data flywheels
- Create systems where more data improves our models.
- Use better models to discover and understand more data.
- Continuously expand NationGraph's underlying knowledge graph.
- Set technical direction
- Define the architecture for how NationGraph acquires and represents public-sector information.
- Make decisions that will shape the platform over the next several years.
- Help determine which technical investments create the strongest long-term data advantage.
You Might Be a Good Fit If- You're an unusually strong engineer who genuinely enjoys working with data.
- You've owned significant production data systems end-to-end.
- You enjoy the detective work of making sense of unfamiliar, messy datasets.
- You're strong in Python, Go, or another systems/backend language.
- You're highly proficient with SQL.
- You understand distributed data systems, including:
- Orchestration
- Idempotency
- Backfills
- Retries
- Observability
- Lineage
- Failure recovery
- You have experience with one or more of:
- Large-scale external data
- Crawling
- Information retrieval
- Entity resolution
- Knowledge graphs
- Document processing
- You're excited about using LLMs and modern ML as components of data infrastructure.
- You care deeply about data quality, correctness, and reliability.
- You have strong product judgment and can reason about what data is actually worth acquiring, not just how to acquire it.
- You thrive in ambiguity and would rather create the architecture than be handed one.
We're particularly interested in backgrounds spanning:
- Alternative data
- Quantitative research infrastructure
- Search and crawling
- AI data infrastructure
- Knowledge graphs
- Large-scale document processing
- Data aggregation
None of these are requirements.
Our Engineering Stack- Backend: Python, Go, PostgreSQL
- Infrastructure: Redis, Docker, Kubernetes
- Frontend: React, TypeScript
- AI / ML: LLMs, agents, proprietary models, and emerging frontier-model research
Our stack will evolve. At Staff level, you'll help decide how.