Software Engineer, Data Mining

NationGraph Business Corporation

• $100K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in data engineering or related field
  • Strong proficiency in Python or Go and SQL
  • Experience with distributed data systems and architecture
  • Familiarity with AI applications in data processing
  • Proven track record of owning significant production data systems
  • Ability to navigate and process messy datasets effectively
  • Strong product judgment regarding data value acquisition

Responsibilities

  • Own and design the end-to-end external data platform
  • Map and comprehensively understand the landscape of government data
  • Build robust systems that handle real-world, messy data
  • Leverage AI to innovate and enhance the traditional data stack
  • Create proprietary data flywheels that enhance knowledge graph expansion
  • Set the technical direction for acquiring and representing public-sector information
  • Influence long-term architectural decisions for data infrastructure

Benefits

  • Opportunity to work on impactful, large-scale data challenges
  • Collaborative environment working with cross-functional teams
  • Exposure to cutting-edge AI and ML technologies
  • Potential to shape the future of public-sector data acquisition and usage
  • Culture that values innovation and technical excellence
Full Job Description
Software Engineer, Data Mining
The Role

We're looking for a Software Engineer, Data Mining to own one of the most important technical problems at NationGraph: building the systems that acquire public-sector information from across the internet at massive scale.

Our goal is to operate hundreds of thousands, and eventually millions, of scrapers covering every level of government across the U.S. and Canada, and eventually worldwide.

This is not a role focused on manually building individual scrapers. You'll own the infrastructure, abstractions, and automation that allow us to create, deploy, monitor, and maintain an enormous fleet of scrapers reliably.

You'll work across:
  • Web crawling and scraping
  • Browser automation
  • Distributed systems
  • Data extraction
  • Infrastructure and orchestration
  • LLMs and agents
  • Monitoring and observability
What You'll Do
  • Own our scraping infrastructure end-to-end
    • Build systems for creating, deploying, scheduling, monitoring, and maintaining hundreds of thousands of scrapers.
    • Design abstractions that allow us to scale toward millions of sources without scaling engineering effort linearly.
  • Build for the messy internet
    • Work across government websites, APIs, procurement systems, PDFs, spreadsheets, meeting records, and legacy systems.
    • Handle changing websites, undocumented APIs, rate limits, broken sources, and countless edge cases.
  • Make scraping a distributed systems problem
    • Build for orchestration, concurrency, retries, backfills, change detection, observability, cost management, and failure recovery.
    • Ensure we know when sources break, data disappears, or extraction silently becomes incorrect.
  • Use AI to rethink scraping
    • Work with our ML Research team to use LLMs and agents to:
      • Discover new sources
      • Understand unfamiliar websites
      • Generate scraping logic
      • Detect source changes
      • Diagnose and repair failures
      • Validate extracted data
  • Build systems that get better with scale
    • Identify common platforms and patterns that can unlock thousands of government agencies at once.
    • Make new sources increasingly cheap and automated to onboard.
  • Expand our coverage globally
    • Help comprehensively map public-sector information across the U.S. and Canada.
    • Build the foundation to eventually acquire public-sector information worldwide.
You Might Be a Good Fit If
  • You're an unusually strong engineer who enjoys figuring out how things work.
  • You've built production web crawlers, scraping systems, browser automation, or large-scale external data pipelines.
  • You're strong in Python, Go, TypeScript, or another backend/systems language.
  • You understand the realities of scraping modern websites, including:
    • JavaScript rendering
    • Sessions and cookies
    • Rate limits
    • Proxies
    • Authentication
    • Changing schemas and websites
  • You understand distributed systems, including:
    • Orchestration
    • Queues and concurrency
    • Idempotency
    • Retries
    • Backfills
    • Observability
    • Failure recovery
  • You care deeply about data quality, correctness, and reliability.
  • You're excited about using LLMs and agents to automate traditionally manual scraping work.
  • You naturally think about leverage: not how to scrape one website, but how to build a system capable of scraping the next 10,000.
  • You thrive in ambiguity and would rather build the system than be handed one.

We're particularly interested in backgrounds spanning:
  • Alternative data
  • Quantitative research infrastructure
  • Search and crawling
  • AI data infrastructure
  • Knowledge graphs
  • Large-scale document processing
  • Data aggregation

None of these are requirements.

Our Engineering Stack
  • Backend: Python, Go, PostgreSQL
  • Infrastructure: Redis, Docker, Kubernetes
  • Frontend: React, TypeScript
  • AI / ML: LLMs, agents

Our stack will evolve, you'll help decide how.

If the idea of building the data infrastructure to map and understand how government works sounds exciting, we'd love to talk.

Similar Jobs

More Jobs at NationGraph Business Corporation

More Information Technology Jobs

Find similar Software Engineer, Data Mining jobs: