Data Engineer

Vitol

$110K — $130K *
Energy & Utilities
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of hands-on data engineering experience with production data pipelines
  • Strong Python skills with a focus on clean, maintainable code and proficiency in data-handling libraries like pandas
  • Depth in SQL for modeling, querying, and performance tuning
  • Experience with Snowflake or comparable cloud data warehouse technologies, particularly in dimensional modeling and cost-aware design
  • Broad familiarity with data technologies including streaming, caching, timeseries, and relational systems, with specific tools like Kafka and Redis
  • Experience in pipeline orchestration, preferably with Prefect, Airflow, or similar tools
  • Solid engineering fundamentals in design patterns, testing, and version control

Responsibilities

  • Build, test, and operate production data pipelines using Python and Prefect
  • Create and maintain warehouse models in Snowflake with a focus on documentation and cost efficiency
  • Migrate legacy pipelines and components to modern frameworks, addressing technical debt
  • Work across various platform technologies, adhering to established patterns
  • Acquire and reliably ingest data from external sources like APIs and feeds
  • Implement data quality and reconciliation checks proactively
  • Support and monitor deployed components, responding to incidents and maintaining reliability

Benefits

  • Medical insurance
  • Dental coverage
  • Vision insurance
  • Paid vacation time
  • 401k plan with company contributions
  • Life insurance
  • Short-term and long-term disability coverage
Full Job Description
Job Description

The role

We are looking for a Data Engineer to build and run the pipelines and data models behind our core data platform - the centralized data layer that serves global energy traders and analysts in real time.

This is a hands-on delivery role. You will write the Python that moves market and fundamental data from source through to the analysts, applications, and AI systems that consume it, and you will support what you ship.

Python is the core skill, but it is not the whole job. We are standing up a Snowflake warehouse, migrating legacy pipelines onto a Prefect-based framework, and consolidating how the platform stores and exposes data. The work reaches across more of the stack than pipeline code alone.

The platform is polyglot by design: streaming, caching, timeseries, relational, and warehouse technologies each doing what they are good at. You will build to the established patterns across that stack, and we expect you to understand why each technology sits where it does.

You will work alongside senior engineers who set those patterns, and directly with the analysts and desk users your work serves. The scope grows as you do - this is a role you can build a platform career from.

What You Will Do
  • Build, test, and operate production data pipelines in Python on our modern pipeline framework, orchestrated with Prefect
  • Build and maintain warehouse models in Snowflake using dbt - clean, tested, documented, and cost-aware
  • Migrate legacy pipelines and Oracle-based components onto current frameworks and standards, retiring technical debt as you go
  • Work across the platform stack - Kafka, Redis, InfluxDB, Oracle, and Snowflake - building to the established pattern for each rather than reaching for the tool you already know
  • Acquire data from external sources - vendor APIs, files, feeds, and web sources - and land it reliably
  • Implement data quality, freshness, and reconciliation checks so problems surface before users find them
  • Use AWS data services where our platform patterns call for them
  • Support what you ship - monitoring and alerting on your components, and investigating when something breaks
  • Contribute to making platform data AI-ready, and work with the catalog and steward teams so what you build is documented, classified, and findable
  • Engage directly with analysts and desk users to check that what you are building solves the actual problem


Qualifications
  • 3+ years of hands-on data engineering experience building and operating production data pipelines
  • Strong Python - clean, tested, maintainable code, not scripts that happen to run. Fluency with pandas and the wider data-handling ecosystem
  • SQL depth - you can model, query, and tune, and you know what makes a query expensive before you run it
  • Snowflake or a comparable cloud warehouse - dimensional modeling, performance tuning, and cost-aware design. dbt experience is a strong plus
  • Breadth across data technologies - streaming, caching, timeseries, relational, and warehouse, with a view on where each belongs. Kafka, Redis, InfluxDB, Oracle, and Snowflake are what we run; comparable exposure matters more than an exact match
  • Pipeline orchestration experience - Prefect preferred; Airflow, Dagster, or similar considered
  • Sound engineering fundamentals - object-oriented design, design patterns, testing, code review, and version control as habits rather than requirements
  • Working knowledge of AWS data services and how to compose them into something reliable
  • A build-to-operate mindset - monitoring, alerting, and failure recovery are part of how you design, not something added later
  • Clear communication - you can explain a technical trade-off to an analyst, and turn a vague request into the right set of questions
  • Terraform, Docker, or API development (FastAPI, Flask) exposure is welcome - we use all three
  • Experience in financial services, commodities, or energy trading data is a plus; the data volume, latency requirements, and stakes are real


Additional Information

What Success Looks Like in Year One
  • You own several production data components outright and they run reliably with minimal intervention
  • Warehouse models you built in Snowflake are in active use by analysts and downstream applications
  • A meaningful share of the legacy pipelines you inherited have been migrated onto the current framework and standards
  • Data quality and freshness checks you implemented are catching issues before users report them
  • You build to the right platform pattern without being told which one applies - and you say so when none of them fit
  • Analysts on your projects come to you directly, and trust what you deliver

Comprehensive benefit coverage includes:
  • Medical
  • Dental
  • Vision
  • Paid Vacation
  • 401k with company contributions
  • Life insurance
  • Short-term & Long-term disability

Similar Jobs

More Jobs at Vitol

  • Data Engineer
    $110K — $130K *
    Houston, TX 77084 (Harris County)
    Energy & Utilities
    In-Person
  • Forward Deployed Engineer
    $110K — $130K *
    Houston, TX 77084 (Harris County)
    Enterprise Technology
    In-Person
  • Data Support Engineer
    $80K — $95K *
    Houston, TX 77084 (Harris County)
    Energy & Utilities
    In-Person

More Energy & Utilities Jobs

Find similar Data Engineer jobs: