Data Engineer

The Maven Group, LLC

$100K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2+ years of experience as a Data Engineer or similar role
  • Proficiency in Python or other relevant programming languages
  • Experience with data pipeline orchestration tools (e.g., Apache Airflow, dbt, Prefect, Dagster)
  • Strong understanding of cloud platforms, primarily AWS, with GCP or Azure experience considered
  • Expertise in SQL and familiarity with databases like PostgreSQL, Snowflake, or BigQuery
  • Knowledge of big data processing frameworks (e.g., Apache Spark, Databricks, Apache Kafka)
  • Familiarity with MLOps tools and workflows for machine learning deployment

Responsibilities

  • Write clean, efficient, and scalable code for data solutions using Python
  • Design and build reliable data workflows with tools like Apache Airflow or dbt
  • Work in cloud environments, primarily using AWS
  • Extract and integrate quality data using SQL and various database technologies
  • Process large-scale data workflows with frameworks like Apache Spark
  • Support deployment of machine learning models using Amazon SageMaker or Kubeflow
  • Monitor and troubleshoot data pipeline health and consistency

Benefits

  • Collaborative work environment with cross-functional teams
  • Opportunities to work with cutting-edge technologies
  • Career development and professional growth
  • Emphasis on documentation and best practices
  • Contribution to impactful data-driven projects
  • Possibility of tackling complex data challenges
  • Involvement in machine learning integration and MLOps
Full Job Description
Data Engineer

We are looking for a skilled and passionate Data Engineer to join our team. You will play a critical role in designing, building, and maintaining our data infrastructure to ensure seamless data flow, scalability, and reliability. You will work closely with data scientists, analysts, and other stakeholders to develop efficient data pipelines, manage large datasets, and integrate machine learning models into production environments.

Key Responsibilities:
  • Programming Fundamentals: Write clean, efficient, and scalable code to build and optimize data solutions using programming languages like Python.
  • Data Pipeline Development: Design, build, and orchestrate robust and reliable data workflows using tools such as Apache Airflow, dbt, Prefect, or Dagster.
  • Cloud Platform Familiarity: Work comfortably in cloud environments, with a strong preference for experience in AWS. Experience in GCP or Azure is also highly valued.
  • Database & Querying Skills: Extract, integrate, and ensure the quality of data from various sources using tools and technologies such as SQL, PostgreSQL, Snowflake, Amazon Redshift, or BigQuery.
  • Big Data Processing: Leverage frameworks like Apache Spark, Databricks, or Apache Kafka to process and manage large-scale data workflows with reliability and efficiency.
  • ML Integration / MLOps: Support the implementation, deployment, and scaling of machine learning models in production environments using tools like Amazon SageMaker, MLflow, or Kubeflow.
  • Monitoring & Troubleshooting: Monitor data pipeline health, troubleshoot issues, and ensure data consistency using tools such as Amazon CloudWatch, Datadog, or Great Expectations.
  • Collaboration & Documentation: Work closely with data scientists, analysts, and other stakeholders to understand data requirements, communicate solutions, and document processes using tools like Git, Jira, and Confluence.

Qualifications:
  • 2 years of experience as a Data Engineer or similar role.
  • Strong proficiency in Python or other programming languages relevant to data engineering.
  • Any additional experience with any of the following:
    • Hands-on experience with data pipeline orchestration tools (e.g., Apache Airflow, dbt, Prefect, Dagster).
    • Solid understanding of cloud platforms (AWS strongly preferred; GCP or Azure experience also considered).
    • Expertise in SQL and familiarity with relational and columnar databases (e.g., PostgreSQL, Snowflake, BigQuery).
    • Knowledge of big data processing frameworks (e.g., Apache Spark, Databricks, or Apache Kafka).
    • Familiarity with machine learning workflows and experience implementing MLOps tools (e.g., Amazon SageMaker, MLflow, or Kubeflow) in production environments.
    • Strong troubleshooting skills and experience monitoring data pipelines and system health using tools like Amazon CloudWatch, Datadog, or Great Expectations.
    • Excellent communication skills and a collaborative mindset, with a focus on documentation and best practices.

Preferred Skills:
  • Experience working with large-scale distributed systems.
  • Knowledge of data governance and security best practices.
  • Proven ability to work in cross-functional teams and contribute to problem-solving and innovation.


Clearance:
  • An active TS/SCI federal security clearance is required

Similar Jobs

More Jobs at The Maven Group, LLC

More Information Technology Jobs

Find similar Data Engineer jobs: