Senior Data QA-Onshore

V4C.ai

$120K — $145K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Data Engineering, or a quantitative field.
  • 8+ years of experience in data engineering, data QA, or SDET roles; 2+ years specifically in test automation in Databricks.
  • Expertise in PySpark and Python for large dataset processing.
  • Proficient in advanced Spark SQL, including optimization techniques.
  • Hands-on experience with big-data validation libraries like Great Expectations and pytest.
  • Knowledge of cloud infrastructure deployment on platforms like AWS, Azure, or GCP.

Responsibilities

  • Design and implement automated testing frameworks within Databricks using PySpark and Python.
  • Create automated assertions for Delta Lake tables regarding data drift and validation.
  • Develop complex automated tests for both batch and real-time data pipelines.
  • Ensure governance by validating data lineage, audit logs, and access controls.
  • Lead the integration of automated data quality tests into CI/CD pipelines.
  • Mentor junior engineers and promote data quality principles across teams.
  • Conduct performance testing on Spark jobs and complex query optimizations.

Benefits

  • Opportunities for technical leadership and mentorship.
  • Exposure to cutting-edge big data technologies and architectures.
  • Strong focus on personal development and growth in data quality practices.
  • Collaborative work environment with a commitment to excellence in data engineering.
Full Job Description
Job Summary

We are seeking a Senior Data QA Automation Engineer to lead the quality strategy, design, and implementation of automated testing frameworks for our big data platforms. In this senior role, you will own the end-to-end data validation strategy within our Databricks Lakehouse architecture, ensuring high-quality, reliable, and compliant data across Delta Lakes, ETL pipelines, and enterprise data models. You will work closely with Data Engineering leadership to establish rigorous quality gates and mentor mid-to-junior engineers on data testing best practices.

Key Responsibilities
  • Strategic Framework Design: Architect, build, and scale automated test frameworks from scratch natively within Databricks using PySpark, Python, and SQL.
  • Lakehouse Quality Engineering: Design robust automated assertions for Delta Lake tables, including checking data drift, schema evolution, and historical data validation via time-travel functions.
  • Enterprise Pipeline Testing: Code complex automated scenarios to validate large-scale batch and real-time streaming data pipelines (Structured Streaming), ensuring source-to-target integrity.
  • Governance Validation: Programmatically verify data lineage, audit logs, and access controls implemented via Databricks Unity Catalog.
  • CI/CD & DevOps Ownership: Lead the integration of automated data quality tests into enterprise CI/CD pipelines (e.g., Azure DevOps, GitHub Actions), leveraging Databricks Workflows, APIs, or Airflow.
  • Technical Leadership & Mentorship: Act as the subject matter expert for data quality; mentor junior team members, establish QA standards, and advocate for data quality principles across engineering teams.
  • Performance Assessment: Design and execute automated performance and scalability tests on Spark jobs, large clusters, and complex query optimizations.


Required Skills and Qualifications
  • Education: Bachelor's or Master's degree in Computer Science, Data Engineering, or a related quantitative field.
  • Experience: 8+ years of experience in data engineering, data QA, or software development engineering in test (SDET), with at least 2+ years of dedicated experience architecting test automation in Databricks.
  • Expert PySpark & Python: Mastery of Python and PySpark (DataFrames and SQL APIs) for processing and profiling large datasets.
  • Advanced Spark SQL: Deep expertise in writing advanced SQL queries, optimization techniques, and understanding Spark query execution plans.
  • Advanced Testing Tooling: Hands-on mastery of big-data validation libraries (e.g., Great Expectations, pytest, Delta Live Tables expectations).
  • Cloud Infrastructure: Strong operational knowledge of Databricks deployment on a major cloud provider (AWS, Azure, or GCP).


Preferred Qualifications
  • Certifications: Databricks Certified Data Engineer Professional or Databricks Certified Machine Learning Professional.
  • Streaming Expertise: Experience validating real-time event-streaming architectures (Kafka, Event Hubs, Kinesis).
  • Data Ops: Solid understanding of DataOps culture, testing infrastructure as code, and data observability principles.

Similar Jobs

More Jobs at V4C.ai

More Information Technology Jobs

Find similar Senior Data QA-Onshore jobs: