Full Job Description
Designation : Manager - Data Engineering
Level : L4
Location : San Jose, California , United States
Experience : 10 to 15 Years
Job Role :
We're looking for a Senior Data Engineer with deep expertise in building and scaling modern data platforms. You'll design, build, and maintain robust data pipelines and infrastructure that power analytics and business-critical applications, with a strong emphasis on the Databricks and AWS ecosystem.
Key Responsibilities
Design, develop, and maintain scalable ETL/ELT pipelines using Databricks (PySpark/Spark SQL)
Build, orchestrate, and monitor workflows using Apache Airflow (DAG design, scheduling, dependency management)
Architect and manage data infrastructure on AWS (S3, Glue, Lambda, EMR, Redshift, IAM, etc.)
Optimize data pipelines for performance, reliability, and cost-efficiency
Implement data quality checks, testing frameworks, and observability/monitoring for pipelines
Collaborate with Data Scientists, Analysts, and Product teams to understand data requirements
Design and maintain data models (Lakehouse/Medallion architecture - Bronze/Silver/Gold layers)
Implement CI/CD pipelines for data engineering workflows
Ensure data security, access controls, and compliance across the data platform
Document architecture, pipelines, and processes for team knowledge sharing
Mentor junior data engineers and contribute to engineering best practices
Required Qualifications
7+ years of experience in Data Engineering roles
Extensive hands-on experience with Databricks (Delta Lake, Unity Catalog, cluster management, notebooks, job orchestration)
Strong experience with AWS cloud services (S3, IAM, Glue, EMR, Lambda, Redshift, CloudWatch)
Proven expertise in Apache Airflow for workflow orchestration and scheduling
Strong programming skills in Python and SQL
Solid understanding of Apache Spark (PySpark, Spark SQL, performance tuning)
Experience with data modeling, warehousing concepts, and Lakehouse architecture
Familiarity with version control (Git) and CI/CD practices
Strong understanding of data governance, quality, and lineage principles
Excellent problem-solving and communication skills