Job Summary
We are seeking a Senior Data Engineer with strong expertise in Spark and streaming technologies to build real-time, scalable data pipelines using Apache Spark, Kafka, and cloud services. The role will focus on ingesting, transforming, and delivering data for analytics and machine learning while supporting scalable data architectures and high-performance data processing.
Key Responsibilities
• Design, develop, and maintain ETL/ELT data pipelines for batch and real-time data ingestion, transformation, and loading using Spark (PySpark/Scala) and streaming technologies such as Kafka and Flink.
• Build and optimize scalable data architectures, including data lakes, data warehouses such as BigQuery, and streaming platforms.
• Optimize Spark jobs, SQL queries, and data processing workflows for speed, efficiency, and cost-effectiveness.
• Implement data quality checks, monitoring, and alerting systems to ensure data accuracy and consistency.
Required Qualifications
• Minimum 8 years of total IT experience.
• Minimum 4 years of recent GCP experience.
• Strong proficiency in Python and SQL, with experience in Scala and/or Java.
• Strong expertise in Apache Spark, including Spark SQL, DataFrames, and Spark Streaming.
• Experience with streaming and messaging technologies such as Apache Kafka and/or Google Pub/Sub.
• Experience with cloud data services, particularly GCP, with familiarity with Azure data services.
• Knowledge of data warehousing technologies such as Snowflake and Redshift.
• Knowledge of NoSQL databases.
• Experience designing and developing scalable batch and real-time data pipelines.
• Strong understanding of data quality, monitoring, and data processing best practices.
Preferred Qualifications
• Experience with Apache Airflow.
• Experience with Databricks.
• Experience with Docker and Kubernetes.
• Experience with Apache Flink.
• Experience optimizing data pipelines and Spark workloads for performance and cost efficiency.
Location
• Sunnyvale, CA
Notes:
Only W2 Candidates required, Glider/prescreening is needed for this role
Mandatory Areas
Must Have Skills - Data Engineer
Skill 1 - big data, spark, pyspark, apache airflow, etl process
Skill 2 - 4+ years in GCP
Good To have Skills -
Skill 1 - Yrs of Exp -N/A
Skill 2 - Yrs of Exp -N/A
Skill 3 - Yrs of Exp -N/A
Skill 4 - Yrs of Exp -N/A
Mandatory if Applicable
Domain Experience (If any ) - N/A
Must have Certifications -N/A
Location - Sunnyvale CA
Onsite Requirement - Y/N- Y
Number of days onsite - 5 Days
If Onsite - Office Address
Only W2 Candidates required, pre-screening is needed for this role
Mandatory Areas
Must Have Skills - Data Engineer
Skill 1 - big data, spark, pyspark, apache airflow, etl process
Skill 2 - 4+ years in GCP
Good To have Skills -
Skill 1 - Yrs of Exp -N/A
Skill 2 - Yrs of Exp -N/A
Skill 3 - Yrs of Exp -N/A
Skill 4 - Yrs of Exp -N/A
Mandatory if Applicable
Domain Experience (If any ) - N/A
Must have Certifications -N/A
Location - Sunnyvale CA
Onsite Requirement - Y/N- Y
Number of days onsite - 5 Days
If Onsite - Office Address