Job Description:The Data Engineer will be based in Fremont, California. In this position, you will report to the Senior Manager of Systems Architecture. We are looking for a talented, early-career data engineer to help us build and maintain the data pipelines that ingest, process, and optimize the data from our existing installations. We have hundreds of thousands of devices in the field in virtually every continent, controlling our solar trackers. Our data and analytics infrastructure supports our maintenance operations and informs new product design. When you work with us, you will be making a positive, material impact on the solar capacity of the planet and learning how to scale outdoor IOT deployments in some very harsh environments
Here is a glimpse of what you'll do...
- Build and maintain reliable, high-quality data pipelines and architecture
- Develop data pipelines, tools, and applications by producing clean, efficient code in Go, Python and SQL
- Automate data ingestion and transformation tasks through appropriate tools and scripting
- Review and debug pipeline code
- Perform data validation and quality testing
- Collaborate with internal teams to improve our data and analytics products
- Document data workflows and monitor pipeline health
- Ensure our data stack is up to date with the latest technologies
Here is some of what you'll need (required)...
- Foundational programming skills, ideally in Go or C.
- Understanding of SQL, databases, and basic data modeling.
- Familiarity with writing and debugging queries and scripts.
- Bachelor's degree in Computer Science, Data Science, or related field
- Familiarity with Linux OS
- Analytical mind with problem-solving aptitude
- Ability to work independently and as part of a team
Here are a few of our preferred experiences...
- Knowledge of common data and programming languages like Python, SQL, and Scala; exposure to Java is a plus
- Experience processing IOT big data, ideally from millions of devices in the field.
- Data technologies like SQL and time-series databases (e.g., Timescale), Databricks, Apache Kafka, Apache Spark, or Apache Hadoop. Docker or Kubernetes
- Familiarity with data orchestration tools like Apache Airflow is a plus
- Working experience with cloud infrastructure as a service, like Microsoft Azure or AWS
- UI development languages, like Angular, or React are a plus
- Exposure to Databricks and Timescale is a strong plus