Gatik AI

Senior/Staff Site Reliability Engineer

Gatik AI$180K — $260K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Site Reliability, DevOps, or Infrastructure roles
  • Strong networking knowledge, including protocols and troubleshooting
  • Hands-on with Docker and its ecosystem tools
  • Expert in Kubernetes and Helm package management
  • Proficient in relational and time-series databases like Postgres
  • Familiar with workflow orchestration tools such as Argo and Airflow
  • Scripting experience in Python and Bash for automation
  • Experience with dashboard tools like Grafana

Responsibilities

  • Upgrade and maintain physical and cloud infrastructures for data offloading
  • Partner with infrastructure teams for monitoring and troubleshooting systems
  • Design and develop BI dashboards and ETL pipelines for insights
  • Architect and deploy test environments for infrastructure solutions
  • Automate deployment and scaling of remote monitoring software
  • Conduct analysis on infrastructure performance to identify optimization opportunities

Benefits

  • Onsite work environment five days a week in Santa Clara, CA
  • Opportunity to work on cutting-edge autonomous vehicle technology
  • Collaboration with infrastructure and platform teams
  • Potential for professional growth in a rapidly expanding field
  • Involvement in critical operations supporting fleet expansion
Full Job Description
About the role

We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring, and scaling of our growing fleet of autonomous vehicles. In this role, you will work closely with our infrastructure and platform teams to manage rollouts of both on-premises and cloud infrastructure in support of expansions to new customer sites.

You will be directly involved in the setup and monitoring of our data offload systems, remote supervision stations, and on-prem continuous integration (CI) environments, ensuring our infrastructure is highly reliable, secure, and optimized for performance. This position plays a critical role in keeping our autonomy operations running smoothly while supporting the rapid growth of our fleet and customer base.

This role is onsite 5 days a week at our Santa Clara, CA office!

What you'll do

  • Upgrade and maintain both physical and cloud infrastructure used for offloading data from our autonomous vehicle fleet.
  • Partner with the infrastructure and platform engineering teams to monitor, maintain, and troubleshoot our on-premises data offload and CI systems.
  • Design, develop, and maintain business intelligence (BI) dashboards and ETL (extract, transform, load) pipelines to provide actionable insights into our infrastructure performance and health.
  • Architect and deploy test environments to validate internal and customer-facing infrastructure solutions.
  • Automate deployment, scaling, and upgrading of our remote monitoring software to ensure operational efficiency.
  • Perform ongoing analysis of infrastructure performance, identifying opportunities for optimization in latency, throughput, and reliability.

What we're looking for

  • 5+ years of experience in a related role such as Site Reliability Engineer, DevOps Engineer, or Infrastructure Engineer.
  • Strong knowledge of networking fundamentals, including protocols, troubleshooting, and optimization.
  • Hands-on experience with Docker and related ecosystem tools (e.g., Docker Compose, Kaniko).
  • Expertise in Kubernetes deployments and package management via Helm.
  • Proficiency with relational and time-series databases (e.g., Postgres, TimescaleDB, InfluxDB).
  • Familiarity with workflow orchestration tools such as Argo and Airflow.
  • Proven experience managing upgrades and rollbacks for customer-facing SaaS environments.
  • Scripting experience in Python and Bash for automation and tooling.
  • Experience building and maintaining dashboards with tools like Grafana.

Salary Range - $180,000- $260,000

About Gatik AI

Gatik AI is a technology company that develops autonomous vehicles for business to business short-haul logistics. The company was founded in 2017 by Gautam Narang and Arjun Narang. Gatik AI's vehicles are designed to operate on fixed routes between distribution centers, warehouses, and retail locations. The company's mission is to deliver goods safely, efficiently, and on time, while reducing road congestion and carbon emissions. Gatik AI has partnerships with Walmart and Loblaw Companies Limited, two of the largest retailers in North America. The company is headquartered in Vancouver, Canada, with offices in Palo Alto, California.
Learn more about Gatik AI
Size
50 employees
Industry
Founded
2017

Similar Jobs

More Jobs at Gatik AI

More Information Technology Jobs

Find similar Senior/Staff Site Reliability Engineer jobs: