Senior Software Engineer, Cloud Reliability

Zilliz • $175K — $225K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience in production cloud systems or online services
  • Bachelor's in Computer Science or equivalent experience
  • Proficient in Kubernetes, Docker, and major cloud platforms (AWS, GCP, Azure)
  • Understanding of distributed systems and operational tradeoffs
  • Familiarity with distributed databases and large-scale systems is beneficial
  • Experience with multi-tenant systems or large infrastructure deployments is valuable
  • Knowledge of modern cloud operations tools like Terraform and Prometheus

Responsibilities

  • Manage the reliability and stability of Zilliz Cloud during growth phases
  • Troubleshoot complex issues related to Kubernetes and cloud infrastructure
  • Create automation tools for incident management and diagnostics
  • Convert recurring problems into reusable tools and documentation
  • Enhance observability metrics like latency and availability
  • Collaborate with engineers to improve cloud system reliability and automation

Benefits

  • High autonomy with a strong ownership culture
  • Fast-paced work environment prioritizing efficiency
  • Remote collaboration with distributed teams across APAC
  • Flexible on-call setup based on timezone rather than disruptive overnight paging
Full Job Description


What you will do:

  • Own the reliability, availability, and production stability of Zilliz Cloud as we scale through the next stage of growth
  • Debug complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems
  • Build automation and diagnostic tooling; log analysis, alert correlation, incident investigation, runbook automation, and remediation workflows so problems get solved once, not repeatedly
  • Turn recurring incidents into reusable tools, automation, documentation, and product improvements
  • Improve observability across latency, availability, throughput, and resource efficiency
  • Partner with database and infrastructure engineers to make Zilliz Cloud more reliable, scalable, and automated


What we are looking for:

  • 3+ years building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services
  • Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience
  • Strong hands-on experience with Kubernetes, Docker, and at least one major cloud platform (AWS, GCP, or Azure)
  • Solid understanding of distributed systems; availability, scalability, performance, failure recovery, and operational tradeoffs
  • Experience with distributed databases, storage systems, search systems, or large-scale online systems is a strong plus
  • Experience operating highly multi-tenant systems or large infrastructure fleets; thousands of nodes, clusters, tenants, or customer deployments is especially valuable
  • Familiarity with modern cloud operations tooling such as Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems
  • Strong bias for action, and the drive to thrive in a fast-paced, rapidly scaling environment


How we operate:

  • High ownership: You own production reliability end-to-end. The whole system, not a slice of it. High autonomy, high trust, minimal process.
  • Fast and focused: We ship often and keep a high bar. This team suits engineers who want velocity and a steep growth curve over red tape.
  • Globally distributed: We work closely with our core engineering teams across APAC. Occasional early morning or evening syncs in exchange for an on-call setup designed around timezone coverage, not overnight pages.


$175,000 - $225,000 a year

About Zilliz

Zilliz is an artificial intelligence company that provides a platform for data analysis and machine learning. The company's platform is based on a proprietary technology called Vector Similarity Search Engine (vSearch) that enables fast and accurate search and analysis of large datasets. Zilliz's technology has the potential to revolutionize the field of data analysis by enabling companies to process and analyze large amounts of data in real-time. The company was founded in 2018 and is headquartered in Beijing, China.
Learn more about Zilliz
Size
200 employees
Industry
Founded
2018

Similar Jobs

More Jobs at Zilliz

More Information Technology Jobs

Find similar Senior Software Engineer, Cloud Reliability jobs: