BS or MS in Computer Science or related field, or equivalent experience.
Experience managing scalable, fault-tolerant Linux/Kubernetes/JVM infrastructure in public clouds.
Proficient in Linux OS, Networking, and Database concepts.
Skilled in deploying and troubleshooting Kubernetes clusters.
Familiarity with NoSQL databases like Cassandra.
Expertise in cloud platforms such as AWS, Azure, and GCP.
Knowledge of configuration management tools like Puppet.
Proficient in Bash or Python for automation.
Experience with Infrastructure as Code (IaC) tools like Ansible or Terraform.
Strong problem-solving and communication skills.
Responsibilities
Maximize system uptime and availability to meet SLAs.
Establish comprehensive monitoring and alerting for critical systems.
Resolve complex issues for key services and automate solutions.
Influence design and architecture standards for platform support.
Lead automation efforts for system updates and upgrades.
Set up infrastructure and tools to enhance deployment processes.
Collaborate with Services and Engineering teams across functions.
Benefits
Excellent benefits package.
Generous equity plan.
Full Job Description
We are looking for Associate Site Reliability Engineer/Site Reliability Engineer to join our team at our HQ in Redwood City, CA.
Responsibilities:
Maximize system uptime and availability, ensuring functional and performance SLAs.
Establish end-to-end monitoring and alerting on all critical aspects.
Solve complex problems for critical services and build automation to prevent problem recurrence.
Influence and create new designs, architectures, standards, and methods for supporting the platform.
Initiate and lead scripting and automation to streamline system updates and upgrades.
Set up critical infrastructure, tools, and framework to streamline the deployment cycle.
Work cross-functionally with Services and Engineering teams.
Qualifications:
BS or MS in Computer Science, related field, or equivalent professional experience.
Demonstrated experience in deploying, managing, and operating scalable and fault-tolerant Linux/Kubernetes/JVM-based infrastructure in AWS, GCP, and other public clouds.
Expertise in Linux Operating Systems, Networking, and Database concepts.
Experience deploying, upgrading, and troubleshooting Kubernetes clusters and workloads.
Experience with Cassandra (or another NoSQL alternative).
Expertise in cloud providers, such as Amazon Web Services, Azure, and GCP.
Experience with configuration management systems such as Puppet.
Experience in Bash or Python; to automate and monitor systems.
Experience with IaC tools like Ansible or Terraform.
Excellent problem-solving, critical thinking, and communication skills.
Experience supporting as a DevOps or sys admin for commercial SaaS solutions.
C3 AI provides excellent benefits, a competitive compensation package and generous equity plan.
California Base Pay Range
$90,000-$160,000 USD
About C3.ai
C3.ai is a software company that provides a platform for developing and deploying enterprise-scale AI applications. The company was founded in 2009 by Tom Siebel and is headquartered in Redwood City, California. C3.ai's platform is used by a range of industries, including oil and gas, manufacturing, and healthcare. The company has raised over $500 million in funding and was valued at $3.3 billion as of 2020.