ServiceNow

Staff Engineer, Machine Learning Systems & Reliability - Moveworks

ServiceNow$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of experience in software or platform engineering, SRE, production engineering, or ML infrastructure.
  • Proficiency in Python and another production systems language like Go, Java, C++, or Rust.
  • Experience with distributed production systems, including troubleshooting and capacity planning.
  • Hands-on knowledge of cloud infrastructure, containers, Kubernetes, and CI/CD practices.
  • Understanding of the ML lifecycle, capable of collaborating with ML engineers or researchers.
  • Ability to identify whether issues are service-health or data/model-quality related.
  • Familiarity with SRE practices including SLIs, SLOs, and blameless postmortems.
  • Strong automation mindset, creating reliable platforms for internal users.

Responsibilities

  • Design the production path for the entire ML lifecycle, from data preparation to retraining.
  • Build workflows for continuous delivery and automate checks for quality and performance.
  • Implement rollout strategies like canary releases and operational rollbacks.
  • Create systems for controlled feedback loops and model refreshes.
  • Define SLIs and SLOs for infrastructure and model performance.
  • Link product telemetry with operational data for comprehensive problem analysis.
  • Optimize scalability and cost efficiency of data workloads and training processes.
  • Participate in the entire production lifecycle, from architecture reviews to incident response.
  • Develop self-service automation tools to minimize operational tasks for ML engineers.
  • Integrate LLMs into operational workflows for measurable enhancements.
  • Establish cloud infrastructure and observability standards.

Benefits

  • Flexible work arrangements, including remote options.
  • Opportunities for technical leadership and mentorship.
  • A culture that fosters collaboration across ML, data, product, and platform teams.
  • Access to resources for professional growth and skill enhancement.
Full Job Description
Job Description

We are building AI-enabled product capabilities that improve through data, feedback, and real-world use. We need the production systems that make those capabilities dependable: repeatable delivery, measurable quality, controlled learning loops, and reliable operation at scale.

We're looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems.

This role sits at the intersection of ML systems, platform engineering, and site reliability engineering. You will partner with ML, data, product, and infrastructure teams to create a paved path from experimentation to production-and take ownership of how those systems perform and evolve once deployed.

What you'll do
  • Design and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining.
  • Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety, performance, and compatibility checks.
  • Implement safe rollout patterns such as shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches.
  • Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails.
  • Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior.
  • Connect model analytics and product telemetry with traditional operational signals so teams can understand whether a problem originates in infrastructure, data, model behavior, or the surrounding product.
  • Improve the scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads. Own capacity planning and resource optimization, including GPU resources where applicable.
  • Participate in production ownership across the service lifecycle: architecture reviews, deployment, on-call, incident response, blameless postmortems, and systemic remediation.
  • Build self-service platforms and automation that reduce operational toil and shorten the time required for ML engineers and data scientists to reach production.
  • Apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows where they produce reliable, measurable improvements.
  • Establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance.
  • Provide technical leadership across ML, data, product, and platform teams, mentoring engineers and influencing architecture without relying on formal authority.


Qualifications

To be successful in this role you have:
  • A track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructure.
  • Strong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or Rust.
  • Experience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimization.
  • Hands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability.
  • Practical understanding of the ML lifecycle-including training, evaluation, model deployment, serving, monitoring, versioning, and retraining-and the ability to collaborate effectively with applied ML engineers or researchers.
  • Experience distinguishing service-health problems from data-quality or model-quality problems.
  • Familiarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortems.
  • A strong automation and internal-customer mindset: you build platforms that are reliable, understandable, and pleasant for other engineers to use.
  • Excellent technical judgment and communication skills, especially when navigating ambiguity and coordinating across teams during production incidents.


Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

About ServiceNow

ServiceNow provides cloud-based solutions that define, structure, manage, and automate services for enterprise operations in North America, Europe, the Middle East, Africa, the Asia Pacific, and other countries. The company offers service management solutions, including incident, problem, change, request, and cost management as well as service catalogs; and IT, HR, facilities, and field service management solutions. It also provides IT operations management solutions covering service mapping, delivery, and assurance solutions; business management solutions such as financial management, project portfolio suite, vendor performance management, and performance analytics as well as governance, risk, and compliance; and application development services.

ServiceNow Careers

Join the dynamic team at ServiceNow, a global leader in digital workflow solutions, where innovation and leadership converge to shape the future of work. At ServiceNow, we offer more than just job opportunities; we provide a platform for professional growth and a chance to be part of a culture that values diversity, creativity, and continuous learning.

Work You’ll Do

Embark on a career journey with ServiceNow and contribute to the world’s leading enterprises' digital transformation. Our team is at the forefront of developing cutting-edge technologies that improve how people work. With ServiceNow, you will use your skills to impact businesses and industries profoundly, driving efficiency and innovation.

Join Our Market-Leading Team

ServiceNow is not just another technology company. We are a team that thrives on diversity and leadership, fostering an inclusive environment that promotes growth and development. Our commitment to diversity training ensures that every team member can achieve their potential.

Innovative Work

ServiceNow is home to more than 10,000 dedicated professionals who lead the charge in digital workflows and enterprise solutions. As part of our team, you will engage in projects that merge technology with practical applications, creating revolutionary products that advance how services are delivered and managed.

Career Development

At ServiceNow, your career trajectory is filled with boundless opportunities. We support your growth with robust training programs, leadership development courses, and access to global challenges. Whether you are looking for an internship, full-time position, or leadership role, ServiceNow equips you with the tools to excel.

Be Part of a Great Team

Working at ServiceNow means being part of a community that values teamwork and innovation. Our collaborative environment encourages networking and sharing ideas, making our workplace vibrant and dynamic. The benefits of joining ServiceNow extend beyond comprehensive health and wellness; they include fostering professional connections and friendships that last a lifetime.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional or a recent graduate, ServiceNow offers a range of employment options to suit your career goals. From internships that provide real-world experience to full-time positions that challenge you to leverage your expertise, we are committed to hiring the best talent.

Stay Connected

Join Our Team Search open positions that match your skills and interests. At ServiceNow, we look for passionate, curious, and solution-driven team players. Explore the possibilities that await you at a company that is committed to your professional success.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at ServiceNow.

ServiceNow Careers

Empowering professionals to achieve more, ServiceNow is where careers are future-proofed, and ambitions are realized. Join us in our journey of growth and innovation.
Learn more about ServiceNow
Size
16,881 employees
Market Cap
$76.5 billion
Industry
Net Income
$118.5 million
Founded
2004
5 Year Trend
+33.5%
Revenue
$4.5 billion
NASDAQ

Similar Jobs

More Jobs at ServiceNow

More Information Technology Jobs

Find similar Staff Engineer, Machine Learning Systems & Reliability - Moveworks jobs: