Newfold Digital

Senior DevOps Engineer, AI Platform

Newfold Digital$130K — $155K *
US-AnywhereRemote in Canada
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of experience in DevOps, SRE, or related roles.
  • Extensive hands-on experience with Kubernetes, particularly AKS and networking.
  • Strong knowledge of Microsoft Azure and Oracle Cloud Infrastructure (OCI).
  • Proficient in CI/CD tools like Jenkins, Bitbucket, and container technologies.
  • Solid understanding of cloud networking concepts such as DNS and load balancing.
  • Experience with observability tools like Grafana and Prometheus.
  • Hands-on with databases and caching systems like PostgreSQL and Redis.

Responsibilities

  • Translate technical designs into production-ready cloud infrastructure.
  • Design and manage Kubernetes environments on Azure and OCI.
  • Support varied AI workloads and asynchronous processing pipelines.
  • Manage networking and security components for cloud services.
  • Create and maintain CI/CD pipelines integrating various tools.
  • Automate infrastructure provisioning with Terraform and related technologies.
  • Ensure complete observability and troubleshooting readiness of infrastructure.

Benefits

  • Work in a cutting-edge AI environment with advanced technologies.
  • Collaborate with a diverse team of AI and application engineers.
  • Opportunities for professional development in cloud infrastructure and AI.
  • Access to modern tools and frameworks for efficient work.
  • Potential for contributing to high-impact AI projects with wide-reaching applications.
Full Job Description
The impact you'll make

We are looking for a hands-on Senior DevOps Engineer to build and operate the infrastructure powering our AI platforms, agent runtimes, web applications, backend services, APIs, and shared platform capabilities. Our environment includes a centralized LLM gateway, Python-based agent runtimes, RAG workers, MCP services, asynchronous processing, databases, caches, queues, and observability services.

You will work with AI engineers, application engineers, and architects who define technical designs, then independently translate those designs into reliable, scalable, secure, and observable production infrastructure across Microsoft Azure and Oracle Cloud Infrastructure.

What you'll do
  • Translate application and platform technical designs into production ready cloud infrastructure with minimal supervision.
  • Design, provision, operate, and troubleshoot Kubernetes environments, primarily Azure Kubernetes Service and Oracle Kubernetes Engine.
  • Support AI workloads including LiteLLM based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
  • Design and manage ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service to service communication.
  • Build and operate infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event driven workloads.
  • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
  • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tooling.
  • Implement end to end observability using metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
  • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
  • Create reusable infrastructure patterns that allow engineering teams to launch new services quickly and consistently.


What we're looking for
  • 7 or more years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a related role.
  • Strong hands on experience operating production Kubernetes environments and deep knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.
  • Strong Microsoft Azure experience, including AKS, networking, identity, storage, and monitoring. OCI experience is preferred, or demonstrated ability to work across cloud providers.
  • Strong cloud networking knowledge across virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.
  • Strong experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
  • Proven experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
  • Hands on experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or equivalent technologies.
  • Experience implementing production observability using OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar tools.
  • Strong Linux, systems, and production troubleshooting skills.


Application engineering knowledge

You do not need to be a full time application developer, but you should understand how modern backend systems work and be able to troubleshoot across application and infrastructure boundaries.
  • Working knowledge of Python, especially backend services built with frameworks such as FastAPI.
  • Familiarity with at least one additional language such as C#, Java, Go, JavaScript, or TypeScript.
  • Understanding of HTTP, HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead letter queues, and asynchronous processing.
  • Ability to read application logs and stack traces and diagnose latency, memory, CPU, connection, and dependency issues.


How you'll work

You will frequently receive a technical design for a new AI workload, application, backend service, or platform capability. From that design, you should be able to independently determine and implement the infrastructure needed to run it in production.
  • Determine the required cloud resources, Kubernetes configuration, namespaces, scaling model, and supporting services.
  • Configure ingress and egress, private connectivity, DNS, TLS, service communication, identities, and secrets.
  • Provision and operate dependencies such as PostgreSQL, Redis, RabbitMQ, storage, and other shared services.
  • Build the Jenkins and Bitbucket CI/CD flow for build, test, container publishing, deployment, validation, and rollback.
  • Define observability, health checks, dashboards, alerts, capacity monitoring, and operational runbooks before production launch.
  • Own infrastructure delivery through UAT and production, partnering with architects and engineers when design tradeoffs require discussion.


What will make you stand out
  • Experience supporting AI or machine learning platforms, LLM gateways, agent runtimes, RAG pipelines, or MCP services.
  • Experience with Cloudflare, Envoy, ArgoCD, GitOps, and OpenTelemetry.
  • Experience building reusable infrastructure platforms for high scale SaaS or customer facing applications.
  • Experience operating distributed backend systems using RabbitMQ, Redis, PostgreSQL, and similar technologies.


AI or machine learning infrastructure experience is helpful, but not required. Strong experience with Kubernetes, web and backend infrastructure, networking, queues, databases, CI/CD, observability, and production cloud operations is the foundation for this role.

#NetworkSolutions

About Newfold Digital

Newfold Digital is a leading web technology company that provides website design, domain registration, hosting, and other web services to small and medium-sized businesses. The company was founded in 2000 and is headquartered in Burlington, Massachusetts. Newfold Digital has a global presence with offices in the United States, Canada, Europe, and Asia. The company's mission is to help businesses succeed online by providing them with the tools and resources they need to build and maintain a professional online presence. Newfold Digital's services are designed to be easy to use and affordable, making them accessible to businesses of all sizes.
Learn more about Newfold Digital
Size
3,762 employees
Market Cap
$1.3 billion
Industry
Net Income
$18.5 million
Founded
1997
5 Year Trend
+12.1%
Revenue
$1.1 billion
NASDAQ

Similar Jobs

More Jobs at Newfold Digital

More Enterprise Technology Jobs

Find similar Senior DevOps Engineer, AI Platform jobs: