We're looking for a Senior DevOps Engineer to power the reliability, security, and efficiency of the world's most impactful weather platform. Your work will be guided by four core pillars- Security, Cost, SLOs, and Developer Experience- with AI acting as a force multiplier across each. You'll build self-service platforms that give developers and weather scientists true independence, weave AI into how we operate, and work side-by-side with R&D to push performance and scale further. Our environment spans two worlds: cloud-native product services on Kubernetes, and scientific computing on HPC clusters- and you'll help both thrive. You'll evolve our cloud infrastructure to match the pace of the business, hold the line on cost, and stay close to production through on-call. The people who thrive here bring a product mindset, take ownership without waiting to be asked, and leave the people and systems around them better than they found them.
As a DevOps Engineer at Tomorrow.io, You'll:- Develop and adopt AI-powered tools to make Development and Operations processes more efficient
- Collaborate with weather scientists, engineers, and Spacecraft Mission Operations Engineers to optimize service performance, reliability, scale, security, and cost
- Evolve and maintain adaptive cloud infrastructure to support our business strategy and enable smooth growth at scale
- Build self-service platforms for scientists and developers to work independently
- Support scientific computing workloads on HPC clusters (SLURM) alongside our cloud-native Kubernetes platforms
- Introduce and integrate MLOps practices for GPU-based model deployment on Kubernetes
- Maintain Production availability by participating in DevOps on-call shifts
What you bring:- At least 6 years of experience as a Platform/DevOps/SRE Engineer in a containerized cloud environment experienced with AWS, GCP, or Azure and IaC, such as Terraform or Crossplane
- Experience in fast-growing, cloud-native startup or scale-up environments
- Strong sense of ownership and accountability for service reliability
- Daily, hands-on use of AI coding agents; experience building agentic workflows is a plus
- 10X mindset - always looking for the fastest, smartest path to a high-quality result
- Daily use of AI coding agents (Claude Code, Copilot, etc.)- must; building agentic DevOps workflows- a plus
- Experience with CI/CD tools and deployment methodologies in Kubernetes
- Experience implementing and customizing monitoring systems (Datadog, Prometheus, Grafana, ELK Stack)
- Experience working in an agile environment with high-velocity teams
- Proficiency with scripting languages like Python, Node.js, and Go
- Adaptable problem-solving mindset - thriving in changing environments and requirements
- Excellent written and verbal communication skills, with the ability to collaborate effectively across distributed teams, time zones, and multiple R&D stakeholders
Bonus:- Experience with HPC / scientific computing- e.g., Slurm, AWS ParallelCluster, Azure CycleCloud
- Familiarity with parallel filesystems- e.g., Lustre, NFS
If you take pride in what your infrastructure enables- faster science, reliable operations, a better product and you'd rather ship something useful this week than something perfect this quarter, this is the place for you. You'll join a small global team with real ownership and room to grow, help shape how a fast-moving company builds and operates with AI, and work alongside engineers and scientists solving problems most infrastructure teams never get near- from forecast models to satellites.
If you have reached this point and you are super excited but not sure you check all the boxes - we still want to speak with you! Your passion is priceless. Other things can be learned.
This position requires access to technology that is controlled under U.S. export control laws and regulations. Accordingly, this position is restricted to U.S. citizens, permanent residents and protected individuals unless and until any required licenses are obtained.
The anticipated salary range for this role is $160-180k, subject to local market and candidates skills and experience. Comprehensive health benefits, unlimited paid time off and other benefits included.