7+ years designing and managing production cloud infrastructure, with 5+ years at a Senior level or above.
Expertise in AWS or GCP, with hands-on experience in both.
Proficient in Kubernetes, Terraform, and CI/CD practices, including GitOps with ArgoCD or similar.
Strong understanding of cloud networking and security protocols, including VPCs and zero-trust access.
Experience deploying into enterprise customer clouds in regulated industries like finance or healthcare.
Familiarity with compliance frameworks such as SOC 2 and ISO 27001.
Proficient in Python and Bash for automation tasks.
Responsibilities
Deploy cloud-native platforms into client clouds on AWS, GCP, and Azure, addressing security and architecture concerns.
Co-manage existing infrastructure in AWS and GCP, including EKS and Aurora Postgres.
Operate the research platform for evaluations and training using various managed services.
Build CI/CD pipelines and observability tools to ensure smooth daily operations.
Develop rapid prototypes and enhance existing ones for production readiness.
Share compliance and security documentation to support audit processes.
Contribute to platform development and MLOps initiatives.
Benefits
Opportunity to work with cutting-edge AI technologies in a dynamic environment.
Collaborative culture with cross-functional teams including product development and research.
Engagement in high-impact projects for institutional investors.
Potential for professional growth in a rapidly evolving field.
Involvement in compliance and security practices that enhance career expertise.
Full Job Description
About the role
Institutional investors want AI for their hardest analytical work, running inside their cloud accounts and regions, under their security and compliance rules. Rational Dynamics builds that AI. Infrastructure decides whether we can ship it: each client brings a different cloud, network, identity provider, and security review. You'll be working across client deployments, internal platform, and the research infrastructure, collaborating with product development and science teams.
What you'll do
Deploy our cloud-native platform into client clouds on AWS, GCP, and Azure, including multi-region and data-residency setups. Answer the security and architecture questions client IT teams raise
Co-own existing infrastructure in AWS and GCP: EKS, Aurora Postgres, Cloud Run, networking, IAM, secrets, and GitOps delivery
Run the research platform for evals and training: proprietary, open-source, managed services such as Vertex AI, SageMaker, and Bedrock, with GPU compute and storage infrastructure
Build CI/CD, observability, and alerting that let a small team ship daily and catch problems before clients do
Stand up rapid prototypes and harden existing prototypes for production. Give Forward Deployed Engineers tooling to do the same in client environments
Share compliance (e.g. SOC 2, ISO 27001, CAIQ) and security work so audit evidence comes from how the platform runs
Contribute to platform development, MLOps, and internal GenAI tools
Manage infrastructure as code in Terraform, write down the trade-offs behind each decision, and track reliability, uptime, and cost
What you bring
7+ years designing and running production cloud infrastructure, 5+ of them at Senior level or above
Depth in AWS or GCP, and hands-on experience with both
Strong Kubernetes, Terraform, and CI/CD, including GitOps with ArgoCD or similar
Cloud networking and security: VPCs, private connectivity, IAM, zero-trust access, encryption, secrets, and multi-region deployments with data-residency requirements
Deployments into enterprise customer clouds in finance, healthcare, defense, or another regulated industry
Work under SOC 2, ISO 27001, PCI DSS, or a similar control framework
Core infrastructure built from scratch at an early-stage company, or owned end to end on a small team
Python and Bash for automation, and daily use of agentic coding tools such as Claude Code
Clear communication about designs and trade-offs with engineers, scientists, and client security teams, and comfort with shifting priorities
Nice to have
Azure (AKS, VNets, Entra ID, ML/AI) or Oracle Cloud (OCI)
ML and GPU infrastructure: Kubeflow, Vertex AI, SageMaker, Bedrock, MLflow, Ray, KServe, Slurm
Infrastructure for LLM and agent systems: model serving, vector databases, and training and eval pipelines
Workflow tools such as Temporal, Airflow, Argo Workflows, Tekton
Full-stack or agentic application development, or ML and RL research
Go or another systems language
A record of running production systems against reliability and uptime targets
"Friends of Rational Dynamics" Candidate Referral Program
If you have a great candidate in mind for this role and would like to have the potential to earn $7,500 to $15,000 if your referred candidate is successfully hired and employed by Rational Dynamics, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Rational Dynamics Referral Bonus Program.