7+ years designing and managing production cloud infrastructure, with 5+ years at a Senior level or above
Expertise in AWS or GCP, with hands-on experience in both
Proficient in Kubernetes, Terraform, and CI/CD, including GitOps with ArgoCD or similar
Experience with cloud networking and security, including VPCs, IAM, and multi-region deployments
Familiarity with enterprise deployments in regulated industries like finance or healthcare
Knowledge of compliance frameworks such as SOC 2 or ISO 27001
Strong coding skills in Python and Bash for automation tasks
Responsibilities
Deploy cloud-native platforms into client clouds on AWS, GCP, and Azure, addressing security and architecture concerns
Co-manage existing infrastructure in AWS and GCP, including EKS and Aurora Postgres
Operate the research platform for evaluations and training using various managed services
Develop CI/CD, observability, and alerting systems to enhance deployment efficiency
Create rapid prototypes and refine them for production use
Share compliance and security documentation to support audit processes
Contribute to platform development and internal GenAI tools
Manage infrastructure as code using Terraform, documenting decision trade-offs
Benefits
Participation in the 'Friends of Rational Dynamics' Candidate Referral Program with potential bonuses
Opportunities for professional growth in a cutting-edge AI environment
Collaborative work culture with cross-functional teams
Exposure to diverse cloud technologies and compliance frameworks
Engagement in innovative projects that impact institutional investors' analytical capabilities
Full Job Description
About the role
Institutional investors want AI for their hardest analytical work, running inside their cloud accounts and regions, under their security and compliance rules. Rational Dynamics builds that AI. Infrastructure decides whether we can ship it: each client brings a different cloud, network, identity provider, and security review. You'll be working across client deployments, internal platform, and the research infrastructure, collaborating with product development and science teams.
What you'll do
Deploy our cloud-native platform into client clouds on AWS, GCP, and Azure, including multi-region and data-residency setups. Answer the security and architecture questions client IT teams raise
Co-own existing infrastructure in AWS and GCP: EKS, Aurora Postgres, Cloud Run, networking, IAM, secrets, and GitOps delivery
Run the research platform for evals and training: proprietary, open-source, managed services such as Vertex AI, SageMaker, and Bedrock, with GPU compute and storage infrastructure
Build CI/CD, observability, and alerting that let a small team ship daily and catch problems before clients do
Stand up rapid prototypes and harden existing prototypes for production. Give Forward Deployed Engineers tooling to do the same in client environments
Share compliance (e.g. SOC 2, ISO 27001, CAIQ) and security work so audit evidence comes from how the platform runs
Contribute to platform development, MLOps, and internal GenAI tools
Manage infrastructure as code in Terraform, write down the trade-offs behind each decision, and track reliability, uptime, and cost
What you bring
7+ years designing and running production cloud infrastructure, 5+ of them at Senior level or above
Depth in AWS or GCP, and hands-on experience with both
Strong Kubernetes, Terraform, and CI/CD, including GitOps with ArgoCD or similar
Cloud networking and security: VPCs, private connectivity, IAM, zero-trust access, encryption, secrets, and multi-region deployments with data-residency requirements
Deployments into enterprise customer clouds in finance, healthcare, defense, or another regulated industry
Work under SOC 2, ISO 27001, PCI DSS, or a similar control framework
Core infrastructure built from scratch at an early-stage company, or owned end to end on a small team
Python and Bash for automation, and daily use of agentic coding tools such as Claude Code
Clear communication about designs and trade-offs with engineers, scientists, and client security teams, and comfort with shifting priorities
Nice to have
Azure (AKS, VNets, Entra ID, ML/AI) or Oracle Cloud (OCI)
ML and GPU infrastructure: Kubeflow, Vertex AI, SageMaker, Bedrock, MLflow, Ray, KServe, Slurm
Infrastructure for LLM and agent systems: model serving, vector databases, and training and eval pipelines
Workflow tools such as Temporal, Airflow, Argo Workflows, Tekton
Full-stack or agentic application development, or ML and RL research
Go or another systems language
A record of running production systems against reliability and uptime targets
"Friends of Rational Dynamics" Candidate Referral Program
If you have a great candidate in mind for this role and would like to have the potential to earn $7,500 to $15,000 if your referred candidate is successfully hired and employed by Rational Dynamics, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Rational Dynamics Referral Bonus Program.