Who You AreWe're seeking a senior DevOps Engineering leader to architect and operate secure, scalable Kubernetes platforms across AWS and AWS GovCloud, including FedRAMP-authorized environments. This person will drive zero-trust security, Terraform and GitOps automation, observability, reliability, disaster recovery, and cost governance for both multi-tenant and dedicated enterprise deployments. They'll also provide technical leadership, mentor engineers, and partner closely with Security and Engineering to evolve Ontic's platform strategy.
Key ResponsibilitiesEnterprise Platform Architecture- Lead the design and governance of scalable, resilient, Kubernetes platforms across approved cloud environments and FedRAMP authorization boundaries
- Define standardized blueprints for multi-tenant and dedicated enterprise deployments
- Architect high-availability, multi-AZ, and disaster recovery strategies aligned with RTO/RPO objectives
- System inventory, authorized-boundary ownership, significant-change assessment, and approved-cloud-boundary wording.
DevSecOps & Zero-Trust Security- Establish secure-by-design cloud architecture and workload identity models
- Enforce RBAC, IAM governance, encryption standards, and least-privilege principles
- Integrate supply-chain security, policy enforcement, and compliance validation into CI/CD pipelines.
- Ensure infrastructure meets applicable FedRAMP and enterprise security-control requirements.
- Vulnerability remediation, POA&M/exception tracking, and incident-response exercises.
AI-Enabled Operations & Platform Intelligence- Implement AIOps capabilities for intelligent alerting, anomaly detection, and controlled remediation through approved runbooks, policy gates, and auditable change controls.
- Enhance observability maturity using metrics, logs, and distributed tracing
- Support AI/ML workload infrastructure, including scalable compute and secure model deployment
- Leverage automation and analytics to reduce MTTR and operational toil
Infrastructure Automation & GitOps- Drive infrastructure standardization using Terraform, Helm, and GitOps workflows
- Design safe deployment strategies (blue/green, canary, progressive delivery)
- Improve infrastructure drift detection and environment reproducibility
Reliability & Performance Engineering- Define SLO/SLI frameworks and reliability benchmarks
- Optimize autoscaling (HPA/VPA/Cluster Autoscaler) and workload efficiency
- Improve performance and scalability of distributed stateful systems
- Backup restoration and recurring disaster-recovery testing against RTO/RPO.
FinOps & Cost Governance- Implement cloud cost observability and allocation strategies
- Optimize compute, storage, and networking costs across multi-cloud environments
- Establish cost governance models aligned with business growth
Leadership & Governance- Mentor senior engineers and promote DevOps culture across teams
- Drive architecture reviews and production readiness assessments
- Partner with security and engineering to align platform strategy with enterprise objectives
Qualifications- 8-10+ years of experience in DevOps / Platform Engineering in enterprise or high-scale SaaS environments.
- Proven experience in architecting and operating production environments in AWS and AWS GovCloud; GCP or OCI experience is a plus.
- Demonstrated experience maintaining a FedRAMP Moderate/High production environment, including continuous monitoring, vulnerability remediation, incident response, configuration management, and audit evidence.
- Experience owning infrastructure patterns and changes within the authorized boundary, including system inventory, data-flow implications, inherited controls, and third-party services.
- Strong experience designing and operating AWS Landing Zone architectures, including multi-account governance and guardrails.
- Advanced Terraform experience building reusable, secure infrastructure modules and standardized environments.
- Deep understanding of AWS and AWS GovCloud networking, including VPC design, segmentation, routing, private connectivity, and security controls.
- Deep expertise in Kubernetes-based platform design and large-scale cluster operations.
- Strong experience implementing GitOps workflows (ArgoCD) and CI/CD automation (Jenkins, GitLab).
- Hands-on production experience with distributed systems, including MongoDB, Elasticsearch, Kafka, Redis, and ArangoDB.
- Strong understanding of IAM, encryption at rest and in transit, and zero-trust access models.
- Experience implementing enterprise observability stacks.
- Experience designing multi-tenant SaaS platforms and dedicated workload environments.
- Exposure to service mesh and secure east-west traffic management.
- Exposure to AIOps is a plus
This role is based in Austin, Texas, and requires work from the Austin office three days a week.This role supports a FedRAMP environment and requires that work be performed within the United States by individuals authorized to work in the U.S. and who are U.S. citizens.Ontic Benefits & PerksCompetitive Salary
Medical, Vision & Dental Benefits
401k
Stock Options
HSA Contribution
Learning Stipend
Flexible PTO Policy
Quarterly company ME (mental escape) days
Generous Parental Leave policy
Home Office Stipend
Mobile Phone Reimbursement
Home Internet Reimbursement for Remote Employees
Anniversary & Milestone Celebrations