Role OverviewPlatform Engineers at Fractal are responsible for developing, maintaining, and optimizing the infrastructure that powers enterprise analytics platforms. In this role, you will manage dbt and Airflow instances, handle system upgrades, resolve infrastructure issues, and support the development and deployment of custom applications - all while adhering to rigorous security and compliance standards. This is a high-impact engineering role at the heart of a modern cloud-native data stack.
Key Responsibilities - Develop, deploy, and maintain dbt and Airflow instances in production
- Manage Kubernetes clusters and containerized applications
- Implement and maintain CI/CD pipelines for data applications
- Handle system upgrades, version migrations, and maintenance windows
- Troubleshoot and resolve infrastructure issues affecting pipeline execution
- Implement monitoring, alerting, and logging for infrastructure health
- Manage infrastructure as code (IaC) using Terraform or CloudFormation
- Provide security hardening and ensure compliance with InfoSec policies
- Support ML compute infrastructure and scaling requirements
- Collaborate with data engineers on pipeline optimization and troubleshooting
- Document infrastructure architecture, procedures, and runbooks
- Participate in on-call rotation for infrastructure support and incident response
Required SkillsLatAm Position (G7): 6-10 years of platform/infrastructure engineering experience - 3+ years with Kubernetes and container orchestration
- 3+ years with AWS or similar cloud platforms
- 2+ years managing data infrastructure (Airflow, dbt, or similar)
- Expert-level Kubernetes administration and orchestration
- Advanced AWS knowledge: EC2, S3, RDS, Lambda, VPC, IAM, CloudWatch
- Proficiency in Infrastructure as Code: Terraform, CloudFormation, or Ansible
- Strong Python and Bash scripting skills
- Experience with dbt and Airflow infrastructure and scaling
- Knowledge of networking concepts: proxies, load balancers, DNS, VPNs
- Understanding of security best practices: SSO, OIDC, IAM, encryption
- Monitoring and logging tools: Prometheus, Grafana, CloudWatch, ELK stack
Preferred Skills - Strong problem-solving and debugging skills
- Excellent documentation practices
- Ability to work under pressure during incidents and outages
- Collaborative approach to supporting multiple teams
- Commitment to reliability and operational excellence
- Bachelors degree in Computer Science, Engineering, or related field
- Masters degree or relevant cloud certifications preferred