External candidates: In order for your application to be correctly processed please sign-in before you apply
Internal candidates: Please go to Workday and click "Find Jobs" link under Career
Thank you for considering opportunities with us!
Job TitleSoftware Engineer III - API Hub - Remote
Requisition NumberR7879 Software Engineer III - API Hub - Remote (Open)
LocationPhoenix (Non-Service)
Additional LocationsJob InformationWe are actively hiring for a Software Engineer III - API Hub - Remote.
Your Role: As a Software Engineer III, API Hub Operations, you will be responsible for the operational reliability, engineering, automation, and support of CSAA's enterprise API platform and cloud-native integration services. You will leverage modern software engineering, DevOps, Site Reliability Engineering (SRE), and cloud technologies to ensure highly available, secure, scalable, and observable API services across enterprise applications.
You will serve as a technical resource supporting API platforms deployed on AWS, partnering closely with Software Engineers, Solution Architects, Security Engineers, Infrastructure Engineering, Product Teams, and Business Stakeholders to deliver reliable API solutions and operational excellence.
Your Work: API Platform Operations & Engineering- Support and maintain enterprise API management and gateway platforms.
- Manage API onboarding, deployment, versioning, publishing, and lifecycle activities.
- Configure and maintain API policies, routing, authentication, authorization, throttling, and traffic management.
- Support API governance standards and operational guidelines.
- Analyze API performance, availability, latency, and usage trends.
- Collaborate with development teams to improve API reliability, scalability, and security.
Cloud Engineering & Platform Support- Design, develop, deploy, and maintain cloud-native solutions using AWS services.
- Support platform components running on AWS including API Gateway, EKS, Lambda, S3, CloudWatch, IAM, Route53, SNS/SQS, and EventBridge.
- Automate operational tasks through scripting, infrastructure-as-code, and cloud-native tooling.
- Implement platform resiliency, disaster recovery, backup, and recovery solutions.
- Troubleshoot infrastructure, networking, application, and cloud integration issues.
Site Reliability Engineering (SRE)- Implement and support observability solutions including monitoring, logging, alerting, dashboards, and distributed tracing.
- Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational reliability metrics.
- Lead incident troubleshooting, root cause analysis (RCA), and problem management activities.
- Participate in production support, on-call rotations, and major incident response activities.
- Drive automation efforts to reduce operational toil and improve platform stability.
- Support capacity planning, performance optimization, and reliability improvements.
DevOps & Automation- Build and maintain CI/CD pipelines and deployment automation.
- Implement continuous integration, automated testing, static code analysis, and release automation.
- Support containerized application deployments using Kubernetes and Helm.
- Develop infrastructure automation using Terraform, CloudFormation, or similar technologies.
- Leverage approved AI tools and automation capabilities to improve engineering productivity and operational efficiency.
Design, Development & Technical Leadership- Design, develop, test, and maintain scalable software components and platform services.
- Apply deep technical knowledge of API platform operations and cloud services to support platform reliability, resolve complex issues, and drive continuous improvement. .
- Provide technical guidance during design reviews, implementation planning, and troubleshooting activities.
- Review system architectures and provide recommendations to improve availability, performance, and maintainability.
- Create and maintain operational runbooks, support procedures, and engineering documentation.
Problem Solving & Operational Excellence- Proactively identify operational risks, recurring incidents, and service improvement opportunities.
- Diagnose complex production issues across applications, APIs, cloud infrastructure, and integrations.
- Collaborate with vendors and internal teams to resolve all issues and service disruptions.
- Translate business requirements into scalable, secure, and supportable technical solutions.
- Lead root cause investigations and implement preventative measures.
Innovation & Continuous Improvement- Continuously improve platform reliability, automation, observability, and operational processes.
- Evaluate and adopt emerging cloud, DevOps, AI/ML, and platform engineering technologies.
- Identify opportunities to optimize operational costs, reduce manual processes, and enhance customer experience.
- Contribute to engineering standards, operational guidelines, and platform governance initiatives.
Collaboration & Communication- Collaborate with Architecture, Security, Infrastructure, Product, and Development teams to deliver reliable API services.
- Communicate technical issues, risks, and operational status clearly to both technical and non-technical audiences.
- Influence technical decisions and platform direction through expertise and collaboration.
- Participate in planning, estimation, release readiness, operational reviews, and post-incident reviews.
Required Experience, Education and Skills- 4+ years of proven experience in Software Engineering, Platform Engineering, DevOps, SRE, or Cloud Operations.
- Bachelor's Degree in Computer Science, Engineering, Information Technology, or related field; or equivalent combination of education and experience.
- Experience supporting production applications and customer-facing platforms.
- Success working within Agile and DevOps environments.
Demonstrates knowledge of and experience with relevant technologies and platforms aligned to the role, which may include:
- Cloud-native services (e.g., AWS Lambda, DynamoDB, S3)
- Build, integration, and deployment tools (e.g., Jenkins, Maven)
- Programming languages and frameworks (e.g., Java/J2EE, Python, Node.js, JavaScript/TypeScript, React)
What would make us excited about you?- Experience supporting enterprise API platforms at scale.
- Strong AWS and Kubernetes operational expertise.
- Experience leading production incident response and root cause analysis efforts.
- Demonstrated success improving operational reliability and automation.
- Familiarity with Site Reliability Engineering principles and practices.
- Ability to influence technical decisions across multiple teams.
- Experience leveraging AI and automation technologies to improve operational efficiency.
- Strong communication and team member relationship skills.
- Success working in highly regulated, customer-facing environments.
- Actively shapes our company culture (e.g., participating in employee resource groups, volunteering, etc.)
- Lives into cultural norms (e.g., willing to have cameras when it matters: helping onboard new team members, building relationships, etc.)
- Travels as needed for role, including divisional / team meetings and other in-person meetings
- Fulfills business needs, which may include investing extra time, helping other teams, etc
Please note we are hiring for this role remote anywhere in the United States with the following exceptions: Hawaii and Alaska.
The national average salary range for this position is $105,345.00-$117,050.00. However, we have a location-based compensation structure. Our salary ranges vary and are calculated based on work location. The starting pay range for this position across all the states we hire in is $105,345.00-$140,550.00. This role also includes an opportunity for a company-wide annual discretionary bonus, through our Annual Incentive Plan (AIP), of up to 8% of eligible pay.
This job posting will be unposted on Sat, 22 Aug 2026.