Mid-Level AWS Production Support EngineerDirect Full-Time
Lafayette, LA | Knoxville, TN | Columbia, SC | Birmingham, ALPosition Summary We are seeking an
AWS Production Support Engineer with
2-5 years of production/application support experience to provide 24x7 operational support for applications running in an AWS environment.
The ideal candidate will have strong troubleshooting and incident-management skills, hands-on experience investigating application and infrastructure alerts, and the ability to perform initial triage using AWS services, Splunk, OpenTelemetry, and application logs.
This role requires someone who is proactive, self-motivated, comfortable working in a production environment, and able to quickly determine whether an issue is isolated or part of a broader outage.
Key Responsibilities - Provide 24x7 production support for applications operating in an AWS environment.
- Monitor production alerts and investigate application or infrastructure incidents.
- Perform initial troubleshooting and triage using AWS, Splunk, OpenTelemetry, and application/system logs.
- Investigate batch and scheduled job failures, gather relevant information, and identify potential causes.
- Determine the severity and impact of production issues and escalate to appropriate SMEs or engineering teams when required.
- Troubleshoot application performance, availability, and operational issues.
- Support production deployments and releases.
- Execute and monitor basic CI/CD jobs and deployment processes.
- Review logs and monitoring data to determine whether an issue is isolated or part of a larger production outage.
- Work closely with development, infrastructure, cloud, and support teams during incident resolution.
- Maintain and continuously improve operational documentation and knowledge articles.
- Identify opportunities to improve production-support processes and operational efficiency.
Required Qualifications - 2-5 years of Production Support / Application Support experience.
- Experience supporting applications in a 24x7 production environment.
- Working knowledge of AWS and the ability to navigate AWS services for troubleshooting and investigation.
- Hands-on experience with Splunk for log analysis and troubleshooting.
- Strong experience with incident investigation, initial triage, troubleshooting, and escalation.
- Experience investigating application alerts, system logs, and production failures.
- Experience supporting and troubleshooting batch jobs / scheduled jobs.
- Strong analytical and problem-solving skills.
- Ability to determine whether an issue is isolated or part of a broader outage.
- Ability to work independently and escalate appropriately when deeper technical expertise is required.
- Strong communication skills and a proactive, self-learning mindset.
Preferred / Nice-to-Have Skills - OpenTelemetry
- AutoSys
- Apache
- RMJ
- CI/CD
- Production release support
- Application debugging
- Monitoring and observability tools
- Knowledge-base / operational documentation experience
Work ScheduleThis position supports a
24x7 production environment and requires flexibility to work rotating
12-hour shifts.
Typical schedule:
3.5 days on / 3.5 days off, including a
Sunday-Wednesday rotation, depending on production-support coverage requirements.
Candidates must be comfortable working flexible schedules as required to support production operations.
Ideal Candidate ProfileThe successful candidate will be a hands-on production-support professional who:
- Enjoys troubleshooting production issues.
- Can investigate alerts independently before escalating.
- Is comfortable working with ambiguity during operational incidents.
- Learns new applications and technologies quickly.
- Understands incident severity and escalation procedures.
- Can work effectively under pressure during production outages.
- Continuously looks for ways to improve support processes and documentation.
Interview Focus AreasCandidates should be prepared to discuss:
- A production incident they handled and their troubleshooting approach.
- How they investigate issues using AWS logs or Splunk.
- Steps taken when a batch or scheduled job fails.
- How they determine when an issue should be escalated.
- Experience supporting a 24x7 production environment.
- Examples of operational-process or documentation improvements they have implemented.
#M1
#DI-CB2
#L1 - KB1
Ref: #404-IT Pittsburgh