Job Summary
We are seeking a proactive and technically skilled Site Reliability Engineer (SRE) to support the production operations of a mission-critical healthcare ecosystem. The ideal candidate will be responsible for ensuring the reliability, availability, and operational excellence of healthcare applications, scheduled workloads, and third-party integrations.
The role involves production incident management, Tidal support, troubleshooting healthcare applications, maintaining automation scripts, performing minor application enhancements, and supporting inbound/outbound data exchanges with external vendors. The engineer will work closely with application development, infrastructure, cloud, database, and business teams to ensure uninterrupted healthcare services.
Key Responsibilities:
Production Support
• Provide production support for healthcare applications.
• Monitor production environments and respond to incidents based on SLA.
• Perform root cause analysis (RCA) and implement preventive measures.
• Participate in on-call rotations and major incident management.
• Ensure high availability and operational stability of critical healthcare systems.
• Monitor scheduled Tidal jobs and workflows.
• Investigate and resolve Tidal job failures.
• Restart, rerun, or modify job executions following SOPs.
• Analyze job dependencies and downstream impacts.
• Coordinate with application and infrastructure teams during failures.
• Recommend automation and self-healing opportunities.
• Support GuidingCare, Clinical applications, Care Management platforms & Integration services
• Batch processing applications
Responsibilities include:
• ServiceNow Incident triage
• Experience in Splunk and Dynatrace analysis
• Application troubleshooting (GuidingCare, Clinical applications, Care Management platforms & Integration services)
• Configuration validation
• Performance issue investigation
• Functional verification after fixes
• Maintain and update operational scripts (PowerShell, Python, Bash, SQL, etc.).
• Develop automation to reduce repetitive operational activities.
• Improve monitoring and alerting through scripting.
• Support CI/CD operational automation where applicable.
• Perform small code changes and bug fixes.
• Support configuration changes.
• Participate in release deployments.
• Validate production fixes.
• Maintain technical documentation.
• Monitoring scheduled file transfers
• Troubleshooting failed transmissions
• Validating file integrity
• Managing SFTP connections
• Coordinating with third-party vendors
• Ensuring timely delivery of healthcare data
• Supporting HIPAA-compliant data exchange processes
• Monitor applications using enterprise monitoring tools.
• Investigate alerts before business impact occurs.
• Perform health checks.
• Drive continuous improvements in platform reliability.
• Identify recurring issues and recommend permanent fixes.
• Documentation & maintain SOPs and knowledge articles.
• Document RCA findings and operational procedures.
Required Technical Skills
• Experience supporting healthcare applications
• Knowledge of healthcare workflows
• Understanding of HIPAA compliance
• Experience with healthcare integrations
• Experience with Tidal Enterprise Scheduler
• Job scheduling concepts
• Batch processing
• Workflow automation
• Linux & Windows Server
• PowerShell, Python & Bash/Shell scripting
• Azure Cloud (preferred)
• Storage & Networking fundamentals
• Experience with one or more:
• Dynatrace
• Splunk
• Azure Monitor
• App Insights
• Databases
• SQL Server
• Oracle
• Basic query optimization
• Stored procedures
• File Transfer
• SFTP
• FTPS
• File encryption
• Secure integrations
• Version Control
• Git or Azure DevOps or GitHub
Preferred Skills
• Experience with ServiceNow
• Knowledge of CI/CD pipelines
• REST API troubleshooting
• Experience with healthcare payer/provider applications
• Familiarity with automation and AI-assisted operations (AIOps/Agentic AI)
Soft Skills
• Strong analytical and troubleshooting skills
• Excellent incident management capabilities
• Ability to work under pressure during production outages
• Strong communication and stakeholder management
• Customer-focused mindset
• Continuous improvement attitude
• Excellent documentation skills
• Ability to work independently and collaboratively
Experience
• Experience supporting mission-critical enterprise applications in a healthcare environment is highly preferred.
• Agentic AI-driven production operations
• Auto-remediation and self-healing solutions
• Azure Logic Apps, Azure Functions, and Automation Accounts
• Healthcare EDI transactions & integrations
• Production support for GuidingCare or similar care management platforms