5-7 years of experience in production environment management
Strong knowledge of Application Performance Monitoring and Optimization strategies
Proficiency in SQL and Oracle Database management
Hands-on experience with CI/CD processes and DevOps best practices
Ability to design and implement Monitoring and Alerting mechanisms
Proven track record in incident response and system recovery
Familiarity with system design consulting and capacity planning
Responsibilities
Plan and manage all aspects of the Production Environment
Define and implement strategies for Application Performance Monitoring
Respond to incidents and enhance platform capabilities based on feedback
Support code deployments across multiple lower environments
Design and standardize Monitoring and Alerting for applications
Engage in the entire lifecycle of services from design to refinement
Support services pre-launch through system design and capacity planning
Benefits
Opportunity to work with a modern tech stack
Collaborative and innovative work environment
Focus on career development and learning opportunities
Involvement in strategic decision-making processes
Access to cutting-edge tools and technologies
Full Job Description
Job Description
The Role
Plan, manage, and oversee all aspects of a Production Environment
Define strategies for Application Performance Monitoring, Optimization in Production environment
Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
Support deployment of code into multiple lower environments.
Design, develop and standardize Monitoring and Alerting mechanisms for the supported applications.
Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.
Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation and refinement.
Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.