Enterprise Cloud Engineer IV - Observability

Wellmark, Inc.

$110K — $130K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years in cloud engineering, Site Reliability Engineering, DevOps, or observability engineering.
  • Experience with enterprise-scale observability platforms in complex production environments.
  • Proven track record of technical initiatives enhancing reliability and efficiency.
  • Familiarity with leading observability tools, including Dynatrace and Splunk.
  • Solid grasp of observability concepts and cloud-native systems across AWS, Azure, or GCP.
  • Advanced knowledge of Kubernetes, microservices, and modern architecture.
  • Experience with Infrastructure as Code and CI/CD practices.

Responsibilities

  • Design and implement observability capabilities across the enterprise.
  • Establish observability standards in collaboration with technical and business teams.
  • Lead the development and evolution of monitoring and telemetry platforms.
  • Promote best practices in reliability engineering and cloud operations.
  • Collaborate with stakeholders to drive data-informed decision-making.
  • Provide technical leadership in observability strategy and solution implementation.
  • Troubleshoot and resolve issues across multiple applications and systems.

Benefits

  • Opportunities for professional development and certification.
  • Access to the latest tools and technologies in observability.
  • Mentorship and leadership opportunities within technical communities.
  • Flexible working arrangements with minimal travel requirements.
  • Participation in shaping observability practices across the organization.
Full Job Description
Job Description

The Enterprise Cloud Engineer IV (Observability) serves as a senior technical engineer responsible for designing, implementing, and advancing enterprise observability capabilities across cloud, infrastructure, application, and AI-driven platforms. This role partners with architecture, security, platform engineering, and business stakeholders to establish observability standards, improve operational excellence, and drive reliability across the technology ecosystem.

The engineer will lead the design, implementation, and evolution of enterprise monitoring, logging, tracing, and telemetry platforms while promoting best practices in reliability engineering and cloud-native operations. This position provides technical leadership in defining observability strategy, implementing scalable monitoring solutions, reducing operational risk, and enabling data-driven decision making across the organization.

Qualifications

Preferred:
  • 7+ years of experience in cloud engineering, Site Reliability Engineering, DevOps, observability engineering, or related disciplines.
  • Proven experience designing and operating enterprise-scale observability platforms in complex production environments.
  • Demonstrated success leading technical initiatives that improved reliability, availability, scalability, and operational efficiency.
  • Experience with leading observability platforms including Dynatrace, Splunk, Datadog, Grafana, Prometheus, OpenTelemetry, New Relic, AppDynamics, Elastic, Honeycomb, or equivalent.
  • Strong understanding of observability principles including metrics, logging, distributed tracing, service health, SLIs, SLOs, error budgets, and reliability engineering.
  • Experience architecting observability solutions for cloud-native and distributed systems running in AWS, Azure, or GCP.
  • Advanced knowledge of Kubernetes, containers, microservices, networking, and modern application architectures.
  • Experience implementing Infrastructure as Code and CI/CD practices that integrate observability controls and instrumentation.
  • Experience supporting observability requirements for AI, machine learning, and intelligent agent platforms.
  • Strong consulting, stakeholder management, and influence skills.
  • Demonstrated ability to mentor engineers and lead technical communities of practice.
  • Industry certifications in cloud platforms and/or observability technologies.

Required:
  • Bachelor's Degree or direct and applicable work experience
  • Minimum 5 years' experience with system engineering/architecture with multiple platforms; including advanced scripting and coding experience. Leverage relationships with Engineers, Security and Application Design/Run teams to develop and champion system design standards and best practices. Lead the transformation of design, build and delivery patterns for cloud services. Strong understanding of DevOps / CICD methodologies and tools
  • Advanced experience with multiple scripting tools and concepts and utilization of them into an enterprise environment
  • Proven experience with a variety of virtualization technologies in multi-tenant, private and hybrid cloud environments. Hands on experience creating reusable solutions (patterns) leveraging 'design for failure' approaches
  • Demonstrated ability to collaborate across all technical and business teams in the process of defining IT Solution scope, business/system/non-functional requirements, and writing/executing test cases
  • Strong communication skills (both written and verbal) with the ability to influence and drive change as well as adapt communication to your audience, with experience presenting complex technical concepts to senior technical leaders
  • Advanced problem solving / trouble shooting experience to identify root cause of enterprise issues and provide innovative solutions
  • Understand risk and how enterprise level changes impact an organization. Mentor and coach others within team on impact of changes
  • Take action to assist in design and implementation of new processes, solutions and measurement of impact
  • Proactive in nature and someone that is naturally curious, inquisitive and willing to learn new skills
  • Travel less than 5%


Additional Information

a. Serve as an escalation point for all levels of Engineers regarding troubleshooting and issue resolution for multiple applications or systems. Responsible for troubleshooting and issue resolution for platforms and will work with Cloud Engineers, Developers, Security, Architects and other appropriate stakeholders as part of issue resolution.

b. Drive the enhancement of existing processes and patterns and lead the design for new processes and patterns. Coordinate and lead the efforts for building and installing services and platforms to support overall stability.

c. Develops and applies industry best practice technology, design and methodology approaches to design platform specific technical designs. Researches and recommends new emerging technologies, techniques and tools that will add value to the organization.

d. Lead design of recommended technical enhancements for existing services and processes. Leads process optimization and improvements; including advanced automation scripts to support and maintain a resilient and scalable cloud environment. Be available to assist with on call escalation as needed.

e. Ensure systems availability and performance. Inform leadership and/or Engineering peers of systems health and create incidents for remediation.

f. Engage with business and technical stakeholders and provide consultation and guidance on the design and implementation of both process and technical solutions. Will coordinate efforts with Engineering peers and other technical subject resources to effectively implement solutions that align with the business needs and are cost concise.

g. Act as escalation point for off-shore administration teams regarding technical design and cloud solution implementation; including review and quality assurance. Work with Engineer peers and Procurement to coordinate any vendor issues or escalations; including contract reviews as necessary.

h. Serve as escalation in resolution of issues related to Third Party Risk analysis and Vendor Management contractual obligations.

i. Create architectural patterns to verify alignment with code templates and patterns for Cloud Services. Draft and publish team documentation (e.g. process, run-books, diagrams etc.), and coordinate with other Engineer peers to maintain and ensure all is kept up to date and accurate.

j. Maintain compliance with all applicable legal, regulatory, licensing, governance, and contractual requirements by helping to establish and monitor effective processes and procedures. Stay up to date with technology policies and procedures; including updates for non-enterprise level systems, team processes and diagrams.

k. Other duties as assigned.

Similar Jobs

More Jobs at Wellmark, Inc.

More Enterprise Technology Jobs

Find similar Enterprise Cloud Engineer IV - Observability jobs: