FIS

VP, Site Reliability Engineering (SRE & Observability Platform)

FIS$180K — $220K *
Jay, FL 32565In-Person
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 15+ years in software engineering, with at least 7 in senior engineering leadership roles for SRE, platform, or AI organizations.
  • Proven record of enhancing reliability and decreasing production incidents within large, complex multi-team environments.
  • Experience in building and managing production AI and agentic systems, demonstrating knowledge of LLMs and AI observability frameworks.
  • Established AI governance for production agents, including safety measures and evaluation protocols.
  • Expertise in observability tools (OpenTelemetry), incident management, and cloud infrastructure, particularly Kubernetes.
  • Skilled in operating a platform as a product emphasizing usability and engagement.
  • Ability to influence organizational change within a large matrixed environment.

Responsibilities

  • Set strategic vision for reliability, observability, and agentic engineering within the enterprise.
  • Build and manage a high-performing platform organization with focus on SRE and AI engineering.
  • Execute a phased roadmap for reliability and observability improvements across teams.
  • Drive the platform's agent-native function adoption, ensuring governance and safety measures are in place.
  • Equip product engineers with AI agents for streamlined observability and incident management.
  • Implement measurable reliability standards, including SLOs and error budgets, across the organization.
  • Engage and influence teams through enablement and shared benefits of the paved road, rather than mandates.

Benefits

  • Leadership visibility working directly with C-level executives.
  • Opportunity to define and shape a forward-thinking SRE organization focused on AI-driven solutions.
  • Ownership of significant projects with lasting impact on engineering culture.
  • Access to committed funding for strategic initiatives benefiting the broader engineering community.
  • A collaborative environment with a strong emphasis on professional development and talent growth.
Full Job Description
Job Description

About the role:

Production reliability is a board-level priority, and we are placing it in one hands-on leader. We are hiring a Vice President of Engineering to found and lead our Site Reliability Engineering (SRE) platform organization and own the mission end to end: raise observability maturity and materially reduce production incidents across the engineering estate. You will do it by making agentic engineering the core method of the function - using AI agents to run the platform itself, and putting AI agents in the hands of product engineers so they can meet the company's observability requirements and SRE goals with far less manual effort.

This is a builder's role with executive visibility. You will stand up a central platform team that treats reliability and observability as a product, deliver an aggressive multi-phase roadmap, and change how thousands of engineers instrument, operate, and take ownership of what they ship - in a large, complex, tech-debt-heavy environment where the winning strategy is paved roads over mandates. Central to that is a two-part AI mandate: making this function fully agent-native, and delivering AI agents to product engineers so they can fulfill their observability requirements and the company's SRE goals with far less manual effort.

You will report directly to the SVP of Platform Engineering and partner with leaders across engineering, product, and the executive team. The mission has a named executive sponsor and committed funding; your job is to turn that mandate into outcomes.

The AI mandate: two transformations you will lead:

Agentic engineering is not a side initiative in this role - it is central to how the function operates and to the value it delivers to the enterprise. You will be accountable for two connected transformations:

1. Transform this function to be fully agent-native

Re-found the platform organization around agentic engineering. Agents - not manual toil - should carry the load of instrumenting legacy code, investigating incidents, generating configuration, assisting on-call, and, over time, executing guarded remediation. You will build the agent control plane, guardrails, and evaluation harnesses that make this safe, and turn the function into the company's proof point for what disciplined, agent-first engineering looks like at scale.

2. Put AI agents in the hands of product engineers for observability and SRE outcomes

Deliver AI agents to product engineers as part of the paved road so they can meet the company's observability requirements and SRE goals with far less manual effort - agents that instrument their services, generate SLOs, dashboards, and alerts, investigate incidents, and assist on-call. These agents, and the reliability, evaluation, guardrails, and cost governance behind them, are how product teams hit reliability targets at scale. This role owns the SRE platform: the agents it provides serve observability and reliability, not product-feature development.

What you'll own:

  • The SRE platform organization - a central platform team plus embedded/partner SREs, an enablement function, and a cross-cutting reliability champions network.
  • The observability & reliability platform - telemetry pipeline (OpenTelemetry), metrics/logs/traces backends, dashboards and alerting, the SLO and error-budget system, incident management, and RCA/correlation.
  • The paved road - shared instrumentation SDKs, templates, and dashboards/alerts/SLOs-as-code that make golden-signal observability near-automatic.
  • The agentic engineering stack - internal agents for instrumentation, investigation/RCA, config generation, on-call, and guarded remediation - plus the agent control plane, guardrails, and evaluation that keep them safe.
  • The agentic paved road for product engineers - AI agents delivered to product teams to instrument services, generate SLOs/dashboards/alerts, investigate incidents, and assist on-call - so they meet observability requirements and SRE goals with minimal manual effort, backed by guardrails and evaluation.
  • Reliability & AI governance - error-budget policy, blameless incident and postmortem practice, AI safety and human-in-the-loop controls, and the metrics reported to senior leadership.
  • The roadmap, budget, and vendor strategy - an ~18-month phased plan, the operating budget, tooling selection, and cost governance for both telemetry and AI workloads.


Key responsibilities:

  • Set the strategy and vision. - Own the enterprise strategy for reliability, observability, and agentic engineering. Define what "reliability as a product" and "agent-native engineering" mean here, and keep both tied to business outcomes.
  • Build and lead the organization. - Recruit, structure, and grow a high-performing platform organization spanning SRE, platform, and AI engineering - hiring and developing senior, staff, and principal talent and the managers who lead them.
  • Deliver the roadmap. - Execute the phased plan - mobilize and instrument, build foundations, scale and standardize the paved road with SLO coverage and error-budget policy, then bring proactive and agentic capabilities to production - on aggressive, overlapping timelines.
  • Make the function agent-native. - Drive adoption of internal agents across the platform's own work, with least-privilege access, blast-radius limits, human-in-the-loop controls, and evaluation before any increase in autonomy.
  • Put agents in product engineers' hands. - Deliver AI agents through the paved road that let product engineers meet observability requirements and SRE goals - instrumenting services, standing up SLOs, and resolving incidents - with far less manual effort, backed by guardrails and evaluation.
  • Make reliability measurable. - Stand up SLOs and error budgets, modern incident management, and blameless postmortems; establish error-budget policy in partnership with leadership.
  • Drive adoption across the estate. - Win teams over with paved roads and lighthouse wins, not mandates - through reliability reviews, enablement, office hours, and published before/after results.
  • Own the tooling and its economics. - Select and evolve an industry-leading, OpenTelemetry-native and AI-native tool stack, with cost and cardinality governance built in from the start for both telemetry and inference.
  • Operate as an executive partner. - Manage stakeholders across engineering and product, report progress and reliability/AI metrics to the CTO and executive team, and steward budget and headcount.


What success looks like:

  • First 90 days - organization mobilized and sponsor alignment confirmed; reliability baseline established; 2-4 lighthouse services selected; instrumentation underway with the first golden-signal dashboards and SLOs live; the first internal agents piloted.
  • By 12 months - the paved road is self-service and adopted by tier-1 teams; SLOs and error-budget policy are in effect; on-call load and alert noise are measurably down; the function is operating agent-first; and product engineers are using the SRE agents to instrument and meet SLOs with far less manual effort.
  • By 18 months - proactive and guarded agentic capabilities are in production; the program hits its targets - a significant reduction in Sev1/Sev2 incidents, mean-time-to-resolution cut substantially, full SLO coverage on tier-1 services, and a healthier on-call - and the company's SRE goals are being met at scale through agents that product engineers rely on.


Required Qualifications:
  • 15+ years in software engineering, including 7+ years in senior engineering leadership leading SRE, platform, infrastructure, or AI organizations at scale - including managing managers.
  • A proven track record of improving reliability and reducing production incidents across a large, complex, multi-team estate (thousands of engineers and/or services).
  • Demonstrated experience building and operating production AI and/or agentic systems at scale - with a working command of LLMs, agent frameworks and orchestration, retrieval, evaluation (evals), guardrails, and AI/agent observability.
  • Experience establishing AI governance and safety for production agents - least-privilege access, human-in-the-loop controls, blast-radius limits, and evaluation-gated autonomy.
  • Deep expertise in observability (OpenTelemetry; metrics, logs, traces), SLOs and error budgets, incident management, and cloud-native infrastructure (Kubernetes, infrastructure-as-code).
  • Experience running an internal platform as a product - paved roads, developer experience, and adoption measured by usage, not decree - and enabling other teams to build on it.
  • A demonstrated ability to drive org-wide change through influence in a large, matrixed organization rather than through mandate.
  • Excellent executive communication - able to translate reliability and AI strategy into business terms and present to C-level stakeholders.
  • Bachelor's degree in Computer Science or a related field, or equivalent practical experience.


Preferred Qualifications:
  • A track record of building internal AI/agent developer tooling that other engineering teams adopted at scale.
  • Experience in financial services or another regulated, high-availability, high-compliance environment.
  • Familiarity with a modern reliability and AI stack - e.g., OpenTelemetry, Prometheus/Grafana or a major observability SaaS, PagerDuty or Slack-native incident tooling, SLO tooling, and LLM/agent observability and evaluation platforms.
  • A track record establishing error-budget policy and a genuinely blameless postmortem culture.
  • Advanced degree in a relevant field.


Leadership competencies:

  • Technical credibility - deep enough to earn the trust of senior engineers and make hard architecture, tooling, and AI calls, while leading through others.
  • AI fluency & responsible innovation - moves fast on agentic capability while insisting on evaluation, guardrails, and safety.
  • Systems thinking - sees the whole sociotechnical system - tooling, incentives, culture, and cost.
  • Change leadership - moves a large organization through influence, evidence, and paved roads.
  • Talent magnet - attracts, grows, and retains exceptional engineering and AI leaders and specialists.
  • Bias for delivery - ships outcomes on aggressive timelines and is accountable for measurable results.
  • Psychological safety - builds a blameless, high-trust culture where failure is analyzed, not punished.


About FIS

Fidelity National Information Solutions is the nation’s most comprehensive source for real estate-related data, technology solutions.

FIS Careers

Join the vibrant team at FIS, a global leader in financial services technology. With a workforce of over 55,000 worldwide, FIS fosters a culture of innovation and leadership, offering unmatched opportunities for growth and professional development. Work You’ll Do At FIS, we empower the financial world with cutting-edge technology solutions, and we need your unique skills to help us redefine the future of finance. Whether you're looking for a position in software development, project management, or client services, FIS offers a dynamic range of job opportunities. Transform the Industry Be part of a team that is at the forefront of driving innovation in banking, payments, capital markets, and wealth management. FIS is where finance meets technology. By joining us, you will not only transform businesses but shape the way the world uses money. Innovative Work Environment FIS is committed to fostering a workplace where diverse perspectives and experiences are welcomed and respected. With diversity training and leadership programs, employees are encouraged to thrive both personally and professionally. Our commitment to diversity and inclusion is integral to our mission to empower the world with better financial technology. Career Growth and Development FIS believes in nurturing talent and leadership from within. With a plethora of training programs, workshops, and seminars, FIS is dedicated to your career advancement. From internships to leadership roles, your journey at FIS is filled with endless possibilities. Our robust networking events and professional development sessions ensure you have the tools to succeed in your career. Benefits and Culture Choosing a career at FIS means joining a family that values work-life balance and offers substantial benefits. Enjoy health, vision, and dental insurance, competitive retirement plans, and generous paid time off. Our culture promotes not just professional growth but personal well-being. Join Our Team Explore the numerous job opportunities at FIS where your skills can lead to meaningful outcomes. We are continuously hiring curious, creative, and motivated individuals who are ready to drive change in the fintech industry. Stay Connected Keep up to date with the latest in career opportunities and company news by joining our FIS Careers Network. Tailor your job search with personalized alerts and get insider tips that can help you craft the perfect resume and ace your interviews. Discover how your career can flourish by joining a team that’s dedicated to redefining the financial services landscape. Search FIS jobs today and see where your ambition can take you! [SEARCH FIS JOBS] Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here. [READ CAREERS BLOG] Job Alert Emails Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Explore the exciting and rewarding opportunities that await at FIS.
Learn more about FIS
Size
65,000 employees
Market Cap
$39.7 billion
Industry
Net Income
$158 million
Founded
1968
5 Year Trend
+9.5%
Revenue
$12.5 billion
NASDAQ

Similar Jobs

More Jobs at FIS

More Information Technology Jobs

Find similar VP, Site Reliability Engineering (SRE & Observability Platform) jobs: