JOB DESCRIPTION
Join a team that builds secure, scalable data pipelines and platforms for collection, access and analytics.
As a Lead data Engineer, within the Corporate Sector you will lead architecture and engineering for scalable backup, recovery, immutable-data protection, and recovery-assurance services across data platforms and storage tiers and deliver data collection, storage, access, and analytics platform solutions in a secure, stable, scalable way, with clear SLOs/SLAs and operational readiness gates.
Job Responsibilities
Observability (Dashboards: Grafana, Dynatrace, Tableau):
- Design, build, and maintain role-based observability dashboards for data platforms and backup/recovery services using Grafana, Dynatrace, and Tableau, providing real-time and historical visibility into platform health, performance, and risk posture.
- Define and govern dashboard standards and operating model (golden dashboards, drill-down paths, naming/tagging conventions, ownership, and lifecycle management) to drive consistent adoption across Engineering, SRE, Operations, and Risk.
- Ensure end-to-end instrumentation and telemetry quality so dashboards surface actionable signals (availability, latency, throughput, error rates, saturation) and platform KPIs (job success/failure, lag/backlog, data freshness, restore success, retention coverage).
- Establish and publish SLIs/SLOs and operational health KPIs; translate these into Grafana/Dynatrace visualizations and Tableau reporting for executive and governance stakeholders.
Data Architecture, Modeling, and Governance:
- Generate and govern data models using firmwide tooling; apply linear algebra, statistical, and geometric algorithms where relevant for modeling and optimization.
- Own and continuously improve the strategy for database backup, recovery, archiving, retention, and restore testing across relational and NoSQL estates.
Platform Engineering & Automation:
- Drive an automation-first, API-led, self-service approach that reduces operational toil and improves resilience and customer experience.
- Build cloud-native capabilities using AWS services, infrastructure as code, CI/CD, and modern software engineering practices.
Reliability Engineering (SRE) & Operational Excellence:
- Embed SRE principles by defining and tracking reliability metrics (SLIs/SLOs), implementing observability standards, leading incident learning, and maintaining runbooks and recovery playbooks.
- Participate in an on-call and incident leadership rotation, acting as an escalation point for complex platform and data reliability issues.
Security, Controls, and Recovery Assurance:
- Assess and report on access control effectiveness and data asset security posture; partner with security/risk to remediate gaps.
- Design preventive/detective controls, policy guardrails, automated validation, and audit-ready evidence for platform controls and recovery readiness.
- Deliver telemetry, reporting, and operational intelligence for backup health, compliance, and recovery assurance (coverage, success rates, RPO/RTO attainment).
Technical Leadership & Delivery Management:
- Set engineering standards, perform high-quality code reviews, mentor engineers, and influence stakeholders across product, architecture, security, and SRE.
- Coordinate cross-team delivery with explicit dependencies, milestones, and measurable outcomes; manage technical debt and balance reliability/security work alongside feature delivery.
Core Engineering Scope (Hands-on):
Required Qualifications, Capabilities, and Skills
- Typically8+ yearsof applied engineering experience (software engineering or related discipline), including leadingproduction-critical systems; experience with bothrelational and NoSQLdatabases; Strong capability in: observability & operations (monitoring, lineage, incident response/runbooks), security/privacy/governance (least privilege, encryption concepts, auditability), system design & tradeoff analysis (scalability, latency, reliability, cost), and technical leadership (standards, mentoring, stakeholder alignment).
- Proficient across the data lifecycle (ingestion, modeling, storage, serving, governance, operations) with demonstrated experience buildingdistributed systems,platform services,APIs, andautomation frameworks.
- Strong knowledge ofenterprise data protection, backup/recovery, cyber resilience, immutable storage, retention, and recovery testing.
- AdvancedAWSexperience spanning backup/data protection services,IAM/security, networking fundamentals, observability, and infrastructure automation.
- Hands-on experience with observability/monitoring platforms such asDynatraceandGrafana and proficiency inPythonand at least one additional language (e.g.,Java, Go, C#).
- Experience with CI/CD platforms (e.g.,Jenkins, GitLab, GitHub Actions) and implementing database backup/recovery/archiving strategies with measurableRPO/RTOtargets.
- Proficient knowledge oflinear algebra, statistical, and geometric algorithms; ability to translate security, control, and regulatory requirements into engineered solutions and operational processes.
Preferred Qualifications, Capabilities, and Skills
- Experience withKubernetes/OpenShiftand containerized platform operations.
- Experience with database platforms andcyber-recovery / isolated recoverysolutions.
- Experience building operational analytics (health/compliance dashboards, evidence automation) forregulated financial servicesenvironments.