In this role you will lead the operational assurance of our mission-critical platform supporting VaR, FRTB, stress testing, sensitivities, limits, and regulatory reporting. This system operates on a rigorous daily end-of-day (EOD) cycle relied upon by Risk, Finance, and Front Office teams. The platform is engineered for SLO-driven, observability-led operations, with a focus on minimizing manual toil and maximizing reliability.
Currently, the platform is undergoing a migration from SSIS to Databricks workflows. The support function is responsible for maintaining stability and performance on both legacy and new systems throughout the migration, including parallel-run reconciliation and eventual decommissioning of old components.
Key Responsibilities- Drive stability and incident command for EOD batch processes using Tidal, Airflow, Kubernetes, and Python orchestration, as well as intraday operations, data feeds, calculation engines, APIs, and reporting.
- Manage the technology plant including clusters, hosts, databases, storage, networks, service accounts, and cloud infrastructure, ensuring optimal capacity, patching, credential management, and secure access posture.
- Maintain rigorous hygiene of both infrastructure and applications, actively burning down backlogs and enforcing operational discipline.
- Lead L2/L3 support operations across regions, interfacing effectively with engineering teams.
- Oversee deployment, change management, audit, and service communication processes to meet regulatory requirements without unnecessary bureaucracy.
- Champion observability as a product, developing business-relevant SLOs, meaningful signals, and reducing alert noise.
- Define and enforce the boundary between development and support, ensuring recurring issues are resolved permanently rather than documented in runbooks.
Required Experience- 10+ years of production support experience in financial services, with at least 5 years leading L2/L3 support across multi-region teams.
- Hands-on expertise with Market Risk platforms (VaR, sensitivities, FRTB, stress testing, P&L, market data management, EOD batch) at scale where daily cycles are critical.
- Proficient in Python, capable of confidently reading, debugging, and extending orchestration and support tooling.
- Working knowledge of enterprise schedulers (Tidal, Airflow) and container orchestration (Kubernetes) within batch-driven environments.
- Experience supporting data pipelines on both SSIS and Databricks workflows, and comfortable managing a mixed technology estate during migration.
- Demonstrated ownership of the technology plant and a proactive approach to operational hygiene.
- Proven track record in incident command and advancing support operations toward SLO-driven, observability-led practices.
The expected base salary ranges from $200k-$250k. Salary offers are based on a wide range of factors including relevant skills, training, experience, education, and, where applicable, certifications and licenses obtained. Market and organizational factors are also considered. In addition to salary and a generous employee benefits package, successful candidates are eligible to receive a discretionary bonus.
#LI-Hybrid
Other requirementsMizuho has in place a hybrid working program, with varying opportunities for remote work depending on the nature of the role, needs of your department, as well as local laws and regulatory obligations. Roles in some of our departments have greater in-office requirements that will be communicated to you as part of the recruitment process
Mizuho Americas offers a competitive total rewards package.
#LI-MIZUHO