Job SummaryThe AI Platform Engineering Specialist will help build and scale a firmwide AI Development Platform while driving adoption of AI capabilities across the enterprise. This role focuses on delivering secure, scalable, production-grade platform solutions across cloud infrastructure, Kubernetes, APIs, data engineering, and Generative AI technologies.
The ideal candidate will have strong hands-on experience with Python, Kubernetes, Azure, AWS, API-based development, infrastructure as code, and CI/CD. The role requires close collaboration with cloud platform, security, network, and engineering teams to build and operate enterprise AI services from proof of concept through production.
Key Responsibilities- Design, build, deploy, and operate AI Gateway solutions across Azure and AWS, taking solutions from proof of concept through production.
- Develop and enhance Python services using FastAPI and Flask for inference, onboarding, and administrative APIs.
- Integrate new model providers and model families, including Azure AI Foundry, Azure OpenAI, and AWS Bedrock.
- Implement request signing, streaming responses, failover, and quota-handling capabilities.
- Implement secure cloud-native authentication and secrets management using Entra ID, Managed Identity, workload federation, AWS IAM roles, and STS.
- Build and maintain entitlement and authorization data layers using SQL Server and PostgreSQL, including schema changes, migrations, and data-quality controls.
- Develop platform governance capabilities, including rate limiting, token accounting, content guardrails, audit logging, and chargeback reporting.
- Deploy and operate services across Kubernetes environments, including on-premises, AKS, and EKS, using Helm, GitOps, and Terraform.
- Maintain and enhance CI/CD pipelines using Jenkins and GitHub Actions.
- Build comprehensive monitoring and observability capabilities using Prometheus, Grafana, Loki, and Snowflake.
- Partner with cloud, network, and security teams on connectivity, egress policies, network controls, architecture reviews, and supporting documentation.
- Participate in on-call production support, investigate incidents, and implement fixes and platform hardening improvements.
- Develop tests and technical documentation as part of the delivery process.
- Review peer code and contribute to overall engineering quality and best practices.
Required Qualifications- Strong production-level Python development experience, including FastAPI or Flask.
- Strong software testing practices and experience developing reliable production services.
- Hands-on experience deploying, configuring, and troubleshooting workloads in Kubernetes environments.
- Practical knowledge of OIDC and OAuth 2.0, including token validation, JWKS, client-credentials flows, claims, and audience handling.
Hands-on Microsoft Azure experience with at least three of the following:- AKS
- Entra ID, including app registrations, service principals, Managed Identity, or Workload Identity
- Azure OpenAI or Azure AI Foundry
- Key Vault
- Azure Database for PostgreSQL
- Azure Cache for Redis
- Azure Monitor
- Hands-on AWS experience with at least three of the following:
- IAM and STS/assume-role
- SigV4 request signing
- AWS Bedrock
- EKS
- VPC endpoints and private networking
- Secrets Manager
- CloudWatch
- Experience with infrastructure as code using Terraform, Bicep, or CDK.
- Experience with CI/CD using Jenkins or GitHub Actions.
- Strong SQL and relational data modeling experience, including schema migrations.
- Strong written and verbal communication skills.
- Ability to collaborate directly with security, network, cloud, and platform engineering teams.
Preferred Qualifications- Experience building or operating API gateways, reverse proxies, or multi-tenant platforms.
- Experience with LLM platform engineering concepts, including streaming, server-sent events, token accounting, prompt and response guardrails, and model evaluation.
- Experience with Kafka and Snowflake for audit and consumption data pipelines.
- Advanced observability experience with Prometheus, PromQL, Grafana, Loki, or OpenTelemetry.
- Experience with Redis or Valkey, including counters, TTLs, and distributed rate-limiting concepts.
- Experience delivering technology solutions within a regulated enterprise environment involving corporate proxies, private networking, and strict change-control processes.
- Knowledge of AI/ML technologies and hands-on experience implementing Generative AI solutions.