Job Family Group:
IT&S Group
Job Description:
Role Overview
bpx energyis building an enterprise AI capability that can scale safely and deliver real operational value. The Senior AI/ML Platform Engineer will help build andoperatethe technical foundationrequiredto move AI/ML capabilities from project-based implementations into governed, observable, production-grade enterprise capabilities.
This is a hands-on platform engineering role focused on the systems, patterns, environments, controls, and automationrequiredforproductionAI/ML delivery. The role will work acrossPalantir, Snowflake,Databricks, AWS, and related AI/ML services to create the paved roads that allow teams to be versatile and quick-moving.
This role will not focus on building one-off AI use cases. It is focused on making AI/ML engineering repeatable, reliable, secure, and scalable across the enterprise.
WhatYoullDo
Build,operate, and evolve AI/ML platform capabilities across Palantir, Databricks, AWS,MLflow, model registries, model serving, feature management, vector stores, and related services.
- Create reusable platform patterns for model development, deployment, serving, monitoring, access controls, and production support.
- Implement CI/CD, infrastructure automation, environment management,secretsmanagement, access controls, and deployment templates for AI/ML workloads.
- Partner with security, infrastructure, data, and enterprise architecture teams to ensure AI/ML platforms are secure, observable, auditable, and operationally reliable.
- Support batch, real-time, streaming, and API-based model deployment patterns.
- Establish standard engineering patterns for experiments, notebooks, jobs, pipelines, model serving, and production promotion.
- Help define platform usage standards, tiered access models, cost controls, observability requirements, and operational support patterns.
- Ensure AI/ML workloads are designed for reliability, scalability, performance, maintainability, and governance.
- Support future federated AI/ML engineering by creating reusable templates, reference architectures, and enablement materials for domain teams.
MinimumRequirements
Bachelors degree in engineering, computer science, information systems, or related field, or equivalent work experience.
- Provenexperiencebuilding,operating, or enabling production AI/ML engineering platforms in a cloud environment.
- Hands-on experience with at least one modern AI/ML platform such as Databricks, AWS SageMaker,MLflow, Azure ML, Vertex AI,orequivalent.
- Practical experience with CI/CD, infrastructure automation, environment management,secretsmanagement, access controls, and production deployment patterns.
- Experience supporting model development and deployment workflows beyond experimentation or notebooks.
- Strong understanding of cloud-native architecture, APIs, containers, compute patterns, storage patterns, and runtime observability.
- Ability to build reusable engineering patterns, templates, reference architectures, and platform paved roads.
- Experience partnering with data engineering, security, infrastructure, and architecture teams to move AI/ML workloads into governed production environments.
- Proven track record to troubleshootplatform, deployment, performance, integration, or reliability issues in sophisticated technical environments.
Strongly Preferred
- Databricks platform engineering experience, including workspaces, clusters/serverless, Unity Catalog,MLflow, model serving, jobs/workflows, permissions, and cost controls.
- AWS experience with IAM, networking, security groups, S3, Lambda, ECS/EKS, API Gateway, Bedrock, SageMaker, or related services.
- Experience supporting regulated, safety-sensitive, industrial, energy, financial, healthcare, or other high-consequence operating environments.
- Experience with platform cost management and workload optimization.
- Experience creating reusable platform enablement materials for engineers, data scientists, or domain technical teams.
Additional Role Scope Information
This is not a traditional software engineering, application development, BI, or data engineering role. It is also not a notebook-only experimentation role.
This role is not a fit for candidates whose experience is primarily:
- Traditional application/software engineering without hands-on AI/ML platform,MLOps, orModelOpsexperience.
- Generic cloud or DevOps engineering withoutproductionAI/ML deployment or platform experience.
- Data science experimentation without responsibility for production deployment patterns.
- Data pipeline engineering without exposure to model development, model serving, or AI/ML lifecycle operations.
- Single-use-case delivery without experience creating reusable platform capabilities.
Adjacent backgrounds are welcome when the candidate candemonstratedirect experience helping AI/ML workloads move into governed, observable, production-grade environments.
Salary and BenefitsWe offer a reward and wellbeing package to enable your work to fit with your life. These can include, but not limited to, access to health, vision and dental insurance, flexible working schedule, paid time off policy, discretionary annual bonus program, long-term incentive program, and a generous 401K matching program.How much do we pay (Base)? $135,000 - $175,000
*Note that the pay range listed for this position is a good faith and reasonable estimate of the range of possible base compensation at the time of posting.
Travel Requirement:
Negligible travel should be expected with this role
Relocation Assistance:
Relocation may be negotiable for this role
Remote Type:
This position is a hybrid of office/remote working
Skills:
Cloud Platforms, Cloud Platforms, Collaboration, Communication, Configuration management and release, Continuous deployment and release, Creating a high performing team, Database Design, Digital Project Management, Documentation and knowledge sharing, Emerging technology monitoring, Facilitation, Information Security, Mentoring, Metrics definition and instrumentation, NoSql data modelling, Problem Solving, Relational Data Modeling, Risk Management, Scripting, Secure development, Service operations and resiliency, Software Design and Development, Solution Architecture, Source control and code management {+ 5 more}
.