Ernst & Young

Manager - AI Inference Engineer

Ernst & Young$125K — $230K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's Degree in a relevant field
  • Strong experience in AI inference system design and operation
  • Deep familiarity with multi-GPU serving
  • Hands-on experience with Kubernetes and cloud-native technologies
  • Strong understanding of production AI workload scaling and reliability
  • Experience integrating AI platforms into enterprise environments

Responsibilities

  • Design and deploy secure private AI inference architectures
  • Translate workload requirements into technical architecture decisions
  • Define reference architectures for high-throughput, low-latency inference
  • Design the model-serving stack including runtime and API layers
  • Evaluate and recommend production use of inference technologies
  • Build and run benchmark suites for evaluating enterprise workloads
  • Support integration with client systems and validate security requirements

Benefits

  • Comprehensive compensation and benefits package based on performance
  • Flexible work environment and hybrid model
  • Flexible vacation policy to support personal circumstances
  • Time off for designated paid holidays and personal leave
  • Opportunities for future-focused skill development and world-class experiences
Full Job Description
Location: Anywhere in Country

The opportunity

We are seeking an experienced AI Inference Engineer to help design, deploy, and operate private AI inference infrastructure at rack scale. This role is focused on the engineering of a production-grade inference platform: architecture, serving stack design, benchmarking, integration, reliability, and operational readiness.

The ideal candidate has deep hands-on experience building AI inference systems and the platform layers that support them, including multi-GPU inference, Kubernetes-based deployment, autoscaling, observability, routing, access control, and production support.

Your key responsibilities

Private AI inference architecture
  • Design and deploy secure private inference architectures for enterprise AI inference.
  • Translate workload requirements into technical architecture decisions across compute, caching, interconnect, storage & networking.
  • Define reference architectures that support high-throughput, low-latency, enterprise inference at scale.

Inference platform engineering
  • Design the model-serving stack, including inference runtime, orchestration, model registry, artifact management, API layer, routing, observability, deployment and fallback.
  • Evaluate and recommend inference technologies and platform components for production use.
  • Configure and harden the target environment for secure, reliable inference workloads.

Benchmarking and performance evaluation
  • Build and run benchmark suites for enterprise workloads.
  • Measure latency, throughput, concurrency, utilization, reliability, and cost efficiency.
  • Tune serving configurations and system parameters to improve production performance.
  • Produce evaluation results and recommendations based on objective testing.

Enterprise integration
  • Support integration of inference platforms with client IT systems.
  • Work with infrastructure and security stakeholders to validate technical requirements and remediate issues.

Model onboarding and operationalization
  • Onboard models into the target inference platform.
  • Configure serving endpoints, routing behavior, monitoring, and operational controls.
  • Support deployment readiness, runbooks, and operational handoff for production use.


Skills and attributes for success
  • Ability to turn business and workload needs into concrete infrastructure and software design decisions.
  • Clear technical communication and strong architecture documentation skills.


To qualify you must have
  • Bachelor's Degree in a relevant field
  • Strong experience designing and operating AI inference systems in production.
  • Deep familiarity with multi-GPU serving and high-performance deployment patterns.
  • Hands-on experience with Kubernetes and cloud-native platform technologies.
  • Strong understanding of scaling, routing, autoscaling, observability, and reliability for production AI workloads.
  • Experience integrating AI platforms into enterprise security and access environments.


Ideally, you'll also have
  • Experience with private, sovereign, or enterprise AI infrastructure.
  • Experience benchmarking large-model inference on GPU clusters.
  • Familiarity with service mesh, policy enforcement, secrets management, and RBAC.
  • Experience working across infrastructure, security, and application teams in a regulated environment.
  • Experience supporting pilot-to-production AI platform rollouts.


What we offer you
At EY, we'll develop you with future-focused skills and equip you with world-class experiences. We'll empower you in a flexible environment, and fuel you and your extraordinary talents in a diverse and inclusive culture of globally connected teams. Learn more.
  • We offer a comprehensive compensation and benefits package where you'll be rewarded based on your performance and recognized for the value you bring to the business. The base salary range for this job in all geographic locations in the US is $125,500 to $230,200. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $150,700 to $261,600. Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography. In addition, our Total Rewards package includes medical and dental coverage, pension and 401(k) plans, and a wide range of paid time off options.
  • Join us in our team-led and leader-enabled hybrid model. Our expectation is for most people in external, client serving roles to work together in person 40-60% of the time over the course of an engagement, project or year.
  • Under our flexible vacation policy, you'll decide how much vacation time you need based on your own personal circumstances. You'll also be granted time off for designated EY Paid Holidays, Winter/Summer breaks, Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.


Are you ready to shape your future with confidence? Apply today.
EY accepts applications for this position on an on-going basis.

For those living in California, please click here for additional information.

EY focuses on high-ethical standards and integrity among its employees and expects all candidates to demonstrate these qualities.

About Ernst & Young

Ernst & Young (EY) is a multinational professional services firm that provides audit, tax, consulting, and advisory services to clients in a wide range of industries. The firm was founded in 1989 through the merger of Ernst & Whinney and Arthur Young & Co., and has since grown to become one of the largest professional services firms in the world. EY is committed to building a better working world by helping its clients solve their toughest challenges, and by creating a positive impact on the communities it serves.
Learn more about Ernst & Young
Size
300,000 employees
Industry
Founded
1989

Similar Jobs

More Jobs at Ernst & Young

More Enterprise Technology Jobs

Find similar Manager - AI Inference Engineer jobs: