JP Morgan Chase & Co.

Principal Software Engineer - LLM Optimization

JP Morgan Chase & Co.$162K — $195K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of practical experience in software engineering.
  • Hands-on experience with LLM inference systems like vLLM or TensorRT-LLM.
  • Strong understanding of GPU memory architecture and dynamics.
  • Expertise in quantization techniques and their trade-offs at scale.
  • Familiar with speculative decoding and production workload variables.
  • Proven benchmarking instincts with relevant tools or custom solutions.
  • Experience with cloud GPU infrastructure management.

Responsibilities

  • Establish systematic benchmarking and performance metrics for LLM workloads.
  • Design and execute quantization experiments for optimization measures.
  • Drive speculative decoding strategies across various model architectures.
  • Maintain a GPU efficiency scorecard for leadership insights.
  • Benchmark platform performance against industry standards.
  • Evaluate inference engine upgrades systematically before production.
  • Collaborate on optimization strategies with other engineering teams.

Benefits

  • Opportunity to shape technical direction and AI capabilities at a major financial firm.
  • High-visibility role with direct impact on the firm's AI initiatives.
  • Collaboration with senior engineering leadership and cross-functional teams.
  • Access to cutting-edge tools and technologies in AI and GPU management.
Full Job Description
JOB DESCRIPTION

As a Principal Software Engineer at JPMorganChase within the AI/ML Data Platform team, you will serve as the firm's deepest technical voice on LLM inference performance — owning optimization strategy, benchmarking rigor, and efficiency at scale. You will work directly with senior engineering leadership to shape how our platform evolves, ensuring every model we serve is fast, cost-efficient, and production-ready. This is a high-visibility individual contributor role where your technical decisions will have direct, measurable impact on the firm's AI capabilities

Job Responsibilities

 

  • Own systematic benchmarking and performance characterization across all production LLM workloads. Establish reproducible baselines, catch regressions early, and quantify the impact of every configuration change before it touches production

  • Design and execute quantization experiments — FP8, INT8/INT4 (GPTQ/AWQ), next-generation precision formats on current hardware — measuring accuracy delta, throughput improvement, memory reduction, and cost-per-token impact

  • Drive speculative decoding strategy across the model portfolio: draft model, n-gram, and multi-token prediction approaches. Own acceptance rate measurement and per-workload configuration recommendations

  • Build and maintain a GPU efficiency scorecard: utilization, memory headroom, cost per 1K tokens, and waste identified — giving leadership a data-driven view of platform efficiency at all times

  • Benchmark our platform against external providers and published industry numbers — know what good looks like, and close the gap

  • Lead inference engine upgrade evaluations: new scheduler architectures, async tensor parallelism, disaggregated prefill/decode, advanced speculative decoding — systematic validation before production promotion

  • Collaborate with the EKS and disaggregated serving teams on KV-cache optimization, prefix caching strategies, and multi-node serving architecture

  • Design and run GPU chaos engineering: induced failure scenarios, hardware diagnostic monitoring, detection and recovery measurement

  • Architect and govern agentic AI-enabled engineering workflows (using enterprise-authorized tools within the work environment) to improve delivery speed, code quality, and operational outcomes at scale (e.g., AI-driven PR review assistance, test generation/maintenance, release readiness checks, incident triage and root-cause acceleration), while defining guardrails for validation, security, resiliency, and reuse across teams. 

  • Apply knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation at scale. 

 

Required qualifications, capabilities, and skills

 

  • Formal training or certification on software engineering concepts and 7+ years applied experience 

  • Deep, hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D or equivalent production serving engines

  • Strong grasp of GPU memory architecture: KV cache sizing and dynamics, memory-bandwidth vs compute bottlenecks, the practical implications of quantization at inference time

  • Experience with quantization techniques and their real-world tradeoffs at scale

  • Familiarity with speculative decoding and the variables that drive acceptance rates in production workloads

  • Rigorous benchmarking instincts — GuideLLM, custom harnesses, or equivalent. Every claim has a number behind it

  • Comfort operating in cloud GPU infrastructure at scale (AWS; EKS, managed inference services)

  • Demonstrated awareness of the LLM inference competitive landscape, with a track record of applying industry benchmarks to drive platform improvements communicate technical trade-offs clearly to senior engineering and business stakeholders — this role presents upward regularly

  • Demonstrated experience designing and leading adoption of agentic AI-enabled development practices (using enterprise-authorized tools within the work environment) across teams, including setting standards for human-in-the-loop validation, auditability/traceability of changes, and secure handling of sensitive data. 

  • Strong understanding of responsible AI use and control expectations in engineering workflows, including security/resiliency implications, data sensitivity, and risk-based governance; ability to influence senior technical leaders on safe scaling patterns and reuse.

 

 Preferred qualifications, capabilities, and skills

 

  • Experience with disaggregated prefill/decode serving architectures, GPU hardware diagnostics (DCGM/NVML/XID event tracking), ML observability and production monitoring

 

About JP Morgan Chase & Co.

JP Morgan Chase & Co. stands at the forefront of the global financial services industry. They offer an expansive array of products and services to a diverse clientele, including individuals, corporations, governments, and institutions. Ever since the merger of J.P. Morgan & Co. and Chase Manhattan Corporation in 2000, this industry-leading entity has become renowned for its comprehensive portfolio encompassing consumer and community banking, corporate and investment banking, commercial banking, as well as asset and wealth management. Headquartered in the vibrant city of New York, JP Morgan Chase & Co. boasts a formidable presence across over 100 countries worldwide.

Unveiling Employment Opportunities at JP Morgan Chase & Co.

Vacancies and Hiring Initiatives

JP Morgan Chase & Co. is continuously on the lookout for talented individuals eager to contribute to its legacy of excellence. The company's recruitment efforts are geared towards identifying candidates with the right blend of skills and qualifications to drive forward its various business segments. Whether you are a seasoned professional or a recent graduate, JP Morgan Chase offers a plethora of job openings across multiple disciplines.

High-Demand Positions

Among the myriad of roles, certain positions stand out for their attractive compensation packages and career advancement prospects. Notably, high-paying jobs at JP Morgan Chase & Co. include Relationship Manager, Branch Manager, and Software Engineer. These roles are critical to the firm's operations and offer lucrative opportunities for those with the requisite expertise.

Navigating the Job Market at JP Morgan Chase & Co.

Leveraging Job Portals and Job Alerts

For job seekers aiming to tap into the opportunities at JP Morgan Chase, staying updated through job portals and subscribing to job alerts is crucial. These tools can provide timely information about job openings, job fairs, and recruitment events, enabling candidates to apply promptly and prepare adequately for interviews.

Preparing Your Job Application

Your job application, comprising your resume and cover letter, is your ticket to securing an interview at JP Morgan Chase. Highlight your qualifications, skills, and experiences that align with the job listing, ensuring you stand out in the competitive job market.

Acing the Interview

Preparation is key to succeeding in your interview with JP Morgan Chase. Familiarize yourself with the company's business segments, values, and recent achievements. Demonstrating how your background and aspirations match the company's goals can significantly increase your chances of employment. A World of Job Opportunites in the Financial Services Industry JP Morgan Chase & Co. offers a world of job opportunities for those seeking to make their mark in the financial services industry. With competitive salaries, comprehensive benefits, and endless possibilities for growth, positions at JP Morgan Chase are highly coveted. By staying informed through job sites, tailoring your applications, and preparing thoroughly for interviews, you can enhance your prospects of joining the esteemed ranks of JP Morgan Chase employees. Explore the job board, seize the job opportunities, and embark on a rewarding career journey with one of the world's leading financial institutions.
Learn more about JP Morgan Chase & Co.
Size
661 employees
Market Cap
$384.5 billion
Industry
Net Income
$29.1 billion
Founded
1823
5 Year Trend
+0.7%
Revenue
$261.5 million
NASDAQ

Similar Jobs

More Jobs at JP Morgan Chase & Co.

More Information Technology Jobs

Find similar Principal Software Engineer - LLM Optimization jobs: