Ampere Computing

AI Accelerator, Software Engineer- Graph Optimization/Compilers

Ampere Computing$159K — $239K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Mathematics, or related field with 3-5 years of relevant experience.
  • Strong foundation in algorithms, data structures, and computational complexity.
  • Proficiency in Python and C/C++, with practical implementation experience.
  • Strong reasoning ability regarding execution dependencies and numerical correctness.
  • Experience with profiling, benchmarking, or hardware analysis is advantageous.
  • Familiarity with deep learning concepts and neural network architectures.
  • Experience with GPU/NPU programming or accelerator architectures is a plus.

Responsibilities

  • Optimize deep learning computational graphs for performance and efficiency on Ampere AI accelerators.
  • Enhance models and frameworks like PyTorch and Llama.cpp based on performance metrics.
  • Develop advanced graph-level optimizations for memory and compute efficiency.
  • Analyze performance across software frameworks and hardware components.
  • Build profiling and benchmarking infrastructure to ensure high performance standards.
  • Collaborate with cross-functional teams on hardware/software co-design issues.
  • Participate in architecture reviews, documentation, and best practices.

Benefits

  • Premium medical, dental, and vision insurance, including income protection and a 401K plan.
  • Unlimited flextime and over 10 paid holidays for work-life balance.
  • Access to healthy snacks and beverages in the workplace.
Full Job Description
Description

About the Role:

As a Software Engineer on Ampere's AI Accelerator team, you will optimize deep learning computational graphs to maximize the performance, efficiency, and scalability of Ampere's AI accelerator hardware. You will work across the software stack, from model frameworks and inference-serving systems to graph optimization, compiler infrastructure, runtimes, and compute kernels.

What You'll Achieve:
  • Optimize computational graphs for performance, throughput, latency, memory efficiency, and power efficiency on Ampere AI accelerators
  • Enable and optimize models, frameworks, and inference platforms, including PyTorch, Llama.cpp, vLLM, and SGLang
  • Develop graph-level optimizations such as operator fusion, pattern matching, redundancy elimination, constant folding, layout optimization, memory planning, quantization, and accelerator offload
  • Optimize transformer and LLM workloads, including dynamic shapes, attention mechanisms, KV-cache management, and mixed-precision execution
  • Analyze end-to-end performance across frameworks, compilers, runtimes, kernels, and hardware
  • Build profiling, benchmarking, validation, and performance-regression infrastructure
  • Identify bottlenecks using traces, compiler diagnostics, microbenchmarks, and hardware performance data
  • Collaborate with compiler, runtime, kernel, architecture, hardware, and applications teams on hardware/software co-design
  • Contribute to architecture, design reviews, code reviews, documentation, and engineering best practices

About You:
  • Bachelor's degree in Computer Science, Computer Engineering, Mathematics, or a related technical field & 5 years of relevant experience; or a Master's degree with & 3 years of relevant experience
  • Strong foundations in algorithms, data structures, graph algorithms, computational complexity, and systems programming
  • Proficiency in Python and C/C++, demonstrated through internships, research, coursework, open-source contributions, or personal projects
  • Strong ability to reason about execution dependencies, memory movement, numerical correctness, and hardware execution behavior
  • Experience diagnosing performance issues through profiling, benchmarking, tracing, or hardware-level analysis is a plus
  • Familiarity with deep learning concepts, neural-network architectures, tensor operations, numerical precision, quantization, and memory layout
  • Experience with CUDA, ROCm, OpenCL, SYCL, Triton, GPU programming, NPU programming, or other accelerator architectures is a plus
  • Familiarity with transformer models, LLM inference, attention mechanisms, KV-cache optimization, speculative decoding, mixed-precision execution, or sparsity is a plus
  • Demonstrated exceptional problem-solving ability-IOI medal, ACM ICPC medal, Codeforces Grandmaster, USACO Platinum, or equivalent achievement in research or production engineering is a strong plus
  • Strong analytical and debugging skills, with the ability to investigate ambiguous technical problems and deliver robust solutions
  • Fast learner who can quickly understand new architectures, frameworks, compilers, and workloads
  • Experience using AI-assisted development tools to accelerate implementation, testing, debugging, and code review while maintaining technical ownership and code quality

What We'll Offer:

At Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $159,000 and $239,000. Our benefits include health, wellness, and financial programs that support employees through every stage of life.

Benefit highlights include:
  • Premium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.
  • Unlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.
  • A variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.

And there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.

#LI-Hybrid #LI-DR

#LI-Hybrid

About Ampere Computing

Ampere Computing is a semiconductor company that designs and manufactures high-performance processors for cloud and edge computing. The company's processors are based on the Arm architecture and are optimized for power efficiency and performance. Ampere Computing was founded in 2017 by former Intel president Renee James and is headquartered in Santa Clara, California.
Learn more about Ampere Computing
Size
200 employees
Industry
Founded
2017

Similar Jobs

More Jobs at Ampere Computing

More Information Technology Jobs

Find similar AI Accelerator, Software Engineer- Graph Optimization/Compilers jobs: