Create Alert
Email me similar jobs

Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

We are building AWS Neuron, an SDK that accelerates deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. As part of the Acceleration Kernel Library team, you will craft high‑performance kernels for ML functions, optimizing performance across hardware, compiler, runtime, and framework layers.

Key Responsibilities
  • Design and implement high‑performance compute kernels for ML operations on Neuron hardware
  • Analyze and optimize kernel‑level performance across multiple generations of Neuron accelerators
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
  • Implement compiler optimizations such as fusion, sharding, tiling, and scheduling
  • Work directly with customers to enable and optimize their ML models on AWS accelerators
  • Collaborate across teams to develop innovative kernel optimization techniques

Basic Qualifications
  • 5+ years of non‑internship professional software development experience
  • 5+ years of programming in at least one software programming language
  • 5+ years of leading design or architecture of new and existing systems (design patterns, reliability, scaling)
  • 5+ years of full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations
  • Experience as a mentor, tech lead, or leading an engineering team

Preferred Qualifications
  • Bachelor’s degree in computer science or equivalent
  • 6+ years of full software development experience
  • Expertise in accelerator architectures for ML or HPC (GPUs, CPUs, FPGAs, or custom)
  • Experience with GPU kernel optimization and GPGPU computing (CUDA, NKI, Triton, OpenCL, SYCL, or ROCm)
  • Demonstrated experience with NVIDIA PTX and/or AMD GPU ISA
  • Experience developing high‑performance libraries for HPC applications
  • Proficiency in low‑level performance optimization for GPUs
  • Experience with LLVM/MLIR backend development for GPUs
  • Knowledge of ML frameworks (PyTorch, TensorFlow) and their GPU backends
  • Experience with parallel programming and optimization techniques
  • Understanding of GPU memory hierarchies and optimization strategies

Amazon is an equal‑opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Base salary range: $193,300.00 – $261,500.00 annually (location‑specific).

#J-18808-Ljbffr
Similar jobs

Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Apply Now
Back to search page