We’re looking for a performance-focused ML Engineer to help speed up large-scale model training by optimizing our internal stack and compute infrastructure. You’ll work across the full training pipeline — from GPU kernels to system-level throughput — applying profiling, CUDA-level tuning, and distributed systems techniques. The goal is to reduce training time, boost iteration speed, and use compute more efficiently.
This is a key role in a growing team building deep technical expertise in ML training systems.
Responsibilities
Requirements
What we offer
By continuing you agree to our Terms & Privacy Policy.