Principal engineer – GPU Memory Systems & Scale-Up/Out Systems
We are seeking a highly motivated Principal engineer to advance the state-of-the-art in GPU memory systems and scale-up/out computing architectures . This role involves designing and optimizing memory hierarchies, interconnects, and distributed systems to improve performance, efficiency, and scalability for next-generation GPU-accelerated workloads (e.g., AI/ML, HPC, and large-scale data analytics).
The ideal candidate will have deep expertise in computer architecture, memory systems, parallel computing, and distributed systems , with a strong publication record in top-tier conferences (e.g., ISCA, MICRO, ASPLOS, HPCA, SC, NeurIPS).
Key Responsibilities
Research and develop novel GPU memory architectures (e.g., cache hierarchies, near-memory computing, disaggregated memory, link/switch optimizations).
Design and evaluate scale-up and scale-out systems for GPU clusters, focusing on interconnect topologies, communication protocols, and load balancing .
Optimize memory bandwidth, latency, and capacity utilization for large-scale GPU workloads.
Collaborate with hardware and software teams to prototype new architectures (e.g., using simulators, FPGA emulation, or real hardware).
Publish cutting-edge research in top conferences/journals and contribute to patents.
Engage with industry and academic partners to drive innovation in GPU-accelerated computing.
Required Qualifications
PhD in Computer Science, Electrical Engineering, or related field (or equivalent research experience).
Strong background in computer architecture, memory systems, and parallel/distributed computing .
Hands-on experience with GPU architectures, CUDA, RDMA, or high-performance interconnects (CXL) .
Proficiency in performance modeling and simulation (e.g., GPGPU-Sim, SST, Gem5, NS3).
Strong programming skills in C/C++, Python, or SystemVerilog/HDL for prototyping.
Track record of publications in ISCA, MICRO, ASPLOS, HPCA, SC, or related venues .
Preferred Qualifications
Experience with disaggregated memory, cache coherence protocols, or memory-centric computing .
Knowledge of AI/ML workloads and their memory/system bottlenecks .
Contributions to open-source hardware/software projects (e.g., LLVM, PyTorch, OpenMPI).
Familiarity with datacenter-scale systems (e.g., distributed training, high-performance storage).
Why Join Us?
Work on cutting-edge research with real-world impact in AI, HPC, and cloud computing.
Collaborate with world-class researchers and engineers .
Competitive compensation, equity, and publication/patent incentives.
This job description balances technical depth with broad applicability, attracting candidates with expertise in GPU memory systems and large-scale computing . Let me know if you'd like any refinements!
#J-18808-Ljbffr