Key Responsibilities Design, deploy, maintain, and troubleshoot enterprise-scale Linux and Windows server infrastructure. Support hyperscale GPU and compute environments used for AI training and inference workloads. Troubleshoot NVIDIA and AMD GPU platforms, including GPU memory errors, driver