We are hiring for a Senior DevOps / Site Reliability Engineer AI Platform. The role focuses on building, operating, monitoring, and scaling an enterprise AI platform across Development, QA, and Production environments.
A strong fit will have advanced Azure and Kubernetes experience along with strong expertise in observability, dashboards, monitoring, alerting, and automated scaling.

What you bring

  • Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure, or platform engineering.
  • Advanced hands-on experience with Microsoft Azure.
  • Strong experience deploying and operating Kubernetes environments.
  • Strong knowledge of Docker and containerization technologies.
  • Experience supporting containerized applications across Development, QA, and Production environments.
  • Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights, or comparable tools.
  • Strong experience with monitoring, observability, logging, alerting, and operational health checks.
  • Experience with application load, infrastructure capacity, performance, and automated scaling.
  • Experience with CI/CD pipelines and automated application deployment.
  • Experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates.
  • Strong troubleshooting skills across applications, containers, infrastructure, networking, and cloud services.

What you'll do

  • Design, deploy, configure, and maintain infrastructure within Microsoft Azure.
  • Deploy and operate containerized applications using Kubernetes.
  • Monitor container and cluster health, resource consumption, capacity, and performance.
  • Configure scaling policies and develop intelligent scaling approaches based on workload and resource utilization.
  • Design and build operational dashboards covering platform health, performance, capacity, errors, latency, and container health.
  • Implement monitoring and alerting across infrastructure, applications, containers, integrations, and AI platform services.
  • Establish actionable alerts, health checks, anomaly detection, and automated remediation where appropriate.
  • Support production platform stability, availability, and operational readiness.
  • Investigate platform, deployment, infrastructure, monitoring, and performance issues.
  • Participate in root-cause analysis and implement preventative improvements.
  • Create operational runbooks and troubleshooting guidance.

Nice to have

  • Experience supporting AI, machine learning, data, or high-compute platforms.
  • Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues, or model performance.
  • Experience implementing automated remediation, predictive monitoring, or AI-assisted platform operations.
  • Familiarity with AWS services and cloud operations.
  • Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus, or similar observability technologies.
  • Experience defining service-level indicators, service-level objectives, and reliability standards.
  • Experience with security, identity, secrets management, and cloud governance within Azure.

More from J M Group Inc
J M Group Inc 12 hours ago
J M Group Inc 12 hours ago
J M Group Inc 12 hours ago

Cloud DevOps Engineer

Apply Now
Back to search page