Role Summary:
As a Senior DevOps & MLOps Engineer, you will design, automate, and optimize the infrastructure that powers our agentic AI platform and Specialized Model. You will build scalable CI/CD systems, manage model deployment workflows, implement observability and monitoring frameworks, and ensure secure, efficient model delivery across cloud, hybrid, and on-prem environments. This role is highly collaborative – you’ll work closely with ML engineers, platform engineers, and product teams to streamline the journey from research to production. We value engineers with a learning mindset, deep ownership, strong systems thinking, and a passion for building reliable infrastructure that supports advanced AI applications at scale.
Key Responsibilities:
MLOps & Model Lifecycle Management: Build and maintain end-to-end machine learning pipelines, including data processing, automated training workflows, validation, and deployment.Implement model versioning systems and experiment tracking. Manage vector database integrations and embedding-based search pipelines. Build automated evaluation and monitoring workflows for model health and performance.
CI/CD & Automation: Design and manage CI/CD pipelines for services and models. Automate infrastructure provisioning using IaC tools. Ensure robust testing frameworks across services and pipelines.
Cloud Infrastructure & Scalability: Architect and manage cloud-native infrastructure across AWS/GCP/Azure. Implement scaling, high availability, and fault tolerance for workloads. Manage Kubernetes workloads, service mesh, and workload optimization.
Inference Optimization & Deployment: Deploy optimized ML workloads using ONNX/TensorRT or similar. Enable secure model deployments on edge or enterprise on-prem servers. Manage distributed training orchestration with frameworks like Ray or PyTorch Distributed.
Monitoring, Observability & Reliability: Build dashboards for monitoring infra, APIs, inference pipelines. Implement alerting systems and SLOs/SLIs. Improve reliability via logging, tracing, auto-recovery.
Security & Governance: Ensure secure handling of sensitive enterprise data. Implement IAM, encryption, network controls, and compliance mechanisms. Build workflows that adhere to enterprise governance and auditability.
Cross-Functional Collaboration: Partner with ML engineers and data teams to translate requirements. Improve engineering processes, documentation, and internal tooling.
Continuous Learning & Team Development: Stay updated on advancements in DevOps/MLOps and cloud systems. Mentor junior engineers in best practices. Contribute innovative solutions that enhance platform reliability and scale.
Required Skills & Experience:
Education & Experience: 8-12 years of experience in DevOps, MLOps, Cloud Engineering, or related fields. Bachelor’s/Master’s degree in Computer Science, Engineering, or equivalent practical experience.
Core Technical Skills: Proficiency in Python, Bash, and automation scripting. Deep experience with Docker and Kubernetes in production. Expertise with CI/CD tools such as GitHub Actions, Jenkins, GitLab CI, or ArgoCD. Strong understanding of cloud ecosystems (AWS/GCP/Azure). Experience with IaC tools (Terraform, CloudFormation).
MLOps Expertise: Experience building ML pipelines for training, validation, and deployment. Familiarity with MLflow, Kubeflow, SageMaker, or Vertex AI. Understanding of distributed training and GPU/TPU optimization. Experience with vector databases such as FAISS, Pinecone, or Milvus.
Observability & Reliability: Experience with Prometheus, Grafana, ELK, or OpenTelemetry. Strong debugging and incident-response skills.
Data & Distributed Systems: Experience handling large datasets and training jobs. Comfortable with data pipelines (ETL processes) and tools for distributed computing or parallel processing. Understanding of how to optimize training and inference on hardware (GPUs/TPUs) to improve throughput.
MLOps and Deployment: Experience deploying ML models into production. Familiar with containerization (Docker) and possibly orchestration (Kubernetes) for services, CI/CD pipelines for ML (e.g., automating model builds and deployments), and monitoring/observability tools to track model performance post-deployment. Knowledge of version control for models/data (MLflow or similar) is valuable.
Software Engineering Practices: Strong CS fundamentals (algorithms, data structures) and ability to write production-quality code. Experience with code reviews, unit/integration testing, and collaborative development (Git or similar version control workflows).
Problem Solving: Demonstrated ability to tackle ambiguous or open-ended problems in a structured way. Comfortable formulating experiments, evaluating results (using appropriate metrics), and iterating to improve model performance.
Communication & Teamwork: Excellent communication skills. Able to explain complex technical concepts to both technical colleagues and non-technical stakeholders. Collaboration is central – you should be able to work effectively in a team, share knowledge, and also take initiative independently when needed
By continuing you agree to our Terms & Privacy Policy.