Role Summary: As a Senior DevOps & MLOps Engineer at Avataar, you will design, automate, and optimize the infrastructure that powers our agentic AI platform and Specialized Model Avataars (SMAs). You will build scalable CI/CD systems, manage model deployment workflows, implement observability and monitoring frameworks, and ensure secure, efficient model delivery across cloud, hybrid, and on-prem environments. This role is highly collaborative – you’ll work closely with ML engineers, platform engineers, and product teams to streamline the journey from research to production. We value engineers with a learning mindset, deep ownership, strong systems thinking, and a passion for building reliable infrastructure that supports advanced AI applications at scale.
Key Responsibilities:
- MLOps & Model Lifecycle Management: Build and maintain end-to-end machine learning pipelines, including data processing, automated training workflows, validation, and deployment.Implement model versioning systems and experiment tracking. Manage vector database integrations and embedding-based search pipelines. Build automated evaluation and monitoring workflows for model health and performance.
- CI/CD & Automation: Design and manage CI/CD pipelines for services and models. Automate infrastructure provisioning using IaC tools. Ensure robust testing frameworks across services and pipelines.
- Cloud Infrastructure & Scalability: Architect and manage cloud-native infrastructure across AWS/GCP/Azure. Implement scaling, high availability, and fault tolerance for workloads. Manage Kubernetes workloads, service mesh, and workload optimization.
- Inference Optimization & Deployment: Deploy optimized ML workloads using ONNX/TensorRT or similar. Enable secure model deployments on edge or enterprise on-prem servers. Manage distributed training orchestration with frameworks like Ray or PyTorch Distributed.
- Monitoring, Observability & Reliability: Build dashboards for monitoring infra, APIs, inference pipelines. Implement alerting systems and SLOs/SLIs. Improve reliability via logging, tracing, auto-recovery.
- Security & Governance: Ensure secure handling of sensitive enterprise data. Implement IAM, encryption, network controls, and compliance mechanisms. Build workflows that adhere to enterprise governance and auditability.
- Cross-Functional Collaboration: Partner with ML engineers and data teams to translate requirements. Improve engineering processes, documentation, and internal tooling.
- Continuous Learning & Team Development: Stay updated on advancements in DevOps/MLOps and cloud systems. Mentor junior engineers in best practices. Contribute innovative solutions that enhance platform reliability and scale.
Required Skills & Experience: - Education & Experience: 8-12 years of experience in DevOps, MLOps, Cloud Engineering, or related fields. Bachelor’s/Master’s degree in Computer Science, Engineering, or equivalent practical experience.
- Core Technical Skills: Proficiency in Python, Bash, and automation scripting. Deep experience with Docker and Kubernetes in production. Expertise with CI/CD tools such as GitHub Actions, Jenkins, GitLab CI, or ArgoCD. Strong understanding of cloud ecosystems (AWS/GCP/Azure). Experience with IaC tools (Terraform, CloudFormation).
- MLOps Expertise: Experience building ML pipelines for training, validation, and deployment. Familiarity with MLflow, Kubeflow, SageMaker, or Vertex AI. Understanding of distributed training and GPU/TPU optimization. Experience with vector databases such as FAISS, Pinecone, or Milvus.
- Observability & Reliability: Experience with Prometheus, Grafana, ELK, or OpenTelemetry. Strong debugging and incident-response skills.
- Data & Distributed Systems: Experience handling large datasets and training jobs. Comfortable with data pipelines (ETL processes) and tools for distributed computing or parallel processing. Understanding of how to optimize training and inference on hardware (GPUs/TPUs) to improve throughput.
- MLOps and Deployment: Experience deploying ML models into production. Familiar with containerization (Docker) and possibly orchestration (Kubernetes) for services, CI/CD pipelines for ML (e.g., automating model builds and deployments), and monitoring/observability tools to track model performance post-deployment. Knowledge of version control for models/data (MLflow or similar) is valuable.
- Software Engineering Practices: Strong CS fundamentals (algorithms, data structures) and ability to write production-quality code. Experience with code reviews, unit/integration testing, and collaborative development (Git or similar version control workflows).
- Problem Solving: Demonstrated ability to tackle ambiguous or open-ended problems in a structured way. Comfortable formulating experiments, evaluating results (using appropriate metrics), and iterating to improve model performance.
- Communication & Teamwork: Excellent communication skills. Able to explain complex technical concepts to both technical colleagues and non-technical stakeholders. Collaboration is central at Avataar – you should be able to work effectively in a team, share knowledge, and also take initiative independently when needed.
Cultural Fit (What We Value):
- Curiosity & Learning Mindset: You have a natural curiosity and passion for learning. You stay updated on new developments in AI and are eager to experiment with new techniques, continually growing your expertise.
- Ownership & Innovation: You take initiative and feel a sense of ownership over your projects. We love team members who propose creative solutions and are excited about building innovative products that have tangible impact.
- Collaboration & Communication: You work well with others, knowing that building complex AI systems is a team sport. You communicate openly, respect diverse viewpoints, and enjoy brainstorming and solving problems in a group setting.
- Pragmatic Impact: You are motivated by applying AI to solve real business problems. You balance theoretical excellence with a pragmatic approach to deliver solutions that work reliably in the real world. An applied impact focus – aiming to create value for users – is at the heart of our mission.
- Adaptability: In the fast-evolving AI landscape, you are adaptable and resilient. When faced with a new challenge or a change in direction, you embrace it as an opportunity to learn and contribute in new ways.