About The Role
The role owns the reliability, scalability, and security of core infrastructure supporting high-traffic production environments.
Key Responsibilities
- Design, build, and maintain production infrastructure using Terraform and Ansible across AWS and Kubernetes
- Implement comprehensive monitoring, logging, and alerting systems using Prometheus, Grafana, and Datadog to ensure high availability
- Automate CI/CD pipelines using GitHub Actions or GitLab CI to streamline zero-downtime deployments
- Manage container orchestration platforms, optimizing cluster performance, resource utilization, and auto-scaling policies
- Participate in an on-call rotation to troubleshoot and resolve production incidents efficiently
- Enforce security best practices, identity access management, and compliance standards across all cloud environments
What We Are Looking For
- 3-6 years of experience in DevOps, Site Reliability Engineering, or infrastructure engineering roles
- Strong proficiency in Infrastructure as Code using Terraform and configuration management tools
- Hands-on experience with Kubernetes, Docker, and container networking in production environments
- Solid scripting skills in Python, Bash, or Go for automation and tooling
- B.S. in Computer Science, related technical field, or equivalent practical experience
- Bonus: Experience with service mesh technologies like Istio, or hands-on migration projects to cloud-native architectures
#J-18808-Ljbffr