This is a highly hands-on role focused on platform reliability, production support, and infrastructure.
The ideal candidate enjoys solving complex operational problems, automating repeatable work, improving system resilience, and supporting service health in a fast-moving environment.
Responsibilities: · Own and support the reliability of Voice Assistant services and the underlying cloud infrastructure.
· Participate in an on-call rotation to support production systems outside normal business hours.
· Lead and participate in incident response, including triage, escalation, mitigation, and restoration of service.
· Drive blameless postmortems and ensure corrective actions are tracked through completion.
· Work to improve customer experiences by strengthening service availability, latency, stability, and resilience against agreed service-level expectations.
· Design, implement, and maintain infrastructure as code using tools such as Terraform and Atlantis.
· Manage and extend GitOps and deployment workflows using ArgoCD and related CI/CD tooling.
· Support and improve cloud and container platforms across AWS and Azure.
· Independently manage virtual servers, containers, and orchestration platforms such as Kubernetes.
· Build, refine, and operate automation that reduces toil and improves operational efficiency.
· Utilize and extend existing observability and reliability capabilities, including monitoring, alerting, logging, and diagnostics.
· Review system performance and capacity trends to identify bottlenecks, support forecasting, and improve scalability.
· Troubleshoot complex infrastructure, networking, and application runtime issues across distributed systems.
· Assist with disaster recovery planning, validation, and recovery readiness.
· Contribute to performance tuning and resilience improvements across infrastructure and services.
· Document operational procedures, support runbooks, and engineering knowledge to improve team effectiveness.
· Coach and support other engineers by sharing operational best practices and reliability engineering approaches.
Required Qualifications: · 3+ years of production experience working as a Site Reliability Engineer, DevOps Engineer, Infrastructure Engineer, or Software Engineer · Experience working with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools.
· Strong experience with AWS and Azure.
· Expertise in Terraform to create, modify, and manage infrastructure configurations or IaC templates · Expertise in containerization technologies (Docker & Kubernetes) to build, package, and deploy optimized container images · Proficiency in designing, implementing, and maintaining complex CI/CD pipelines that span with increasing complexity and integrating across multiple environments · Knowledge in cloud platforms (AWS) to optimize cloud resource utilization and costs throughout product lifecycle · Expertise in version control systems to perform branching, merging, and resolving merge conflicts · Versed within InfoSec policies and procedures to adhere to security standards/regulations and identify gaps in security architecture · Experience in monitoring and analytics platforms to set up monitors, alerts, and diagnostic tools for proactive issue detection, root cause analysis, and performance optimization across distributed systems · Understanding of cloud billing and cost management tools to recognize total costs · Ability to learn and apply new technologies, programming practices, patterns, and methods · Organized and detail-oriented · Ability to develop healthy working relationships and collaborate with peers and leaders · Exhibits integrity and high standards in work quality · Excellent verbal and written communication skills · Values diversity and differences amongst individuals in interactions
By continuing you agree to our Terms & Privacy Policy.