Create Alert
Email me similar jobs

Remote Site Reliability Engineer(DevOps)

Remote Friendly Full-time Linux Hybrid GitHub Trailblazer Product Engineering
Nexthink is the leader in digital employee experience management software. The company provides IT leaders with unprecedented insight allowing them to see, diagnose and fix issues at scale impacting employees anywhere, with any applicationor network, before employees notice the issue. As the first solutionto allow IT to progress from reactive problem solving to proactive optimization, Nexthink enables its more than 1,300 customers to provide better digital experiences to more than 18 millionemployees. Dual headquartered in Lausanne, Switzerland and Boston, Massachusetts, Nexthink has 9 offices worldwide.
#LI-Hybrid
We deliver unmatched visibility across all environments, so IT teams can consistently see, diagnose, and fix digital workplace issues. We are looking for an experienced, proactive and innovative professional that is keen to join as a Senior Site Reliability Engineer! They work closely with over 50 Product Engineering teams that develop our products and services, as well as with the Technical Platform Engineering, Security and Architecture teams to understand the reliability requirements, design and implement solutions, and promote them for adoption and usage.
Be a part of Nexthink's Digital Employee Experience technological revolution, ensuring our global customers enjoy a seamless user experience. As a Senior Site Reliability Engineer, you will:
Implement and manage cloud-native systems (AWS) using best-in-class tools and automation.
Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles.
Define and maintain SLOs, SLAs, and error budgets, and proactively address availability and performance issues.
Develop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning.
Build internal platform tools and automation to support provisioning, monitoring, and operational efficiency.
Monitor infrastructure and applications ensuring high-quality user experiences.
Work closely with software engineers to embed observability, fault tolerance, and reliability principles into service design.
Support automated testing, canary deployments, and rollback strategies to ensure safe, fast, and reliable releases.
Minimum Bachelor’s degree in Computer Science or equivalent practical experience.
~5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices.
~ Strong hands‑on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS product.
~ Python, Go, Bash...), Terraform).
~ Proficiency with Kubernetes, container‑based deployment (e.g., Docker) and related ecosystems (e.g., Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane).
~ Experience with managing monitoring solutions (e.g. Datadog).
~ Deep understanding of Linux systems, networking, and common troubleshooting practices.
~ Solid understanding of the network stack (e.g., cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (e.g., Istio) and storage (e.g., S3, EBS, etc).
~ Knowledge of zero‑downtime deployment strategies, blue/green and canary releases.
~ Experience with chaos engineering or resilience testing practices.
~ Excellent problem‑solving skills, collaborative mindset, and a strong grasp of agile, iterative development.
~ Excellent written and verbal skills in English.
We are the pioneers and trailblazers of a global IT Market Category (DEX) that is shaping the future of how the world works, giving our customers’ IT Teams total digital visibility across their enterprise. This enables our IT teams to solve complex technical challenges, create ever more productive workplaces, and deliver happy, satisfied employees in the digital workplace.
Permanent Contract and a competitive compensation package.
Health insurance through our partnership with ACKO, including OPD coverage for dental, vision, health check‑ups, consultations, and pharmacy expenses.
Hybrid work model balancing office and remote work, with a structured approach for new hires to foster connections and onboarding.
Flexible Hours and unlimited vacation (employees have unlimited paid time off on top of the 22 days of holidays we offer). Plus, company‑paid bank holidays (12), sick days (10‑30), bereavement leave (5), and 3 days per year for volunteering.
Free access to professional training platforms to explore your interests and enhance your skills.
Gratuity is payable at retirement or resignation based on your last drawn basic pay.
Please note that not all the benefits listed above are available for temporary, contract, and internship roles.
Similar jobs

Remote Site Reliability Engineer(DevOps)

Apply Now
Back to search page