Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.

1. About the Company/Team

Our team builds and operates scalable, reliable, and secure cloud-based distributed systems that support mission-critical business and customer needs. We focus on engineering resilient platforms, improving service availability, strengthening operational excellence, and enabling seamless growth through automation, observability, and modern system design practices.

We value collaboration, technical ownership, continuous learning, and a strong customer-first mindset. Team members work across engineering, operations, security, and product stakeholders to deliver high-quality systems that are performant, secure, and designed for long-term reliability.

2. Job Summary

We are seeking a Software Engineer with strong experience in software design, development, and operations of large-scale distributed systems. In this role, you will contribute to the design, implementation, scalability, reliability, and security of cloud infrastructure components and services.

You will work on distributed system components, automation tooling, monitoring dashboards, incident response processes, and change management practices. This role is ideal for an engineer who enjoys solving complex technical problems, writing high-quality code, improving system resilience, and supporting production-grade cloud services.

3. Key Responsibilities

  • Design, implement, and enhance components of distributed systems that support horizontal and vertical scalability, including distributed state management and large-scale data processing.
  • Build fault-tolerant system components by implementing redundancy, replication, automatic failover, retry mechanisms, circuit breakers, and timeout strategies.
  • Develop tests, dashboards, telemetry, alarms, and alerting mechanisms to proactively detect, monitor, and resolve component-level health and performance issues.
  • Implement functional requirements, fault-injection tests, brown-out scenarios, data replication, and synchronization techniques to maintain correctness, integrity, and availability.
  • Diagnose, debug, and resolve operational issues in production systems, while contributing to incident response, root cause analysis, and operational support rotations.
  • Create and maintain automation scripts, troubleshooting tools, and Infrastructure as Code solutions to support cloud infrastructure management and operational efficiency.
  • Apply security controls such as encryption, access management, and remediation plans to protect applications and data in multi-tenant cloud environments.
  • Follow change management practices for patching, upgrades, rollbacks, and releases while minimizing customer impact and avoiding unnecessary maintenance windows.

  • 4. Qualifications & Skills

    Mandatory Skills

  • BS or MS degree in Computer Science or equivalent practical experience.
  • 4+ years of experience in software design and development, including building and operating large-scale distributed systems.
  • Strong understanding of microservices, data structures, algorithms, operating systems, and distributed systems concepts.
  • Solid knowledge of relational databases, NoSQL systems, storage platforms, and distributed persistence mechanisms.
  • Strong coding skills in Java or another object-oriented programming language.
  • Experience implementing scalable, reliable, and fault-tolerant system components.
  • Ability to diagnose and troubleshoot production issues using logs, metrics, dashboards, and operational tools.
  • Strong collaboration, problem-solving, and communication skills.
  • Self-Assessment Questions

    Before applying, candidates may reflect on the following questions:

  • Do I have 4+ years of hands-on experience designing, developing, and operating large-scale distributed systems?
  • Am I confident in my understanding of microservices, data structures, algorithms, operating systems, and distributed systems principles?
  • Have I worked with relational databases, NoSQL systems, storage platforms, or distributed persistence mechanisms in production environments?
  • Can I write strong, maintainable code in Java or another object-oriented programming language?
  • Have I diagnosed and resolved production issues using logs, metrics, dashboards, alerts, or other operational troubleshooting tools?
  • Career Level - IC3


    Senior Core Infrastructure Engineer - Java

    Apply Now
    Back to search page