Create Alert
Email me similar jobs

Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE)

a { text-decoration: none; color: #464feb; } tr th, tr td { border: 1px solid #e6e6e6; } tr th { background-color: #f5f5f5; }

Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE)

Location: Austin, TX or Sunnyvale, CA (Onsite)

About Us

STAFFXPERT LLC is a trusted talent solutions partner connecting highly skilled professionals with leading organizations across technology, engineering, cloud infrastructure, and digital transformation initiatives. We are committed to matching exceptional talent with innovative opportunities that drive business success and career growth.

Job Summary

STAFFXPERT LLC is seeking a Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE) on behalf of our client in Austin, TX or Sunnyvale, CA.

This role is ideal for an experienced database and reliability engineering professional who excels in managing large-scale distributed database platforms. The successful candidate will be responsible for designing, administering, optimizing, and supporting highly available CockroachDB environments while driving automation, observability, performance, and operational excellence across mission-critical systems.

Key Responsibilities

  • Design, deploy, administer, and maintain production-grade CockroachDB clusters in cloud and hybrid environments.

  • Monitor database health, performance, availability, and resource utilization to ensure reliable operations.

  • Perform database performance tuning, query optimization, indexing strategies, and capacity planning.

  • Implement and manage backup, recovery, disaster recovery, and business continuity solutions.

  • Develop automation and Infrastructure-as-Code (IaC) solutions to streamline provisioning, upgrades, and operational tasks.

  • Establish and maintain Site Reliability Engineering (SRE) practices, including SLIs, SLOs, and Error Budgets.

  • Lead incident response, troubleshooting, root cause analysis (RCA), and post-incident remediation activities.

  • Build and maintain monitoring, logging, and alerting solutions using industry-standard observability tools.

  • Collaborate with engineering, DevOps, and infrastructure teams to improve platform reliability, scalability, security, and performance.

  • Support database migrations, production releases, version upgrades, and modernization initiatives.

  • Participate in on-call support for critical production environments.

  • Implement database security, access controls, auditing, and compliance best practices.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.

  • 7+ years of experience in database engineering, administration, or site reliability engineering.

  • Strong hands-on experience administering and supporting CockroachDB or similar distributed SQL database platforms.

  • Deep understanding of distributed systems, cluster management, replication, and high-availability architectures.

  • Proven expertise in database performance tuning, query optimization, and capacity planning.

  • Experience designing and implementing backup, recovery, and disaster recovery strategies.

  • Strong knowledge of Site Reliability Engineering (SRE) principles and operational best practices.

  • Experience with incident management, reliability engineering, and production support.

  • Hands-on experience with AWS, Azure, or Google Cloud Platform (GCP).

  • Proficiency with Infrastructure-as-Code tools such as Terraform or Ansible.

  • Strong Linux/Unix administration skills.

  • Scripting experience using Python, Shell, Go, or similar languages.

  • Experience with CI/CD pipelines and automation frameworks.

  • Excellent analytical, troubleshooting, communication, and collaboration skills.

Preferred Qualifications

  • Experience supporting large-scale, mission-critical distributed systems.

  • Knowledge of Kubernetes and containerized platforms.

  • Experience with observability tools such as Prometheus, Grafana, Datadog, ELK, Splunk, or similar solutions.

  • Understanding of database security, governance, compliance, and auditing requirements.

  • CockroachDB certification or equivalent expertise in distributed database technologies.

  • Experience with PostgreSQL internals and PostgreSQL-compatible ecosystems.

  • Knowledge of multi-region architectures, distributed consensus mechanisms, and cloud-native platforms.

  • Experience in FinTech, Retail, E-Commerce, SaaS, or other high-scale environments.

Similar jobs

Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE)

Apply Now
Back to search page