Role Senior Database Reliability Engineer
Location Remote
Fulltime
Job Description
Must Have Technical/Functional Skills
- 5+ years of experience designing, operating, and troubleshooting PostgreSQL in production environments with previous work with cloud infrastructure and managed data services, including PostgreSQL on Kubernetes, Amazon RDS, AWS, Terraform, service discovery, and secrets management.
- 5+ years of experience managing production database or distributed data systems across application, database, operating system, storage, and network layers.
- 3+ years of experience in Linux systems engineering (performance tuning, memory management, I/O tuning, configuration, security, and networking) and automating infrastructure or database operations with tools such as Terraform, Ansible, Chef, or Puppet.
- 2+ years of experience using at least one scripting or programming language such as Python, Bash, Go, Ruby, or Perl for automation and operational tooling.
- Experience working in a polyglot production data environment, including at least one non-PostgreSQL system such as Kafka/MSK, Click House, Redis, MySQL, Cassandra, Elasticsearch, or a similar distributed data system.
Roles & Responsibilities
- As a Senior Database Reliability Engineer, you will help make Meraki's database and data systems reliable, scalable, secure, and operable. You will combine database engineering, reliability engineering, and platform automation to support relational databases, streaming systems, analytical stores, and low-latency data services.
- You will work closely with SRE, application engineering, security, and infrastructure teams to improve the availability and performance of our data systems, reduce operational toil, and guide safe architectural changes as the platform grows.
- Plan, administer, maintain, and secure the PostgreSQL infrastructure in collaboration with Site Reliability Engineering (SRE) teams to ensure high performance and reliability.
- Design, build, and maintain ETL pipelines for PostgreSQL, as well as develop procedures and scripts for data migration.
- Perform operational database administration tasks such as installation, upgrades, patching, backup/recovery, monitoring, capacity planning, and architectural changes in cloud environments.
- Participate in production operations such as on-call rotation, incident response, monitoring, alerting, and post-incident review processes.
Good to Have:
Preferred Qualifications
- Experience building database platform tooling, self-service workflows, and paved paths that enable application teams to use data systems safely and efficiently.
- Ability to operate MSK/Kafka, Click House, or Redis at scale, covering cluster operations, replication, partitioning or sharding, retention, capacity planning, and workload tuning.
- Previous responsibility for refining reliability practices for production data systems, such as SLOs, disaster recovery plans, backup validation, and failover testing.
- Deep knowledge of PostgreSQL internals and operational behaviour, including concurrency, transaction consistency, replication, maintenance, backup and recovery, indexing, and query performance.
- Ability to communicate effectively in writing and verbally by producing design documents, leading.