Start Your Search Here

push notification bell

Would you like to receive notifications about Computer and Mathematical Occupations jobs in Sunnyvale?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Resource Logistics

Sunnyvale / Global

Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE)

Job Description

Job Title: Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE)

Location: Sunnyvale, CA

Job Summary

We are seeking a highly skilled CockroachDB Database Engineer with strong Site Reliability Engineering (SRE) experience to design, implement, manage, and optimize large-scale distributed database platforms. The ideal candidate will have hands-on expertise in CockroachDB administration, performance tuning, high availability, disaster recovery, automation, observability, and operational reliability. The role requires close collaboration with development, infrastructure, and platform engineering teams to ensure highly available, resilient, and scalable database services.

Key Responsibilities

Design, deploy, administer, and maintain production-grade CockroachDB clusters across cloud and on-premises environments.

Monitor database health, performance, latency, throughput, and resource utilization to ensure service reliability and availability.

Implement and manage backup, restore, disaster recovery, and business continuity strategies.

Perform database capacity planning, performance tuning, indexing, and query optimization.

Develop automation scripts and Infrastructure-as-Code (IaC) solutions to streamline provisioning, upgrades, and operational tasks.

Establish and manage SRE practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

Drive incident management, root cause analysis (RCA), postmortems, and preventive re,tion activities.

Build and maintain monitoring, logging, and alerting solutions using tools such as Prometheus, Grafana, ELK, Datadog, or similar platforms.

Collaborate with DevOps and Engineering teams to improve platform reliability, scalability, security, and operational excellence.

Support production releases, database migrations, version upgrades, and platform modernization initiatives.

Participate in on-call rotation and provide support for critical production incidents.

Implement database security controls, access governance, auditing, and compliance best practices.

Apply Now

Similar Opportunities

View all jobs

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location