Start Your Search Here

push notification bell

Would you like to receive notifications about jobs in San Diego?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Jobrapido

San Diego / Global

Principal Software Engineer, Distributed Systems

Job Description

Background

We are hiring a Staff / Principal Site Reliability Engineer focused on distributed systems and platform reliability . This is not a traditional DevOps role.

You will own the architecture and engineering practices that ensure Sigma workflows execute reliably even when individual components fail.

Processes crash. Networks become unavailable. APIs time out. Workers hang. Messages may be delivered more than once. Databases experience contention. Kubernetes pods restart.

Your job is to design Sigma so those conditions are expected, detected, and automatically recovered from.

The fundamental principle is simple:

An accepted Sigma job may fail, but it should never disappear.

You will work closely with engineering leadership to evolve Sigma's execution architecture, improve production reliability, establish observability standards, and eliminate classes of distributed-system failures before they affect customers.

What You'll Own

Design and improve the systems responsible for executing long-running and scheduled infrastructure workflows.

Build patterns for:

Durable workflow execution

Idempotent operations

Automatic retries and exponential backoff

Worker leases and heartbeats

Stuck-job detection and recovery

Workflow timeouts and cancellation

Failure isolation

Dead-letter handling

Checkpointing and workflow resumption

Concurrency control

Distributed locking and lease management

Reconciliation loops

Graceful degradation during dependency failures

Evaluate and implement workflow technologies such as Temporal or equivalent systems where appropriate.

Messaging and Event Infrastructure

Own and improve Sigma's asynchronous execution architecture.

Work extensively with technologies such as:

Kafka

Distributed consumers

Consumer groups

Message delivery semantics

Partitioning

Consumer lag

Backpressure

Retry strategies

Event ordering

Duplicate message handling

Poison-message isolation

Ensure event-processing failures cannot silently cause workflows to stop progressing.

Database Reliability

Help design application patterns that remain reliable under high concurrency.

Deeply understand and troubleshoot:

PostgreSQL transactions

Row and advisory locking

Deadlocks

Lock contention

Transaction isolation

Connection pooling

Long-running transactions

Query performance

Database failover

Schema and migration safety

Partner with application engineers to identify and eliminate patterns that create production contention or reliability risks.

Kubernetes and Platform Reliability

Improve the reliability of Sigma deployments running in containerized environments.

Responsibilities include:

Kubernetes workload architecture

Pod lifecycle and failure recovery

Resource limits and capacity planning

Autoscaling

Health checks

Graceful shutdown

Rolling deployments

High availability

Infrastructure-as-code

Disaster recovery

Backup and restore validation

Design systems that tolerate node, pod, process, and dependency failures without losing customer work.

Observability

Make the state of Sigma understandable in real time.

Build and improve:

Metrics

Distributed tracing

Structured logging

Dashboards

Alerting

Workflow-level telemetry

Kafka consumer monitoring

Database performance monitoring

Worker health monitoring

Queue depth and backlog monitoring

Technologies may include:

OpenTelemetry

Grafana

Prometheus

Sentry

CloudWatch

Engineers should be able to answer questions such as:

Why is this workflow still running?

Which worker owns this task?

When did it last make progress?

Which dependency is causing the delay?

Did this operation execute once or multiple times?

Can the workflow safely resume?

Are scheduled jobs starting when expected?

Reliability Engineering

Help establish engineering practices expected of mature enterprise platforms.

This includes:

Service-level objectives and indicators

Error budgets

Capacity planning

Load testing

Failure-mode testing

Chaos testing

Incident response

Root-cause analysis

Production readiness reviews

Runbooks

Automated recovery

Disaster-recovery testing

Move Sigma from detecting incidents to preventing and automatically recovering from them.

What We're Looking For

You have significant experience operating production distributed systems where reliability matters.

Strong candidates will have deep experience with several of the following:

Kubernetes

Kafka or similar event-streaming systems

PostgreSQL

Distributed systems

Asynchronous worker architectures

Workflow orchestration

Temporal, Cadence, Step Functions, Conductor, or similar systems

Python backend systems

Cloud infrastructure

Infrastructure-as-code

Observability platforms

High-availability architectures

You understand concepts such as:

At-least-once delivery

Idempotency

Leases and fencing tokens

Distributed locks

Leader election

Retry semantics

Backpressure

Eventual consistency

Transaction boundaries

Failure domains

Reconciliation

Split-brain scenarios

Distributed tracing

Most importantly, you have personally diagnosed and improved production systems experiencing real distributed-system failures.

What Success Looks Like

Within your first several months, you will help Sigma establish an architecture where:

Scheduled workflows reliably begin when expected

Accepted work cannot silently disappear

Worker crashes do not result in lost workflows

Infrastructure failures automatically recover where possible

Long-running workflows can safely resume

Duplicate execution is prevented or safely handled

Stuck workflows are automatically detected

Database and messaging contention is visible before becoming an outage

Engineers can trace a workflow from request through every execution step

Customer environments can scale without requiring manual infrastructure babysitting

Apply Now

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location