Start Your Search Here

push notification bell

Would you like to receive notifications about Computer and Mathematical Occupations jobs in Cupertino?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

LingaTech

Cupertino / Global

Software Engineer, ML Infrastructure

Job Description

Senior Software Engineer, ML Infrastructure

Location: Cupertino, San Franciso

Salary: $209,000 - $235,000 USD

Duration: FullTime

The Role

We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the foundational infrastructure that powers. Our platform challenges are shaped by the nature of energy markets: forecasts and trading decisions run on tight schedules, battery dispatch commands must execute reliably in real time, and ML models need to train and deploy continuously as new data arrives.

What You'll Do

Design and build the compute platform that runs clients production services, batch jobs, and ML training workloads

Help own our production workflow orchestration end to end, from the cluster and node pools it runs on to the abstractions our teams build on top of it

Make our ML iteration cycle faster and cheaper by profiling workflows, cutting latency, and improving how we use compute

Build the observability and developer tooling that lets engineers understand and troubleshoot their own workloads, from post-run cost reporting to monitoring and alerting

Improve the reliability and resiliency of production workloads, including how we handle capacity constraints and multi-region routing

Manage GPU and accelerator capacity: node pools, drivers, spot vs. on-demand tradeoffs, and scheduling and quota so training and batch jobs get the compute they need without overspending

Own autoscaling and quota management so the platform scales up under load and down to zero when idle

Work closely with the ML team to find pain points and quickly ship solutions

Make architectural decisions that shape how we build software as we grow

What We're Looking For

Significant experience building and operating production infrastructure on a public cloud platform (we run on GCP, but AWS or Azure experience translates well)

Strong distributed systems and infrastructure skills: standing up services, scaling and debugging Kubernetes (we run on GKE and it's foundational to our platform), writing Terraform, and comfort with cloud networking, IAM, and secrets management

Hands-on experience with workflow orchestration tools like Flyte, Temporal, or Airflow

Proficiency in Python, and either already know Go or have experience with a similar systems language (C++, Java, Rust) and are excited to work in Python and Go day-to-day

Experience working closely with ML engineers or data scientists as your customers, and genuine excitement to keep doing it

Clear communication, whether writing a design doc, reviewing code, or explaining a complex system to someone new to it

Nice to Have

Experience running GPU or accelerator workloads on Kubernetes (node pools, drivers, scheduling, quota) is strongly preferred for this role

Familiarity with observability tooling (we use Grafana + Google Cloud Monitoring)

Experience building internal developer platforms or tooling that other engineers rely on

Prior work in domains where latency and reliability have direct business consequences

Apply Now

Similar Opportunities

View all jobs

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location