LingaTech
Cupertino / Global
You have blocked notifications
Oops! You have blocked notifications. Click here for more info
You have blocked notifications, please check your browser settings.
You're currently subscribed to job notifications
Subscribe to notifications
You will no longer receive notifications
Cupertino / Global
Senior Software Engineer, ML Infrastructure
Location: Cupertino, San Franciso
Salary: $209,000 - $235,000 USD
Duration: FullTime
The Role
We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the foundational infrastructure that powers. Our platform challenges are shaped by the nature of energy markets: forecasts and trading decisions run on tight schedules, battery dispatch commands must execute reliably in real time, and ML models need to train and deploy continuously as new data arrives.
What You'll Do
Design and build the compute platform that runs clients production services, batch jobs, and ML training workloads
Help own our production workflow orchestration end to end, from the cluster and node pools it runs on to the abstractions our teams build on top of it
Make our ML iteration cycle faster and cheaper by profiling workflows, cutting latency, and improving how we use compute
Build the observability and developer tooling that lets engineers understand and troubleshoot their own workloads, from post-run cost reporting to monitoring and alerting
Improve the reliability and resiliency of production workloads, including how we handle capacity constraints and multi-region routing
Manage GPU and accelerator capacity: node pools, drivers, spot vs. on-demand tradeoffs, and scheduling and quota so training and batch jobs get the compute they need without overspending
Own autoscaling and quota management so the platform scales up under load and down to zero when idle
Work closely with the ML team to find pain points and quickly ship solutions
Make architectural decisions that shape how we build software as we grow
What We're Looking For
Significant experience building and operating production infrastructure on a public cloud platform (we run on GCP, but AWS or Azure experience translates well)
Strong distributed systems and infrastructure skills: standing up services, scaling and debugging Kubernetes (we run on GKE and it's foundational to our platform), writing Terraform, and comfort with cloud networking, IAM, and secrets management
Hands-on experience with workflow orchestration tools like Flyte, Temporal, or Airflow
Proficiency in Python, and either already know Go or have experience with a similar systems language (C++, Java, Rust) and are excited to work in Python and Go day-to-day
Experience working closely with ML engineers or data scientists as your customers, and genuine excitement to keep doing it
Clear communication, whether writing a design doc, reviewing code, or explaining a complex system to someone new to it
Nice to Have
Experience running GPU or accelerator workloads on Kubernetes (node pools, drivers, scheduling, quota) is strongly preferred for this role
Familiarity with observability tooling (we use Grafana + Google Cloud Monitoring)
Experience building internal developer platforms or tooling that other engineers rely on
Prior work in domains where latency and reliability have direct business consequences
Palo Alto / Global
Sunnyvale / Global
Mountain View / Global
Mountain View / Global
Palo Alto / Global
Mountain View / Global