Tesla
Austin / Global
You have blocked notifications
Oops! You have blocked notifications. Click here for more info
You have blocked notifications, please check your browser settings.
You're currently subscribed to job notifications
Subscribe to notifications
You will no longer receive notifications
Austin / Global
What to Expect This team manages multiple functions across Tesla that includes Platform Engineering managing fleet of Kubernetes clusters, Devops, MLOps, Cloud Infrastructure, Sandbox platform and Factory SRE as well. Continued development and automation of deployment, monitoring, self-healing, and alerting processes is imperative to the success of our engineering groups.
You will be responsible for the internal platform that lets every team securely spin up sandboxed AI agents, run ephemeral workloads, train models at scale, and serve them in production all as a reliable, self-service experience. This is a high-leverage, high-ownership role where you will design, build, and operate the systems that sit at the intersection of Kubernetes, high-performance networking, GPU infrastructure, and modern ML platforms.
What You'll Do Build and own the end-to-end AI/ML platform (training, inference, experimentation) as a self-service product for all internal users
Write production Kubernetes operators and controllers in Go for GPU workloads, training jobs, model deployments, sandboxes, and cluster lifecycle
Operate large-scale GPU fleets (A100/H100/B200) scheduling, MIG, topology-aware placement, health monitoring
Build and operate inference infrastructure using KServe, Triton, vLLM, and Ray Serve autoscaling, model versioning, and request batching
Build and operate training-as-a-service; distributed training (PyTorch , FSDP), MLflow, and checkpoint management
Design secure, isolated sandbox environments for AI agents and ephemeral execution contexts for untrusted code (gVisor, Kata, Firecracker)
Architect serverless/Lambda-style ephemeral workloads, scale-to-zero, event-driven compute (Knative, KEDA, Firecracker microVMs)
Integrate GPU compute, training, inference, and sandboxes as first-class services in the internal cloud platform
Manage Kubernetes cluster lifecycle at fleet scale using Cluster API (CAPI) provisioning, upgrades, scaling, and decommissioning
What You'll Bring Have a minimum of 8+ years of practical experience as a backend or platform engineer building scalable distributed systems, ideally operating at tech lead level , cloud-native products, developer tools, or external developer facing products
Built agent sandbox/ephemeral compute platforms (gVisor, Kata, Firecracker, or similar)
Strong fundamentals in service-oriented architectures, networking, and systems design, ideally having owned both the technical vision and execution of a foundational platform system end to end
Strong profciency in Python, Go, Rust, or similar systems languages (full stack)
Production experience with KServe, Triton, vLLM, or Ray Serve for model inference
Built or operated training-as-a-service, distributed training, MLflow, job scheduling
Root cause analysis and systems thinking - failure modes, back-pressure, resource contention, blast radius
Deep Kubernetes internals expertise (scheduler, API server, etcd, admission controllers, CRDs, control plane at scale)
Built production Kubernetes operators in Go (controller-runtime / Kubebuilder)
Production Cluster API (CAPI) experience; management clusters, custom providers, ClusterClass, fleet-scale lifecycle
Benefits
Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:
Medical plans > plan options with $0 payroll deduction
Family-building, fertility, adoption and surrogacy benefits
Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
Healthcare and Dependent Care Flexible Spending Accounts (FSA)
401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
Company paid Basic Life, AD&D
Short-term and long-term disability insurance (90 day waiting period)
Employee Assistance Program
Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
Back-up childcare and parenting support resources
Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
Weight Loss and Tobacco Cessation Programs
Tesla Babies program
Commuter benefits
Employee discounts and perks program
?
It doesn’t matter where you come from, where you went to school or what industry you’re in—if you’ve done exceptional work, join us to rethink the future of sustainable energy and manufacturing.
Austin / Global
Austin / Global
Austin / Global
Austin / Global
Austin / Global
Austin / Global