Start Your Search Here

push notification bell

Would you like to receive notifications about jobs in Seattle?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Atlassian

Seattle / Global

Machine Learning System Engineer

Job Description

Machine Learning System EngineerEngineering | Seattle, United States | Remote, Remote | San Francisco, United States |As a ML System Engineer on the AI & ML Platform's Inference team, you will design and optimize large-scale model serving systems end-to-end. You will have the chance to own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding).In this role, you are expected to:Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).Optimize latency and throughput of model inference under real production workloads.Build reliable, high-concurrency serving systems that serve billions of requests reliablyBenchmark, fine-tune, and accelerate inference engines.Create robust CI/CD infrastructure for seamless model deployment and inference engine updates.Partner with senior ML engineers to fine-tune and deploy open-source LLMsOn your first day, we'll expect you to have:3+ years of software engineering experienceDeep low-level systems programming (C/C++ or Rust)Experience with large-scale, high-concurrent production serving.Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).It would be great, but not required if you have:1+ years of system performance optimization experienceLow-level inference optimizations: GPU kernelsAlgorithmic inference optimizations: quantization, speculative decoding, distillationExperience with testing, benchmarking, and reliability of inference services.Experience designing and implementing CI/CD infrastructure for inference.Strong background in system optimizations: batching, caching, load balancing, parallelism.
Apply Now

Similar Opportunities

View all jobs

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location