Start Your Search Here

push notification bell

Would you like to receive notifications about IT & Technology, Engineering jobs in Seattle?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Amazon Web Services (AWS)

Seattle / Global

Systems Development Engineer, AWS Generative AI & ML Servers

  • $129.000 - $175.000

Job Summary

Salary Range:
$129.000 - $175.000
Apply Now

Job Description

Description

Do you want to build the backbone of Generative AI at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price performance improvements for multi-billion variable LLMs at cloud scale? Come join us. We are seeking a Systems Development Engineer to develop automation software, diagnostic tooling, and fleet health infrastructure for our accelerated (AI/ML) server platforms. You will work across multiple teams and organizations to build scalable, reliable systems that keep our fleet healthy — with a vision toward zero-touch operations where automation detects, diagnoses, and resolves issues without human intervention.

Description

Do you want to build the backbone of Generative AI at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price performance improvements for multi-billion variable LLMs at cloud scale? Come join us. We are seeking a Systems Development Engineer to develop automation software, diagnostic tooling, and fleet health infrastructure for our accelerated (AI/ML) server platforms. You will work across multiple teams and organizations to build scalable, reliable systems that keep our fleet healthy — with a vision toward zero-touch operations where automation detects, diagnoses, and resolves issues without human intervention.

Description

Do you want to build the backbone of Generative AI at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price performance improvements for multi-billion variable LLMs at cloud scale? Come join us. We are seeking a Systems Development Engineer to develop automation software, diagnostic tooling, and fleet health infrastructure for our accelerated (AI/ML) server platforms. You will work across multiple teams and organizations to build scalable, reliable systems that keep our fleet healthy — with a vision toward zero-touch operations where automation detects, diagnoses, and resolves issues without human intervention.

What You Will Do

You will solve complex architectural problems that may not be well-defined in advance. You will own your team's systems, proactively identify deficiencies, and write scalable, robust code to solve issues before they impact customers. You will decompose large, difficult server testability, reliability, and diagnosis problems into straightforward tasks and components — delivering yourself and through others in parallel — using a combination of hardware, software, system design, processor architecture, diagnostics, and operations knowledge.

Key job responsibilities

Fleet Health & Predictive Infrastructure

Build and own the automation infrastructure responsible for the health of the accelerator (AI/ML) compute server fleet

Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact

Drive toward zero-touch operations — building automation that detects, diagnoses, triages, and remediates hardware and software faults without human intervention

Develop monitoring tools, dashboards, and alerting systems to provide real-time visibility into fleet health across lab and production environments

Define and track fleet health metrics (failure rates, mean time to detect, mean time to repair, first-time fix rate, predictive accuracy)

Debugging & Troubleshooting

Debug and resolve complex system-level issues across compute, GPU, and networking in production environments

Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including PCIe, power, NIC, NVMe, and GPU subsystems

Perform root cause analysis on hardware failures — correlating across firmware, kernel, driver, and physical layer to isolate faults

Build diagnostic tooling that automates root cause identification and reduces reliance on manual triage

Improve manufacturing throughput and yield through test optimization

Systems Development & Automation

Define and develop software, automation, and enabling tools for server hardware programs; track and report progress

Design and build scalable system-level software with focus on durability, availability, security, and diagnostics

Develop and maintain device drivers for Linux on ARM and x86 architectures

Build automation solutions using modern programming languages (Python, Ruby, Java, C/C++, etc.)

Work with OS internals and accelerator/GPU software stacks in Linux-based environments

Build, manage, and deploy CI/CD pipelines for rapid deployment of code changes to org-owned and customer-owned systems

Cross-Team Collaboration

Work across internal HWEng teams to ensure new server hardware addresses data path and control path functionality needed by dependent service teams

Work closely with internal customers to identify early any potential problems onboarding new accelerated compute servers into their ecosystem

Engage with ODMs and design partners on testability, diagnostic, and automation requirements during hardware design and development (NPI)

Contribute to server design to improve robustness, testability, diagnosability, and reliability

Partner with datacenter operations teams to close the loop between field failures and design improvements

A day in the life

You will collaborate with a variety of roles (SDEs, SDETs, Mechanical/Electrical/Hardware Engineers, TPMs, Managers, Principals) and organizations through server conception, test validation, qualification, launch, and operations — driving high quality and reliability into current and future designs for AWS accelerated server solutions. From orchestration tooling development to hardware integration to kernel driver debugging, you dive deep into problems across the breadth of AWS.

About The Team

The Hardware Engineering AI/ML development team is a group of engineers and technical program managers directly responsible for launching and maintaining server hardware in the fleet — including AI/ML accelerator servers with GPUs. Located in Seattle, Cupertino, and Austin, we work with internal development teams, ODMs, and design partners to deliver servers deployed in datacenters worldwide.

Basic Qualifications

2+ years of non-internship professional software development experience

1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience

Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby

Preferred Qualifications

Familiarity with server hardware architecture, BMC/IPMI, firmware, PCIe topology, and hardware diagnostics

Experience working with ODMs or hardware design partners

Exposure to zero-touch or self-healing automation concepts for large-scale infrastructure

Experience working in large-scale datacenter or cloud environments

Experience with hardware bring-up, validation, or fleet-wide deployment

Familiarity with telemetry pipelines, anomaly detection, or operational metrics at scale

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 148,700.00 - 201,200.00 USD annually

USA, TX, Austin - 129,200.00 - 174,800.00 USD annually

USA, WA, Seattle - 129,200.00 - 174,800.00 USD annually

Company - Amazon.com Services LLC

Job ID: A10525890

#J-18808-Ljbffr

Apply Now

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location