Datacenter Platform Engineering Group (DPEG) System Level Debug EngineerAt AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we're looking for talent who feel the same: people who want to leave the planet better than they found it, those who don't shy away from humanity's challenges but are determined to help solve them.AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you're designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.The Datacenter Platform Engineering Group (DPEG) organization is looking for an experienced system level debug engineer. Individual will be part of a team that will be responsible for Data Center deployments as well as ensuring availability/uptime of a large number of systems. Individual should be familiar with System level debug at all levels as well have RAS knowledge to quickly root cause issues. Person should have not only strong technical skills but be able to direct junior level engineers to solving problems and guide them to be successful in handling service level tickets.Experience in debugging of complex HW/FW issues is a must, understand the flow of a GPU through the different layers of a system and be able to validate the items connecting to the GPU SOC (pcie, vr's, RMs, retimers, HBM, internal networking). Handling of Data Center related issues that come with electrical complexities, direct liquid cooling and complex cluster networks. Communication Is essential in working with different owners of the functional code stack as well as the ability to drive issues via phone calls, chat messages, e-mails. Hands on experience with Hardware in a Data Center environment will be required.Key Responsibilities:Debug / triage engineer and understanding of industry tools for root causing complex issuesUnderstanding of GPU/System level HW and SW flowAbility to probe parts of a board; check electrical and power currents and validate a systemRAS knowledge and flow of issues seen at a GPU levelProvide leadership for driving to root cause issues and guide junior level engineersDocument flows and methods of bring-up, boot-up, system initialization and debugPreferred Experience:Experience in Systems Integration, DebugProven ability to drive resolution of critical problems within a lab, Data CenterRelationship with external customers/partners and able to help resolve problems in a Data CenterRelationship with external customers/partners on ability to work manufacturing issues/failuresRelationship with external customers/partners on ability to define requirements for deployments/debugSignificant experience in SoC and/or System debug of complex issuesDevelop / Document debug capabilities on a given SOC and SystemGo-to-person at a Data Center for resolving issues, looking at complex issuesCollaborate with internal teams on root causing issues, finding optimum resolutionsHands-on experience in using industry debug tools, scopes as well examine board level powerAcademic Credentials:Bachelors or Masters degree in electrical or computer engineeringLocation: Austin, TXThis role is not eligible for visa sponsorship.
Apply Now