Start Your Search Here

push notification bell

Would you like to receive notifications about jobs in Seattle?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Amazon

Seattle / Global

Sr. Software Engineer- AI/ML, AWS Neuron

Job Description

Senior Software EngineerShape the Future of AI Accelerators at AWS Neuron. We build Amazon Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon's custom machine learning accelerators, Inferentia and Trainium. As a Senior Software Engineer on our Machine Learning Applications team, you will optimize the world's most demanding AI models at a scale few engineers ever get to work on. This role offers a unique opportunity to work at the intersection of machine learning, high-performance computing, and distributed architectures, where you'll help shape the future of AI acceleration technology.Key job responsibilities:Design, develop, and optimize machine learning models including GPT, Kimi, and Qwen on custom AI accelerators.Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, low level optimizations, and production deployment.Build infrastructure to systematically analyze and onboard multiple models with diverse architecture.Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models.Analyze and optimize system-level performance across multiple generations of Neuron hardware.Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks.Implement optimizations such as fusion, sharding, tiling, and scheduling.Work directly with customers to enable and optimize their ML models on AWS accelerators.Collaborate across teams to develop innovative optimization techniques.What makes this role unique:Direct influence on AWS's AI Accelerator used by thousands of ML applications.Full-stack optimization from high-level frameworks to low level primitives.Collaboration with both open-source ML communities and hardware architecture teams.Requires passion for performance tuning and system architecture.Basic qualifications:Bachelor's degree.5+ years of non-internship professional software development experience.Knowledge of Python and/or C++ programming.5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.Experience in debugging, profiling, and implementing software engineering best practices in large-scale systems.Knowledge of system performance, memory management, and parallel computing principles.Experience owning a performance optimization roadmap and mentoring engineers on optimization.Preferred qualifications:Master's degree in computer science or equivalent.Knowledge of machine learning model architecture and inference.Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.Hands-on development with PyTorch is preferred.Experience scaling workloads across multi-GPU and multi-node topologies with NCCL and tensor, pipeline, or expert parallelism.Experience writing and optimizing custom CUDA/Triton kernels for tensor operations.Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers.
Apply Now

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location