Talent.com
Advanced Micro Devices, Inc
Principal Software Developer – AI/ML Performance Validation & Systems TestingAdvanced Micro Devices, Inc • San Jose, California, United States
Principal Software Developer – AI/ML Performance Validation & Systems Testing

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, Inc • San Jose, California, United States
13 hours ago
Job type
  • Full-time
Job description


ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.




THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID




Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Create a job alert for this search

Principal Software Developer – AI/ML Performance Validation & Systems Testing • San Jose, California, United States

Similar jobs

ML Engineer for AI Product Tools -- Hybrid/Remote

AndiamoPalo Alto, CA, United States
Remote
Full-time

A leading technology staffing firm seeks a Machine Learning Engineer to work on real products rather than isolated research.This role emphasizes exploring AI techniques to develop intelligent featu... Show more

 • Promoted

Consumer Research Associate (Remote)

GL IncSanta Cruz, California
$15.00 hourly
Remote
Part-time +1

Product Testers are wanted to work from home nationwide in the US to fulfill upcoming contracts with national and international companies.We guarantee 15-25 hours per week with an hourly pay of bet... Show more

 • Promoted

Automation & Controls Engineering Manager

Heron PowerScotts Valley, CA, United States
Full-time

Automation & Controls Engineering Manager.At Heron Power, we are building the future of scalable power electronics.Our mission is to enable high-volume, capital-efficient manufacturing of advanced ... Show more

 • Promoted

Clinical Supervisor (Masters Required)

ACESSanta Cruz, CA, United States
Full-time

Clinical Supervisor (Masters Required).ACES is driven to elevate the standards in the treatment of autism.Our team of Applied Behavior Analysis (ABA) clinicians is deeply committed to helping child... Show more

 • Promoted

AI Enterprise Architect - Remote & AI Solutions Leader

Workday, Inc.Pleasanton, CA, United States
Remote
Full-time

A leading AI platform provider is seeking an experienced Enterprise Architect - AI to drive the adoption of AI-enabled solutions.This role involves acting as a trusted advisor to prospects, leverag... Show more

 • Promoted

Senior Manager, Embodied AI

General MotorsSunnyvale, CA, United States
Full-time

Senior Engineering Manager AI/ML Engineering.We are seeking an experienced and technically strong Senior Engineering Manager AI/ML Engineering for the Data Foundations organization of Embodied AI... Show more

 • Promoted

Lead AI Architect for Agentic Systems (Hybrid / Remote)

Zoom Video CommunicationsSan Jose, CA, United States
Remote
Full-time

A leading video communication platform is seeking a Principal AI Architect to design intelligent agent systems and oversee their development.The role involves establishing technical standards for A... Show more

 • Promoted

Director Engineering - AI/ML

Stanford Health CarePalo Alto, CA, United States
Full-time

We are looking for a rare combination of hands-on technical leader and strategic thinker to own the engineering vision for chatEHR as we mature from a pilot into a tier 1 clinical platform.This is ... Show more

 • Promoted

MLOps and ML Engineer

APTASKSan Ramon, CA, United States
Full-time

Job TitleThe client is a leading digital transformation consultancy and engineering services company that specializes in driving digital transformation for Fortune 1000 enterprises.The company prov... Show more

 • Promoted

Remote Mathematics Expert (PhD) - $100+ per hour

Turing Scotts Valley, California
Remote
Full-time
Quick Apply

Remote contract for PhDs in Mathematics, Statistics, or related fields.Work on cutting-edge projects with top AI labs while earning $100+/hour, fully remote, with flexible weekly hours.Help fine-tu... Show more

 • Promoted

Remote Product Tester - Flexible Studies

BuzzTestersSanta Cruz, CA
Remote
Full-time

Earn up to $400/week, depending on the number and type of studies you complete.Share honest feedback on products and services from independent brands.Assignments include short online surveys (10–20... Show more

 • Promoted

Senior Software Engineer, Flight Simulator

Joby AviationSanta Cruz, CA, United States
Full-time

Joby Flight Training Simulator EngineerImagine a piloted air taxi that takes off vertically, then quietly carries you and your fellow passengers over the congested city streets below, enabling you ... Show more

 • Promoted

Principal EP Mapping Specialist - CAS

Medtronic PlcSanta Cruz, CA, United States
Full-time

Careers that change lives start here.Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life.Our 95,000 employees work across more th... Show more

 • Promoted

Software Engineer, Machine Learning

Meta PlatformsNewark, CA, United States
Full-time

Talented Engineers WantedMeta is seeking talented engineers to join our teams in building cutting-edge products that connect billions of people around the world.As a member of our team, you will ha... Show more

 • Promoted

Senior Manager, Gen AI Software Engineering

Thermo FisherPleasanton, CA, United States
Full-time

Senior Manager, Generative AI Software EngineeringAs part of the Thermo Fisher Scientific team, you'll discover meaningful work that makes a positive impact on a global scale.Join our colleagues in... Show more

 • Promoted

Principal Consultant & AI/ML

Argyle InfotechSan Jose, CA, United States
Full-time

Principal Consultant AI/ML & LLM Engineering.We are looking for a Principal Consultant with strong expertise in AI/ML engineering and Large Language Model (LLM) application development.This role i... Show more

 • Promoted

Remote Principal AI Architect -- Hybrid AI Systems

LenovoSan Jose, CA, United States
Remote
Full-time

A global technology powerhouse is seeking a highly experienced Principal AI Architect to lead the design and implementation of advanced AI systems.This role requires deep knowledge of AI models, sc... Show more

 • Promoted

Principal Machine Learning Engineer

WorkdayPleasanton, CA, United States
Full-time

Your Work Days Are Brighter HereWe're obsessed with making hard work pay off, for our people, our customers, and the world around us.As a Fortune 500 company and a leading AI platform for managing ... Show more

 • Promoted

Software Engineering Manager II, AI/ML GenAI, Google Cloud AI

GoogleSunnyvale, CA, United States
Full-time

Software Engineering Manager II, AI/ML GenAI, Google Cloud AI.Advanced experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders; deep expertise in dom... Show more

 • Promoted

Principal Developer Experience Engineer - Remote

PayNearMeFremont, CA, United States
Remote
Full-time

Company DescriptionPayNearMe develops technology to facilitate the end-to-end customer payment experience, making it easy for businesses to accept, disburse and manage payments.Our modern and relia... Show more