Talent.com

Software testing Jobs in Sunnyvale, CA

Create a job alert for this search

Software testing • sunnyvale ca

Last updated: 8 hours ago

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, IncSan Jose, California, United States
Full-time

At AMD, we believe technology has the power to solve the world’s most important challenges.From advancing healthcare and scientific discovery to powering AI and the technologies people rely on ever... Show more

 • New!

Software Engineer

CYNET SYSTEMSSan Jose, CA, US
Full-time

Job Overview: Pay Range: $66hr - $71hr Requirement/Must Have: Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience).Strong skills in SQL, Python, Java, or ... Show more

 • Promoted

Online Product Testing - $25-$45 per hour

OCPACampbell, California, US
$25.00 hourly
Part-time +1

Product Testers are wanted to work from home in the UK to fulfill upcoming contracts with local and international companies.We guarantee 15-25 hours per week with an hourly pay of between £18/hr.Th... Show more

 • Promoted

Online Product Testing - $25-$45 per hour

OCPASanta Clara, California, US
$25.00 hourly
Part-time +1

Product Testers are wanted to work from home in the UK to fulfill upcoming contracts with local and international companies.We guarantee 15-25 hours per week with an hourly pay of between £18/hr.Th... Show more

 • Promoted

Hardware Validation/ Testing Engineer

VailexaSanta Clara, CA, US
Full-time

At Vailexa, we’re not just hiring — we’re building thinkers, creators, and future leaders.We believe in giving people the space to grow, the freedom to think, and the opportunity to create real imp... Show more

 • Promoted

QA Testing Director /Manager/SME

Tanisha SystemsSanta Clara, CA, US
Full-time

QA Testing Director /Manager/SME REMOTE / Work from Home FTE/Fulltime Salary – Market Start: Immediate Domain: Retail / E-commerce Domain We are seeking a Director - Testing / QA for defining and l... Show more

 • Promoted

Strong Java AI, API Testing

Programmers.ioSunnyvale, CA, United States
Full-time
Quick Apply

Detailed JD:</b></p> <ul type="disc"> <li><b>API Testing:</b> Deep expertise in <b>Java-based REST API</b> automation testing.BDD Framew... Show more

Project Manager (Testing)

IntersourcesSunnyvale, CA, United States
Full-time

Location: Sunnyvale, CA Duration: Long term contract.Looking for a Testing PM that has experience managing testing teams for large cross-functional projects interacting with multiple systems.Prefer... Show more

ROS Software Developer

Raas Info Solutions Pvt LtdSunnyvale, CA, US
Full-time

Job Description: We are seeking a Senior Software Engineer contractor to develop robotics platform software on NVIDIA Jetson (Linux/ROS2) for a manipulation system integrating custom grippers, sens... Show more

 • Promoted

Software Engineer - Software Engineer

TrinusSanta Clara, CA, United States
Full-time
Quick Apply

We are seeking an Software engineer (Contract), specialized in AI/ML applications to independently drive the development, evaluation, deployment and end-to-end lifecycle management of our AI-powere... Show more

Oracle EPM Testing & Defect Lead

Lockheed MartinSunnyvale, CA, United States
Full-time

Oracle Epm Testing & Defect Lead | Lockheed Martin.Join Lockheed Martin's digital transformation journey as we accelerate the OneLM Mission-Driven Transformation through our 1LMX program.This strat... Show more

Vehicle Testing Specialist

Continuum Resource NetworkSunnyvale, CA, US
Full-time
Quick Apply

We are helping our client hire .In this role, you will execute test plans, support a wide range of testing functions, and report on testing outcomes to provide crucial information for the soft... Show more

Remote Customer Service Representative – Product Testing

GLOCPAMilpitas, California
$15.00 hourly
Remote
Part-time +1

Product Testers are wanted to work from home nationwide in the US to fulfill upcoming contracts with national and international companies.We guarantee 15-25 hours per week with an hourly pay of bet... Show more

 • Promoted

Team Manager - Testing Operations

TSMGMountain View, CA, United States
Full-time

The purpose of the role is to drive operational performance in sourcing with a focus on lead generation and validation metrics.The manager is responsible for managing daily operational targets to e... Show more

Staff Software DevQA Engineer(End-to-End Testing)

FortinetSanta Clara, CA, United States
Full-time

Collaborate closely with developers to identify, reproduce, and resolve defects.Maintain test environments and continuous integration pipelines.Document test plans, test cases, and results clearly ... Show more

Software Engineer V

Pinnacle Technical ResourcesCupertino, CA, US
Full-time

Position: Animation Software Engineer/Graphics Engineer V Location: Cupertino, California - Remote Duration: Contract Job ID: 170702 Job Description: Job Description: Keynote Animation Software Eng... Show more

 • Promoted

Hardware Testing Engineer

Simple SolutionsSanta Clara, CA, us
Full-time

Santa Clara, Ca (Onsite daily).Hours of support we currently need:.One day on the weekends as well.Must be will to work OT and weekends as necessary.We need candidates who are strong in Linux, netw... Show more

Packaging Testing Technician

Eurofins USA BioPharma ServicesSan Jose, California, United States
$18.00 hourly
Full-time

Receiving Customer samples and recording information in Shipping and receiving Log.Conduct Transit and Package Integrity Testing as Per ASTM Standards , ISTA standard or per Customer Protocol.Add S... Show more

Online Product Testing - $25-$45 per hour

OCPALos Altos, California, US
$25.00 hourly
Part-time +1

Product Testers are wanted to work from home in the UK to fulfill upcoming contracts with local and international companies.We guarantee 15-25 hours per week with an hourly pay of between £18/hr.Th... Show more

 • Promoted
People also ask
Principal Software Developer – AI/ML Performance Validation & Systems Testing

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, IncSan Jose, California, United States
8 hours ago
Job type
  • Full-time
Job description


ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.




THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID




Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.