Talent.com
Software Engineer - Power Management, Hardware Health
Software Engineer - Power Management, Hardware HealthOpenai • San Francisco, California, United States
Software Engineer - Power Management, Hardware Health

Software Engineer - Power Management, Hardware Health

Openai • San Francisco, California, United States
[job_card.variable_days_ago]
[job_preview.job_type]
  • [job_card.full_time]
[job_card.job_description]

About the Team

OpenAI’s Hardware Health team is dedicated to ensuring the optimal performance and reliability of our custom-built hyperscale supercomputers. We focus on maximizing supercomputing capacity for research and ensuring that our researchers are minimally impacted by hardware faults. This team is critical in maintaining the infrastructure that supports cutting-edge AI research at OpenAI.

The Hardware Health team operates within the broader Platform organization, which is incubated inside OpenAI’s Research team. Our work is on the front lines of innovation, supporting the engineering and research required to train large-scale AI models of unprecedented capability.

About the Role

As a Software Engineer on the Hardware Health team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity.  This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability.

Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale.

In this role, you will :

Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations.

Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability.

Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults.

Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power management solutions.

Drive the development of power throttling mechanisms at the IT system level to dynamically adjust power usage based on workload demands and infrastructure limitations.

Collaborate with hardware design teams to integrate system-level power control requirements into IT hardware design, ensuring seamless coordination between software-driven power management and hardware capabilities.

You might thrive in this role if you have :

7+ years of software engineering experience with a focus on solving large-scale, system-level challenges.

Strong proficiency in Python and familiarity with automation and scripting tools (e.g., shell scripting).

Experience with distributed systems to efficiently aggregate and analyze streaming data.

Knowledge of electrical engineering concepts including digital signal processing, power systems, Fast Fourier Transforms, or related areas.

Experience in system-level investigations and development of automated solutions to address power management, fault detection, and remediation.

Strong analytical skills and the ability to dig into noisy data (experience with SQL, PromQL, Pandas, etc.).

Comfort working with both hardware and software teams to solve multidisciplinary problems.

Bonus points if you have :

Deep expertise with the power characteristics of synchronous workloads (as seen in supercomputing or model training environments).

Knowledge of power control requirements in IT hardware design, with the ability to drive cross-functional collaboration to integrate power management features into hardware systems effectively.

Working knowledge of control system fundamentals and how physical systems respond to control strategies.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or any other legally protected status.

OpenAI Affirmative Action and Equal Employment Opportunity Policy Statement

For US Based Candidates : Pursuant to the San Francisco Fair Chance Ordinance, we will consider qualified applicants with arrest and conviction records.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this  link .

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

[job_alerts.create_a_job]

Hardware Engineer • San Francisco, California, United States

[internal_linking.similar_jobs]
Hardware Systems & Test Engineer

Hardware Systems & Test Engineer

Shyld AI • San Francisco Bay Area, United States
[job_card.full_time]
Our platform combines sensing, embedded systems, AI, and hardware deployed in real clinical environments.Hardware Systems & Test Engineer. This role sits at the intersection of.QA, system integratio...[show_more]
[last_updated.last_updated_variable_days] • [promoted]
Staff Software Engineer, ML Performance & Systems

Staff Software Engineer, ML Performance & Systems

Fal • San Francisco, California, United States
[job_card.full_time]
Help fal maintain its frontier position on model performance for generative media models.Design and implement novel approaches to model serving architecture on top of our in-house inference engine,...[show_more]
[last_updated.last_updated_variable_days] • [promoted]
Senior Software Engineer (GTM)

Senior Software Engineer (GTM)

Toma • San Francisco, California, United States
[job_card.full_time]
We're building the AI platform for underserved industries.LLM usage has seen a meteoric rise in the past year, but there is still a significant gap between agentic innovation and its use in the rea...[show_more]
[last_updated.last_updated_variable_days] • [promoted]
Sr. Software Engineer

Sr. Software Engineer

Mem Protocol • San Francisco, California, United States
[job_card.full_time]
If you think about how to design networked systems for optimal uptime, efficiency, and latency then this may be a good fit for you. Designing and implementing efficient data storage solutions.Creati...[show_more]
[last_updated.last_updated_30] • [promoted]
Software Engineer - Systems

Software Engineer - Systems

Specter • San Francisco, California, United States
[job_card.full_time]
Specter is creating a software-defined “control plane” for the physical world.We are starting with protecting American businesses by granting them ubiquitous perception over their physical assets.T...[show_more]
[last_updated.last_updated_30] • [promoted]
Software Engineer, Fleet Hardware Health

Software Engineer, Fleet Hardware Health

OpenAI • San Francisco, CA, United States
[job_card.full_time]
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, an...[show_more]
[last_updated.last_updated_30] • [promoted]
Local Contract Physical Therapist - $41-48 per hour

Local Contract Physical Therapist - $41-48 per hour

Medical Solutions Allied • San Pablo, CA, United States
[job_card.full_time]
Medical Solutions Allied is seeking a local contract Physical Therapist for a local contract job in San Pablo, California. Job Description & Requirements.We’re seeking talented healthcare profession...[show_more]
[last_updated.last_updated_30] • [promoted]
Software Engineer

Software Engineer

Modern Treasury • San Francisco, California, United States
[job_card.full_time]
This position can be based out of San Francisco, New York, or remote.We're looking for Full-Stack Software Engineers who want to help businesses of any size move money worldwide.You will contribute...[show_more]
[last_updated.last_updated_30] • [promoted]
Physical Therapist - Home Health

Physical Therapist - Home Health

21ST CENTURY HOME HEALTH SERVICES • Richmond, CA, United States
[job_card.full_time]
Physical Therapist - Home Health Full Time.At 21st Century Home Health Services (HHS), we are committed to treating every patient with the same empathy, compassion and understanding that we would s...[show_more]
[last_updated.last_updated_30] • [promoted]
Software Engineer, Core Engine

Software Engineer, Core Engine

Eventual • San Francisco, California, United States
[job_card.full_time]
Eventual is a data platform that helps data scientists and engineers build data applications across ETL, analytics and ML / AI. OUR PRODUCT IS OPEN-SOURCE AND USED AT ENTERPRISE SCALE.Our distributed ...[show_more]
[last_updated.last_updated_30] • [promoted]
Software Engineer, Monetization

Software Engineer, Monetization

Postman • San Francisco, California, United States
[job_card.full_time]
Postman is the world’s leading API platform, used by more than 40 million developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals acro...[show_more]
[last_updated.last_updated_30] • [promoted]
Senior Firmware EngineerSoftware Engineering • Berkeley, CA • Full time • On-site

Senior Firmware EngineerSoftware Engineering • Berkeley, CA • Full time • On-site

Form Energy • Berkeley, CA, United States
[job_card.full_time]
Are you ready to build America's energy future? Form Energy is an American manufacturing and energy technology company.We're revolutionizing energy storage with cost-effective, multi-day technology...[show_more]
[last_updated.last_updated_30] • [promoted]
Ground Software & Systems Manager - Mission Operations (0346U), Space Sciences Laboratory - 83546

Ground Software & Systems Manager - Mission Operations (0346U), Space Sciences Laboratory - 83546

InsideHigherEd • Berkeley, California, United States
[job_card.full_time]
Ground Software & Systems Manager - Mission Operations (0346U), Space Sciences Laboratory - 83546.At the University of California, Berkeley, we are dedicated to fostering a community where everyone...[show_more]
[last_updated.last_updated_variable_days] • [promoted]
Physical Therapist

Physical Therapist

TotalMed • Pinole, CA, United States
[job_card.full_time]
As a Physical Therapist, your responsibility involves assessing and treating patients.You will develop personalized plans tailored to achieve specific physical therapy goals.Collaboration with the ...[show_more]
[last_updated.last_updated_30] • [promoted]
Software Engineer - Solutions - Primefocus Health

Software Engineer - Solutions - Primefocus Health

Lg Nova • San Francisco, CA, United States
[job_card.full_time]
Software Engineer - Solutions - Primefocus Health.Primefocus Health is a seed‑stage healthtech startup headquartered in San Francisco, launched as a venture spin‑out from LG Electronics’ North Amer...[show_more]
[last_updated.last_updated_30] • [promoted]
Physical Therapist

Physical Therapist

Concentra Careers • Richmond, CA, United States
[job_card.permanent]
Are you ready to take your career to new heights? At Concentra, you will be a vital member of our patient care team and play a crucial role in providing exceptional care to our patients.Our mission...[show_more]
[last_updated.last_updated_30] • [promoted]
Software Engineer, Platform Engineering

Software Engineer, Platform Engineering

Asana • San Francisco, California, United States
[job_card.full_time]
We are dedicated to ensuring proactive elimination of entire classes of security risk by engineering the core libraries, platforms and frameworks that provide secure guardrails for all Asanas.We ar...[show_more]
[last_updated.last_updated_30] • [promoted]
Staff Software Engineer

Staff Software Engineer

Bio-Rad Laboratories • Hercules, CA, United States
[job_card.full_time]
This role is both technical and collaborative.You will work closely with cross-functional teams including systems engineers, mechanical designers, assay development scientists, and quality engineer...[show_more]
[last_updated.last_updated_30] • [promoted]