Talent.com
MY HR
Site Reliability Engineer.MY HR • Berkeley, CA, United States
No longer accepting applications
Site Reliability Engineer.

Site Reliability Engineer.

MY HR • Berkeley, CA, United States
21 days ago
Job type
  • Permanent
  • Quick Apply
Job description

Site Reliability Engineer.

Our client is inviting applications for the position of Site Reliability Engineer. Client's mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE Office of Science programs. Client provides critical HPC and data systems and support for client's 11,000 plus users researching energy, physics, materials science, and chemistry and other DOE mission areas. As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure client is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work ensures that client's computational power remains an uninterrupted catalyst supporting fundamental scientific research related to energy.


ESSENTIAL DUTIES & RESPONSIBILITIES*
Describe the duties/functions essential to performing the job.
Works an onsite 5-day weekly schedule consisting of Owl (midnight 8 am) shifts to monitor the client HPC Facility.
Review and respond to alerts from computer systems, storage, network, and other data center/facility-related systems by triaging or calling the appropriate on-call staff.
Create appropriate solutions to improve processes, prevent issue recurrence, and automate responses to all routine service conditions.
Identify issues and propose solutions that will improve monitoring capabilities or provide better automation for triage.
Possess expertise in ServiceNow and its usage to develop and implement customized service management solutions.
Respond to alerts from multiple systems to ensure that data collection continues 24/7, providing real-time information for diagnoses.
Develop and maintain tools within the monitoring pipeline in collaboration with the Operations Team.
Create new software to provide alerts and notifications from HPC system APIs into the monitoring pipeline.
Builds and maintains application/tool configurations to ensure software runs reliably as data and user demands grow.
Collaborate with other groups at client to ensure that communication and workflows are clearly understood.
Work closely with other client groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periods.
Perform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure to ensure peak operational efficiency.
Provide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents so that workflows and protocols can be appropriately tracked by others.
Work on and resolve problems of diverse scope where data analysis requires the evaluation of identifiable factors.
Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
Work on and resolve complex issues where the analysis of situations or data requires an in-depth evaluation of variable factors.


POSITION REQUIREMENTS
Experience in or willingness to work within a 24/7 onsite team environment to support large-scale data centers or critical installations.
Experience on Linux shell and working in a command-line (e.g. SSH) environment.
Experience with developing tools using various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
Motivated, self-starter who can learn technologies that improve data center management in areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative cooling, and power utilization.
Experience with network security: configuring/maintaining ACLs, knowledge of firewalls

Experience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.
Good to Have : Practical experience in developing and deploying Agentic AI or autonomous automation tools to streamline technical tasks.
Experience with ServiceNow implementation is a plus
Familiarity with ITSM best practices and an understanding of how to align service lifecycles with business goals is preferred.
Knowledge, Skills & Abilities
Strong hands-on knowledge of the Linux shell and working in a command-line (e.g. SSH) environment.
Strong hands-on knowledge on various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
Knowledge of and ability to work on large data communications networks/ Network Protocols and IT infrastructure supporting highly available systems and applications.
Strong communication skills and ability to work effectively across multiple technical teams.
Good to Have : Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to automate decision-making, optimize complex workflows, and enhance proactive system monitoring.

About MY HR:
MY HR is an award-winning, woman and minority-owned firm based in Atlanta. We specialize in providing full-service professional HR services, and are proud to be an equal opportunity employer. With a commitment to excellence and a focus on diversity, we strive to help businesses of all sizes achieve their human resources goals.

Follow us for more info:

www.myhrmgmt.com
linkedin.com/company/my-hr/
facebook.com/myhrsupplier
instagram.com/myhrmanagement/

MY HR is an award-winning Full-Service Professional Human Resources Consulting firm offering Staff Augmentation, Project and SOW staffing, Permanent Placement, Recruitment Process Outsourcing (RPO), Payroll Services, and full range HR Services including compliance, training, and workforce development. With our personal touch, we help small to mid-sized companies as well as Fortune 500 companies grow and strengthen in the HR area by providing customized HR solutions.

Check out our website: myhrmgmt.com

Create a job alert for this search

Site Reliability Engineer. • Berkeley, CA, United States

Similar jobs

Site Lead

FlexcarRichmond, CA, United States
Full-time

Title: Site LeadLocation: Richmond , CA Classification - Exempt - Full timeCompensation: $90,000 - $124,000full Benefits Day One A Career with Real Impact Flexcar is completely reimagining car own... Show more

 • Promoted

Manager, Mechanical Design Engineering

Form EnergyBerkeley, CA, United States
Full-time

Are you ready to build America's energy future? Form Energy is an American manufacturing and energy technology company.We're revolutionizing energy storage with cost-effective, multi-day technology... Show more

 • Promoted

Site Manager | COL

HomeRise (formerly Community Housing Partnership)San Francisco, CA, United States
Full-time

Site Manager | Jazzie Collins Apartments.Starting Salary: $74,700 annually.HomeRise believes that home has the power to stabilize a person's life.Built on a simple - but powerful idea called suppor... Show more

 • Promoted

Staff Software Engineer, Reliability

General MotorsSan Francisco, CA, United States
Full-time

Staff Software EngineerThe AV platform team develops the first layers of software on the GM Autonomous Vehicles from working with hardware to moving large amounts of data up the software stack.With... Show more

 • Promoted

Maintenance & Engineering Manager

CHANDON CaliforniaYountville, CA, United States
Full-time

Manage the business's physical assets (buildings, infrastructure and equipment) at the Yountville site.Responsible for the successful facilities and maintenance preventive and corrective work progr... Show more

 • Promoted

Remote Senior SRE & FEDRAMP Reliability Engineer

TenableSan Francisco, CA, United States
Remote
Full-time

A leading cybersecurity company is seeking a Staff Software / Site Reliability Engineer to work on a cloud-based vulnerability management platform.The role involves ensuring product reliability, ad... Show more

 • Promoted

Senior Site Manager

Burbank HousingNapa, CA, United States
Full-time +1

Heritage House - Valle Verde - Napa, CA 94559.Hourly Position Type Full Time Job Shift Day.Heritage House is a 66-unit property located in the southern region of Napa, CA.Of the 66 units, 41 are fu... Show more

 • Promoted

Maintenance Area Lead (Sonoma Region)

MidPen Housing CorporationSonoma, CA, United States
Full-time

Maintenance Area Lead (Sonoma Region).At MidPen, we build communities that change lives.Since 1970, we have been committed to our mission: to provide safe, affordable housing of high quality to tho... Show more

 • Promoted

Director, Site Reliability Engineering

Stellar Development FoundationSan Francisco, CA, United States
Full-time

Director Of Site Reliability Engineering.Interested in working on cutting-edge blockchain technology and creating equitable access to the global financial system? Since 2014, the mission-driven tea... Show more

 • Promoted

CommVault Systems Engineer

Contact Government ServicesSan Francisco, CA, United States
Full-time

Commvault Systems Engineer (Data Protection / Backup).Employment Type: Full-Time, Experienced.CGS is seeking an experienced Commvault Data Protection Engineer with extensive knowledge and experienc... Show more

 • Promoted

Remote Consumer Feedback Participant

GL IncNapa, California
$15.00 hourly
Remote
Part-time +1

Product Testers are wanted to work from home nationwide in the US to fulfill upcoming contracts with national and international companies.We guarantee 15-25 hours per week with an hourly pay of bet... Show more

 • Promoted

Project Site Leader - PHX, LA, San Francisco, Seattle or Irvine

Vertiv HoldingsSan Francisco, CA, United States
Full-time

The Project Site Leader will provide world class jobsite leadership for large, long-duration, high-profile projects for Vertiv power and/or thermal equipment.The Site Leader is the primary Vertiv S... Show more

 • Promoted

Technical Operations Lead

Indee LabsBerkeley, CA, United States
Full-time

Indee Labs is a biotechnology startup backed by SOSV, Y Combinator, Social Capital and Founders Fund.The team at Indee Labs develops and markets Hydropore for more effective engineered cell therapi... Show more

 • Promoted

Senior Site Reliability Engineer

HiveSan Francisco, CA, United States
Full-time

Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content, and is trusted by hundreds of the world's largest and most innovative organizations.The company... Show more

 • Promoted

Remote Senior Site Reliability Engineer (SRE) - Zetachain

Blockchain WorksSan Francisco, CA, United States
Remote
Full-time

Site Reliability Engineer to join our team and run critical infrastructure for our blockchain and web applications.You'll learn to deploy and maintain a fleet of RPC and validator nodes for multipl... Show more

 • Promoted

Remote Site Reliability Engineer -- Cloud & Automation

PatreonSan Francisco, CA, United States
Remote
Full-time

A leading platform for creators is seeking a Site Reliability Engineer to improve the performance, reliability, and cost efficiency of their growing infrastructure.This remote role will involve sig... Show more

 • Promoted

Site Manager

FortrexSan Francisco, CA, United States
Full-time

As a Site Manager II specializing in food processing sanitation, you will lead a dedicated team to ensure rigorous cleanliness and hygiene standards, safeguarding product quality and consumer safet... Show more

 • Promoted

Lab Operations, Facilities and Equipment Manager

nEye SystemsEmeryville, CA, United States
Full-time

Lab Operations, Facilities & Equipment Manager.Eye's MEMS-based silicon photonics optical circuit switches (OCS) eliminate critical bottlenecks in AI processing by enabling direct optical connectio... Show more

 • Promoted

Site Reliability Engineer -- Remote, Flexible PTO

ConductorOne Inc.San Francisco, CA, United States
Remote
Full-time

A leading identity security platform is seeking a Site Reliability Engineer in San Francisco, California.You will design and operate reliable infrastructure across cloud environments while automati... Show more

 • Promoted

Project Manager

ShimmickNapa, CA, United States
Full-time

Shimmick is seeking a dynamic Project Manager to support our Napa Creek Flood Protection Project in Napa, CA.This $30M+ project will provide necessary infrastructure to expand flood protection to t... Show more