Talent.com
Bitdeer Technologies Group
K8 Site Reliability SMEBitdeer Technologies Group • San Jose, California, United States
No longer accepting applications
K8 Site Reliability SME

K8 Site Reliability SME

Bitdeer Technologies Group • San Jose, California, United States
30+ days ago
Job type
  • Full-time
Job description

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Position Overview

You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own

  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.

Feed the AIOps substrate

  • The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.

What success looks like in year 1

  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.

Job Requirement:

  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Create a job alert for this search

K8 Site Reliability SME • San Jose, California, United States

Similar jobs

Site Reliability Engineer

Foxconn Industrial InternetSan Jose, CA, United States
Full-time +1

Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industr... Show more

 • Promoted

Sales Associate (Lead Generator FT or PT)

Bellows Plumbing, Heating, Cooling & ElectricalSanta Cruz, CA, United States
Full-time +1

HVAC, Generators, and Water Treatment Services Lead Generator.Bellows Heating Plumbing Cooling & Electrical is currently seeking a highly motivated individual to generate leads for HVAC, generators... Show more

 • Promoted

Field Associate - Property Showings & Evaluations

DoorsteadSanta Cruz, CA, United States
Full-time

Must be based in or around the Santa Cruz area.The Field Associate role encompasses two primary functions: Showings and Property Evaluations.This is a contracted hourly position that is suited for ... Show more

 • Promoted

Sr. Principal Instrument Replacement Lead

WatersMilpitas, CA, United States
Full-time

Principal, Lifecycle Program Lead.Our Enterprise Transformation team is on a mission to improve business outcomes and employee experiences by driving step-change improvements in critical enterprise... Show more

 • Promoted

Special Needs Caregiver

Alegre Home CareSanta Cruz, CA, United States
Part-time

Great Shifts Available Now Full Time 32 to 40 Hours Per Week.Home Care and Facility Settings: Days, PMs, NOCs and 12 Hours Shifts.Work with Adults and Children with Developmental Disabilities.We se... Show more

 • Promoted

Pre-Sales Lead TECHM-JOB-24881

Keylent, Inc.San Jose, CA, United States
Full-time

Manages relationships with Sales Field team members.Engages with vendor Customer Service team members.Engages with vendor and Cisco Supply Chain and Systems teams.Develops strong overview of Demons... Show more

 • Promoted

Environmental Simulation Lab & Reliability Department Manager / SME

EurofinsSanta Clara, CA, United States
Full-time

Environmental Simulation Lab & Reliability Department Manager / SME.Eurofins Scientific is an international life sciences company, providing a unique range of analytical testing services to clients... Show more

 • Promoted

Senior/Staff Site Reliability Engineer

Gatik AISanta Clara, CA, United States
Full-time

Senior/Staff Site Reliability Engineer.Gatik, the leader in autonomous middle-mile logistics, is revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solutio... Show more

 • Promoted

Site Reliability Engineer (Splunk, Prometheus, Grafana) Hybrid

KaavSunnyvale, CA, United States
Full-time

Site Reliability Engineer Role.This is a Site Reliability Engineer Role for Sam's Cash Application team.Role and responsibilities include:.Production tickets handling and troubleshooting: Requires ... Show more

 • Promoted

Package Reliability EngineerSanta Clara, CA

Celestial AISanta Clara, CA, United States
Full-time

Celestial AI's Photonic Fabric is the next-generation interconnect technology that delivers a tenfold increase in performance and energy efficiency compared to competing solutions.This innovation e... Show more

 • Promoted

Special Needs Caregiver

Addus HomecareAptos, CA, United States
Full-time

Addus Homecare - - Responsibilities: Assist with personal care; Provide occasional house cleaning, laundry, and assist with meal preparation; Transport client to appointments and daily errands. Show more

 • Promoted

Companion

Hope ServicesSanta Cruz, CA, United States
Full-time

Companion Position at Hope Services.Are you a person who enjoys helping others? Are you currently seeking fulfillment in your professional life?.Hope Services is Silicon Valley's leading provider o... Show more

 • Promoted

Remote Site Reliability Engineer -- Aviation Tech

DevOps projectsMountain View, CA, United States
Remote
Full-time

A leading innovator in aviation technology is seeking a Site Reliability Engineer to join their remote team.The role includes managing and improving infrastructure, providing tooling support, and p... Show more

 • Promoted

Essential Program Caregiver

LifespanSanta Cruz, CA, United States
Part-time

Lifespan - Santa Cruz, CA 95062.Hourly Position Type Part Time Job Shift Any.This is a new program that Lifespan is creating.Geared towards clients that need minimum assistance with ADLS.Services t... Show more

 • Promoted

Site Reliability Engineer Remote

PayNearMeSan Jose, CA, United States
Remote
Full-time

As our Site Reliability Engineer you will design build and maintain the systems and infrastructure that power our applications ensuring their reliability scalability and performance.You will bring ... Show more

 • Promoted

Product Infrastructure Engineer - Site Reliability

ZyphraPalo Alto, CA, United States
Full-time

Infrastructure Engineer - Site ReliabilityAs an Infrastructure Engineer - Site Reliability, you'll be responsible for designing and maintaining the systems that keep Zyphra's infrastructure robust,... Show more

 • Promoted

Motivating Caregivers needed! Help us make a difference! - Santa Cruz

ameriCARESanta Cruz, CA, United States
Full-time

The Caregiver role is a crucial position for those who rely on others for basic daily care, such as bathing, eating, and personal hygiene.Caregivers typically work with elderly or disabled individu... Show more

 • Promoted

Site Reliability Engineer

Syntricate TechnologiesSan Jose, CA, United States
Full-time

Extensive experience working with Linux flavors like RHEL/CentOS OS, shells, filesystems and utilities.Knowledge of distributed computing and experience working with container orchestration framewo... Show more

 • Promoted

Caregivers for Clients with Developmental Disabilities

Cambrian HomecareSanta Cruz, CA, United States
Full-time

Cambrian Homecare Care Provider Opportunity.Cambrian Homecare, LLC is hiring individuals with the passion to provide one-on-one care to children or adults with developmental disabilities in their h... Show more

 • Promoted

Senior Technical Program Manager, Reliability

RivianPalo Alto, CA, United States
Full-time +1

Senior Technical Program Manager.Rivian is on a mission to keep the world adventurous forever.This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous soul... Show more