Bitdeer Technologies Group

Bitdeer Technologies Group

Sr. SRE Platform Software Engineer

San Jose, CA · Senior

Sponsorship not specifiedDetected 15 days ago
PythonDistributed SystemsSQLElasticsearchAWSGCPKubernetesHelmCI/CDPrometheusDatadogSite Reliability EngineeringLLMsLogisticsProcurementWriting

About the role

  • You don't only ship code; you ship a service that other squads, cloud-service teams, and tenants depend on.

Responsibilities

  • You will own 1-2 of these: Collection & Storage: collection-agent, customer-sdk-gateway, metrics-store, logs-store, traces-store, profiles-store, analytics-lake, enrichment-service, collection-monitor.
  • SQL & Time-Series Data: Ability to read a Prometheus query plan, build a recording-rule strategy, and write SQL that joins per-tenant telemetry against analytics-lake tables.
  • Experience writing and maintaining your own comprehensive tests.
  • Technical Writing Fluency: Ability to author clear design docs that align with existing platform architecture, create runbooks optimized for 3 AM on-call responses, and write intent-driven PR descriptions.

Requirements

  • 7+ years of production software engineering experience, including 2 or more years operating what you built (real on-call experience, not just shipping code).
  • Ability to evaluate CRDT vs.
  • Experience at production scale with Prometheus, VictoriaMetrics, Mimir, Thanos, Loki, Elasticsearch, Tempo, Jaeger, or OpenTelemetry.
  • Must have built or substantively contributed to the ingest, query, or storage paths of these systems.
  • Hands-on experience with Argo, Flux, Helm, Kustomize, Cosign signing, signed-bundle promotion, and blast-radius-aware rollouts.
  • Experience executing end-to-end mTLS bootstrap with certificate rotation.
  • Hands-on experience with HashiCorp Vault or cloud KMS (AWS KMS / GCP KMS).
  • Hands-on experience with BMC, IPMI, and Redfish at OEM scale (Supermicro, Dell, HPE, Lenovo).
  • Familiarity with Kubernetes GPU Operator, Slurm controller, or Ray GCS.
  • GitOps & CI/CD: Hands-on experience with Argo, Flux, Helm, Kustomize, Cosign signing, signed-bundle promotion, and blast-radius-aware rollouts.

Nice to have

  • Production-depth mastery of at least one systems-grade language-Go (preferred), Rust, or Java.
  • Proficiency in Python for tooling and SDK work.
  • Programming Languages: Production-depth mastery of at least one systems-grade language-Go (preferred), Rust, or Java.
  • (GPU / AI-Infra Context) Experience in at least one of the following areas is a strong plus: NVIDIA Internals: Deep understanding of DCGM and NVIDIA driver internals, including XID semantics and MIG / vGPU partitioning.

Skills

  • About Bitdeer Technologies Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
  • alert-engine-framework, alert-correlation, slo-framework, default M-series alert rules.
  • prediction-engine-framework and built-in predictors (GPU, Link, Disk, XPA, Straggler, SDC, Stranded GPU).
  • cicd-pipeline, gitops-sync, plugin-registry, sre-image-registry.
  • Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series alert rules.

Benefits

  • topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s, Slurm, Ray, Volcano, Kueue, and KubeRay.
  • Experience with InfiniBand or RoCE fabrics, including subnet managers, partitioning, optical health, and NCCL collective tracing.

This listing is sourced directly from Bitdeer Technologies Group's careers page and normalized into a canonical job model.