Bitdeer Technologies Group
Sr. SRE Platform Software Engineer
San Jose, CA · Senior
Sponsorship not specifiedDetected 15 days ago
PythonDistributed SystemsSQLElasticsearchAWSGCPKubernetesHelmCI/CDPrometheusDatadogSite Reliability EngineeringLLMsLogisticsProcurementWriting
About the role
- You don't only ship code; you ship a service that other squads, cloud-service teams, and tenants depend on.
Responsibilities
- You will own 1-2 of these: Collection & Storage: collection-agent, customer-sdk-gateway, metrics-store, logs-store, traces-store, profiles-store, analytics-lake, enrichment-service, collection-monitor.
- SQL & Time-Series Data: Ability to read a Prometheus query plan, build a recording-rule strategy, and write SQL that joins per-tenant telemetry against analytics-lake tables.
- Experience writing and maintaining your own comprehensive tests.
- Technical Writing Fluency: Ability to author clear design docs that align with existing platform architecture, create runbooks optimized for 3 AM on-call responses, and write intent-driven PR descriptions.
Requirements
- 7+ years of production software engineering experience, including 2 or more years operating what you built (real on-call experience, not just shipping code).
- Ability to evaluate CRDT vs.
- Experience at production scale with Prometheus, VictoriaMetrics, Mimir, Thanos, Loki, Elasticsearch, Tempo, Jaeger, or OpenTelemetry.
- Must have built or substantively contributed to the ingest, query, or storage paths of these systems.
- Hands-on experience with Argo, Flux, Helm, Kustomize, Cosign signing, signed-bundle promotion, and blast-radius-aware rollouts.
- Experience executing end-to-end mTLS bootstrap with certificate rotation.
- Hands-on experience with HashiCorp Vault or cloud KMS (AWS KMS / GCP KMS).
- Hands-on experience with BMC, IPMI, and Redfish at OEM scale (Supermicro, Dell, HPE, Lenovo).
- Familiarity with Kubernetes GPU Operator, Slurm controller, or Ray GCS.
- GitOps & CI/CD: Hands-on experience with Argo, Flux, Helm, Kustomize, Cosign signing, signed-bundle promotion, and blast-radius-aware rollouts.
Nice to have
- Production-depth mastery of at least one systems-grade language-Go (preferred), Rust, or Java.
- Proficiency in Python for tooling and SDK work.
- Programming Languages: Production-depth mastery of at least one systems-grade language-Go (preferred), Rust, or Java.
- (GPU / AI-Infra Context) Experience in at least one of the following areas is a strong plus: NVIDIA Internals: Deep understanding of DCGM and NVIDIA driver internals, including XID semantics and MIG / vGPU partitioning.
Skills
- About Bitdeer Technologies Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
- alert-engine-framework, alert-correlation, slo-framework, default M-series alert rules.
- prediction-engine-framework and built-in predictors (GPU, Link, Disk, XPA, Straggler, SDC, Stranded GPU).
- cicd-pipeline, gitops-sync, plugin-registry, sre-image-registry.
- Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series alert rules.
Benefits
- topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s, Slurm, Ray, Volcano, Kueue, and KubeRay.
- Experience with InfiniBand or RoCE fabrics, including subnet managers, partitioning, optical health, and NCCL collective tracing.
Apply directly at Bitdeer Technologies Group →Create a free account for alerts like thisView Bitdeer Technologies Group immigration profile
This listing is sourced directly from Bitdeer Technologies Group's careers page and normalized into a canonical job model.