Chainlink Labs

Chainlink Labs

Senior Site Reliability Engineer, Observability

United States · Senior

Sponsorship not specifiedDetected 54 days ago
PythonJavaGoC++RubySwiftDistributed SystemsCode ReviewGitAWSKubernetesTerraformGitHub ActionsPrometheusGrafanaDevOpsSite Reliability EngineeringCommunication

About the role

  • The Chainlink stack provides the essential data, interoperability, compliance, and privacy standards needed to power advanced blockchain use cases for institutional tokenized assets, lending, payments, stablecoins, and more.
  • Since inventing decentralized oracle networks, Chainlink has enabled tens of trillions in transaction value and now secures the vast majority of DeFi.
  • Chainlink leverages a novel fee model where offchain and onchain revenue from enterprise adoption is converted to LINK tokens and stored in a strategic Chainlink Reserve.

Responsibilities

  • Support multiple telemetry types, like metrics, logs and traces.
  • Define and support modern governance in observability and problems at scale.
  • Lead the design and deployment of monitoring/observability services to detect and alert the team of needed action.
  • Create processes around alert response operations and support the team to ensure the reliable delivery of oracle data.
  • Make recommendations to ensure sufficient metrics are collected to create alerts with every new feature release.

Requirements

  • 7+ years of relevant professional experience.
  • Experience programming in C, C++, Java, Python, Go, Perl, or Ruby
  • Experience with monitoring and logging.
  • You know how to export metrics using Prometheus, have built a Grafana dashboard or two, and have experience with a centralized logging solution like an ELK Stack, Splunk or Grafana Stack.
  • Experience with distributed systems and container orchestration.
  • You have maintained or even built Kubernetes clusters before and feel comfortable deploying completely new services on them
  • You can give and receive constructive feedback, and you do not shy away from planning meetings and code reviews
  • Experience running any infrastructure in the blockchain/web3 space
  • Ability to scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity
  • Experience working remotely in a distributed team

Skills

  • AWS; Terraform/Terragrunt; Kubernetes, Calico and ArgoCD; Prometheus and Grafana; GitHub Actions; Packer
  • Build and orchestrate Modern OTEL-based Observability Platform
  • Ensure reliability, security, and performance exceed our defined SLAs
  • Ingest, aggregate, transform, and utilize data from a multitude of sources in our real time data pipeline.
  • Oversee the availability, performance, and supportability of our observability infrastructure.
  • Champion reliability and security by taking the time to do your work right the first time

This listing is sourced directly from Chainlink Labs's careers page and normalized into a canonical job model.