Chainlink Labs
Senior Site Reliability Engineer, Observability
United States · Senior
Sponsorship not specifiedDetected 54 days ago
PythonJavaGoC++RubySwiftDistributed SystemsCode ReviewGitAWSKubernetesTerraformGitHub ActionsPrometheusGrafanaDevOpsSite Reliability EngineeringCommunication
About the role
- The Chainlink stack provides the essential data, interoperability, compliance, and privacy standards needed to power advanced blockchain use cases for institutional tokenized assets, lending, payments, stablecoins, and more.
- Since inventing decentralized oracle networks, Chainlink has enabled tens of trillions in transaction value and now secures the vast majority of DeFi.
- Chainlink leverages a novel fee model where offchain and onchain revenue from enterprise adoption is converted to LINK tokens and stored in a strategic Chainlink Reserve.
Responsibilities
- Support multiple telemetry types, like metrics, logs and traces.
- Define and support modern governance in observability and problems at scale.
- Lead the design and deployment of monitoring/observability services to detect and alert the team of needed action.
- Create processes around alert response operations and support the team to ensure the reliable delivery of oracle data.
- Make recommendations to ensure sufficient metrics are collected to create alerts with every new feature release.
Requirements
- 7+ years of relevant professional experience.
- Experience programming in C, C++, Java, Python, Go, Perl, or Ruby
- Experience with monitoring and logging.
- You know how to export metrics using Prometheus, have built a Grafana dashboard or two, and have experience with a centralized logging solution like an ELK Stack, Splunk or Grafana Stack.
- Experience with distributed systems and container orchestration.
- You have maintained or even built Kubernetes clusters before and feel comfortable deploying completely new services on them
- You can give and receive constructive feedback, and you do not shy away from planning meetings and code reviews
- Experience running any infrastructure in the blockchain/web3 space
- Ability to scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity
- Experience working remotely in a distributed team
Skills
- AWS; Terraform/Terragrunt; Kubernetes, Calico and ArgoCD; Prometheus and Grafana; GitHub Actions; Packer
- Build and orchestrate Modern OTEL-based Observability Platform
- Ensure reliability, security, and performance exceed our defined SLAs
- Ingest, aggregate, transform, and utilize data from a multitude of sources in our real time data pipeline.
- Oversee the availability, performance, and supportability of our observability infrastructure.
- Champion reliability and security by taking the time to do your work right the first time
Apply directly at Chainlink Labs →Create a free account for alerts like thisView Chainlink Labs immigration profile
This listing is sourced directly from Chainlink Labs's careers page and normalized into a canonical job model.