nscaleoperationsukltd

nscaleoperationsukltd

Operational Data & Observability Engineer

US

Sponsorship not specified$145k-$180kDetected 1 day ago
PythonGoBashAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringPlatform EngineeringData EngineeringIncident ResponseCommunication

About the role

  • In this role, you'll help ensure our infrastructure and applications remain reliable, scalable, and performant by providing engineering teams with actionable operational insights.

Responsibilities

  • Design & Build Observability Solutions
  • Design and implement enterprise observability strategies across infrastructure, services, and applications.
  • Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility.
  • Build and maintain centralized logging and log analysis pipelines.
  • Implement distributed tracing to improve visibility across microservices and complex application workflows.
  • Establish performance baselines and develop anomaly detection strategies.
  • Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems.
  • Design and manage operational data pipelines that support monitoring and analytics.
  • Develop APIs and integrations that enable operational data consumption across teams.
  • Participate in an on-call rotation and support incident response activities.

Requirements

  • 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering.
  • Hands-on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent.
  • Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions.
  • Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent.
  • Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM).
  • Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies.

Nice to have

  • Experience supporting microservices-based architectures.
  • Expertise across multiple observability platforms.
  • Experience with incident management, root cause analysis, and post-incident reviews.
  • Infrastructure as Code experience using Terraform, Ansible, or similar tools.
  • Familiarity with eBPF or low-level Linux performance monitoring.
  • Understanding of security monitoring, audit logging, and compliance requirements.
  • What Success Looks Like
  • Success in this role will be measured by your ability to:

Skills

  • Actual compensation may vary based on job-related factors such as skill set, experience, education, and location.

Compensation

  • $145,000 - $180,000 USD
  • The range below reflects the base salary for the position.

Benefits

  • In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs.
  • Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

This listing is sourced directly from nscaleoperationsukltd's careers page and normalized into a canonical job model.