nscaleoperationsukltd
Operational Data & Observability Engineer
US
Sponsorship not specified$145k-$180kDetected 1 day ago
PythonGoBashAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringPlatform EngineeringData EngineeringIncident ResponseCommunication
About the role
- In this role, you'll help ensure our infrastructure and applications remain reliable, scalable, and performant by providing engineering teams with actionable operational insights.
Responsibilities
- Design & Build Observability Solutions
- Design and implement enterprise observability strategies across infrastructure, services, and applications.
- Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility.
- Build and maintain centralized logging and log analysis pipelines.
- Implement distributed tracing to improve visibility across microservices and complex application workflows.
- Establish performance baselines and develop anomaly detection strategies.
- Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems.
- Design and manage operational data pipelines that support monitoring and analytics.
- Develop APIs and integrations that enable operational data consumption across teams.
- Participate in an on-call rotation and support incident response activities.
Requirements
- 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering.
- Hands-on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent.
- Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions.
- Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent.
- Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM).
- Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies.
Nice to have
- Experience supporting microservices-based architectures.
- Expertise across multiple observability platforms.
- Experience with incident management, root cause analysis, and post-incident reviews.
- Infrastructure as Code experience using Terraform, Ansible, or similar tools.
- Familiarity with eBPF or low-level Linux performance monitoring.
- Understanding of security monitoring, audit logging, and compliance requirements.
- What Success Looks Like
- Success in this role will be measured by your ability to:
Skills
- Actual compensation may vary based on job-related factors such as skill set, experience, education, and location.
Compensation
- $145,000 - $180,000 USD
- The range below reflects the base salary for the position.
Benefits
- In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs.
- Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Apply directly at nscaleoperationsukltd →Create a free account for alerts like thisView nscaleoperationsukltd immigration profile
This listing is sourced directly from nscaleoperationsukltd's careers page and normalized into a canonical job model.