ISHIR

ISHIR

Senior DevOps Engineer

United States · Senior

Sponsorship not specifiedDetected 111 days ago
PythonJavaBashGitAWSDockerKubernetesTerraformHelmCI/CDLinuxNginxPrometheusDevOpsSite Reliability EngineeringJiraCommunication

About the role

  • Assist with service instrumentation across the core observability pillars-tracing, logging, and metrics-with hands-on OpenTelemetry usage (collectors/SDKs) and related telemetry tooling.
  • Contribute to and improve documentation: runbooks, FAQs, onboarding guides, and standard operating procedures.

Responsibilities

  • Operate and improve platform tools so product teams can ship reliably triaging tickets, fix build issues, and handling routine service requests (access, secrets, environment setup).
  • Maintain and extend self-service workflows (templates, golden paths) by updating docs, examples, and guardrails under guidance from senior engineers.
  • Perform day-to-day Kubernetes operations: deploy/update Helm charts, manage namespaces, diagnose rollout issues, and follow runbooks for incident response.
  • Support CI/CD pipelines (e.g., GitLab CI): keep pipelines green, add/adjust jobs, implement basic quality gates, and help teams adopt safer deploy strategies (blue/green, canary).
  • Monitor and operate the observability stack using Prometheus, Alert manager, and Thanos; maintain alert rules, dashboards, and SLO/SLA indicators; help reduce alert noise and improve signal quality.
  • API integration experience (Java, Python, or Go) to build small internal tools or glue code.

Nice to have

  • writing/modifying jobs, artifacts, environments, and basic deployment strategies.
  • Scripting ability in Bash or Python (Go a plus) to automate repetitive tasks and improve runbooks.
  • Familiarity with AWS fundamentals (e.g., IAM, EC2/EKS, S3, CloudWatch/CloudTrail, Parameter Store/Secrets Manager).
  • Practical understanding of monitoring/observability (dashboards, logs, alerts) and how to use them for triage and remediation, including Prometheus/Alertmanager/Thanos and OpenTelemetry basics.
  • Comfortable working from tickets (Jira/ServiceNow), following change-management practices, and communicating clearly with stakeholders.
  • Terraform experience for infrastructure as code (managing AWS, Kubernetes add-ons, and platform components).
  • Deeper Linux fundamentals and container runtime basics for effective debugging and performance tuning.
  • Exposure to insurance/financial services environments, including awareness of compliance and operational controls.

Skills

  • Skilled in automation, scripting (Bash/Python), and observability tools like Prometheus and OpenTelemetry.
  • Focused on reliable deployments, issue resolution, and maintaining efficient platform operations.
  • Experience operating Kubernetes (or similar) and core ecosystem tools (Helm, Docker, Ingress NGINX, Argo Rollouts basics).

Company info

  • WE'RE LOOKING FOR SOMEONE WHO HAS: 8+ years in a platform/SRE/DevOps or infrastructure role, with a strong bias toward automation and support.
  • We are not just consultants, we are partners in our clients' success, assisting them with re(gaining) competitive edge by identifying opportunities for differentiation, industry disruption, scalable innovation, and go-to-market strategies that deliver successful outcomes.
  • At ISHIR, we help bold businesses accelerate innovation through Talent, Speed-to-Market, and AI.

This listing is sourced directly from ISHIR's careers page and normalized into a canonical job model.