Saviynt

Saviynt

Principal Site Reliability Engineer, Google Cloud

Atlanta · Principal

Sponsorship not specified$240k-$250kDetected 20 days ago
PythonGoDistributed SystemsBackend DevelopmentGitMySQLAWSGCPAzureCloud PlatformsKubernetesCI/CDPrometheusGrafanaDatadogSite Reliability EngineeringAPI DevelopmentRESTKafkaIncident ResponseAccount ManagementCommunication

About the role

  • Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI.
  • Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions.
  • For more information, please visit www.saviynt.com.

Responsibilities

  • In this pivotal role, you will be instrumental in designing, building, and maintaining the shared infrastructure services and platforms that our product and application teams will depend on
  • You will focus on creating reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on their core business logic and deliver features faster in a multi-cloud environment
  • Design and build core platform components and shared infrastructure services that other development teams will integrate with and leverage to deploy and operate their applications
  • Architect, implement, and manage highly available and scalable Kubernetes platforms as a service for internal consumers
  • Develop robust, internal-facing tools and automation for infrastructure provisioning and management primarily using Go (Golang)
  • Architect and optimize foundational solutions within Cloud environments (AWS, Azure, etc.), focusing on creating reusable patterns and modules for other teams
  • Design and implement shared Event-Driven Architecture components and messaging platforms using technologies like Kafka or Google Pub/Sub that product teams can easily utilize
  • Develop and maintain robust CI/CD pipelines (e.g., GitLab CI and ArgoCD) as a service, providing standardized and automated deployment workflows for various development teams
  • Design and build resilient Distributed Systems components that serve as building blocks for other applications, focusing on reliability, fault tolerance, and performance
  • Manage and optimize our shared infrastructure across Multi-Region Cloud Environments, ensuring that platform services are globally available and performant for all consumers

Requirements

  • 1+ years of experience as a Principal SRE with a strong focus on building tools and services for other engineers
  • Solid understanding and practical experience with CI/CD pipeline tools (especially GitLab CI) and experience establishing automated delivery processes for other teams
  • Proficiency in establishing and utilizing comprehensive Observability and Monitoring platforms (e.g., Prometheus, Grafana, ELK stack, Datadog) for shared infrastructure
  • Strong experience with RESTful API design principles and building well-documented, consumable APIs
  • Knowledge of Service Mesh concepts and practical experience with solutions like Istio in a platform context
  • Hands-on experience with Relational Databases (e.g., MySQL, PostgresSQL), ideally in managing them as a service

Nice to have

  • Extensive hands-on experience with at least one major Cloud Provider (GCP is a must)

Compensation

  • If required for this role, you will: - Complete security & privacy literacy and awareness training during onboarding and annually thereafter - Review (initially and annually thereafter), understand, and adhere to Information Security/Privac

This listing is sourced directly from Saviynt's careers page and normalized into a canonical job model.