Tekmetric

Tekmetric

Site Reliability Engineer

United States · Junior

Sponsorship not specifiedDetected 8 days ago
PythonJavaGoBashAWSCloud PlatformsDockerTerraformCI/CDPrometheusGrafanaDevOpsSite Reliability EngineeringCybersecurityIncident ResponseLeadershipCommunicationCollaborationProblem SolvingMentoring

About the role

  • We move fast, stay curious, and take full ownership of our results - no excuses, no finger-pointing.
  • If you thrive in ambiguity, take initiative, and view honest feedback as fuel for growth, you'll feel right at home here.
  • We're direct but respectful, ambitious yet grounded, and collaborative at every level.

Responsibilities

  • Design and implement scalable infrastructure: Architect and maintain reliable, scalable, and secure cloud infrastructure that supports positive user experiences and measurable business growth.
  • Monitor and optimize system performance: Develop and maintain monitoring, alerting, and incident response practices to ensure system reliability and performance at scale.
  • Automate everything: Create automated pipelines for deployment, testing, and infrastructure management to improve speed, consistency, and reliability across the organization.
  • Ensure high availability and disaster recovery: Implement and manage solutions for backup, disaster recovery, and failover processes to ensure business continuity.
  • But we're not just building software.
  • We're building a movement.
  • We're empowering repair shops to rise above the daily grind, create meaningful connections with their customers, and lead the industry forward - one interaction at a time.
  • Come build with us.
  • At Tekmetric, we're building a culture where winning matters - not for ego, but because when our customers win, we win together.
  • Everyone leads through impact and is encouraged to speak up, share ideas, and challenge assumptions (even your manager's).

Requirements

  • 3+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related field, with deep knowledge of cloud environments (preferably AWS or GCP.).
  • Hands-on experience with AWS (or similar cloud providers) and infrastructure as code (Terraform, etc.).
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
  • Proficiency in scripting languages like Python, Bash, or similar.
  • Experience with designing and optimizing Continuous Integration and Continuous Deployment (CI/CD) pipelines.
  • Strong communication skills and ability to work cross-functionally, solving complex technical challenges in a collaborative manner.
  • Ability to troubleshoot and resolve critical issues in high-pressure environments, maintaining composure and professionalism.
  • Cloud Infrastructure: Hands-on experience with AWS (or similar cloud providers) and infrastructure as code (Terraform, etc.).
  • Monitoring and Logging: Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).

Nice to have

  • Experience with Infrastructure as Code tools like Terraform.
  • Familiarity with monitoring tools like Prometheus, Grafana, or the ELK stack.
  • Exposure to compliance and security best practices in cloud environments.
  • Experience coding in one or multiple programming languages such as Go, Java, Javascript.

Skills

  • What You'll Bring
  • Strong experience in automation tools
  • Expertise in working with containerized environments like Docker and orchestration tools such as Kubernetes.

Company info

  • Tekmetric is the all-in-one, cloud-based platform helping auto repair shops run smarter, grow faster, and serve customers better.
  • From running a shop, to securing payments to engaging customers, our platform simplifies operations so shop owners can focus on what really matters: delivering exceptional service, earning trust, and growing sustainably.

This listing is sourced directly from Tekmetric's careers page and normalized into a canonical job model.