TP Link USA Corp

TP Link USA Corp

Cloud Operations Engineer - Infrastructure

Irvine, California, United States

No sponsorshipDetected 48 days ago
PythonGoAWSAzureCloud PlatformsKubernetesTerraformHelmCI/CDLinuxSite Reliability EngineeringPlatform EngineeringIncident ResponseCommunicationCollaborationProblem Solving

About the role

  • The company is committed to delivering innovative products that enhance people's lives through faster, more reliable connectivity.
  • We believe technology changes the world for the better!
  • Embracing professionalism, innovation, excellence, and simplicity, we aim to assist our clients in achieving remarkable global performance and enable consumers to enjoy a seamless, effortless lifestyle.

Responsibilities

  • Design, build, and maintain reliable, scalable, and secure cloud-native infrastructure platforms supporting large-scale production workloads.
  • Operate and optimize multi-account AWS environments, ensuring infrastructure is secure, repeatable, and auditable through Infrastructure as Code tools such as Terraform.
  • Manage production Kubernetes clusters, including provisioning, upgrades, autoscaling, networking, observability, capacity planning, and day-to-day operations.
  • Build and operate Kubernetes ecosystem components such as CRDs, Helm, HPA, Cluster Autoscaler, CoreDNS, and Cluster API.
  • Manage and enhance Istio service mesh capabilities, including traffic routing, service discovery, resilience, security, and service-to-service communication.
  • Participate in a scheduled on-call rotation to support production cloud infrastructure and Kubernetes platforms.
  • Drive automation for infrastructure provisioning, configuration management, CI/CD pipelines, observability, and operational workflows using Terraform, Go, Python, or similar technologies.
  • Collaborate with application engineering, architecture, security, and platform teams to improve infrastructure reliability, scalability, and operational efficiency.
  • We believe that diversity fuels innovation, collaboration, and drives our entrepreneurial spirit.
  • Beyond compliance, we strive to create a supportive and growth-oriented workplace for everyone.

Requirements

  • Bachelor's degree or above in Computer Science, Software Engineering, Information Technology, or a related field.
  • 2+ years of hands-on experience in cloud infrastructure, Kubernetes operations, platform engineering, SRE, or related areas.
  • Strong knowledge of AWS services, including EKS, IAM, VPC, EC2, S3, and related networking and security capabilities.
  • Hands-on experience operating Kubernetes in production environments, including cluster architecture, workload orchestration, networking, autoscaling, and troubleshooting.
  • Familiarity with Kubernetes ecosystem tools such as CRDs, Helm, Cluster API, HPA, Cluster Autoscaler, and CoreDNS.
  • Experience with GitOps tools such as FluxCD or ArgoCD.
  • Experience with CI/CD pipelines and infrastructure automation using Terraform, Go, Python, or similar tools.
  • Strong problem-solving skills and ability to diagnose and resolve complex infrastructure issues in distributed systems.

Nice to have

  • Experience with NVIDIA device plugins, GPU scheduling, or GPU workload operations in Kubernetes environments.
  • Experience with additional public cloud platforms such as Azure or Alibaba Cloud.
  • Kubernetes certifications such as CKA, CKAD, or CKS are a plus.

Company info

  • Headquartered in the United States, TP-Link Systems Inc. is a global provider of reliable networking devices and smart home products, consistently ranked as the world's top provider of Wi-Fi devices.
  • With a commitment to excellence, TP-Link serves customers in over 170 countries and continues to grow its global footprint.
  • At TP-Link Systems Inc, we are committed to crafting dependable, high-performance products to connect users worldwide with the wonders of technology.
  • Operate and improve GitOps-based deployment workflows using tools such as FluxCD or ArgoCD.
  • Define and improve reliability practices, including SLOs, Error Budgets, monitoring, alerting, incident response, and post-mortems.

Visa & Work Authorization

  • Please, no third-party agency inquiries, and we are unable to offer visa sponsorships at this time

This listing is sourced directly from TP Link USA Corp's careers page and normalized into a canonical job model.