CyberArk
Principal Engineer Software
Office - USA - CA - Headquarters, USA · Principal · Full-time
No sponsorship$147k-$238kDetected 25 days ago
PythonBashNode.jsKubernetesAnsibleLinuxNginxPrometheusGrafanaMachine LearningCiscoBGP/OSPFLoad BalancingNetwork MonitoringCollaborationProblem Solving
About the role
- Sr Staff Data Center & OpenShift Operations Engineer
- The Senior Data Center Operations Engineer is responsible for the bedrock of our high-availability infrastructure.
- This role bridges the gap between physical hardware and the Red Hat OpenShift Container Platform (OCP).
Responsibilities
- High-Availability (HA) Infrastructure: Monitor and maintain data center systems with a focus on "Zero Single Point of Failure" (ZSPoF) architecture for OpenShift control planes and worker nodes.
- Cluster Reliability Engineering: Implement and manage OpenShift 4.x clusters across multiple power and cooling zones to ensure 99.99% uptime.
- Disaster Recovery & Business Continuity: Design, test, and execute automated failover strategies and backup/restore procedures using tools like OADP (Velero) and Red Hat ACM.
- Automated Maintenance: Perform routine maintenance and upgrades using GitOps (ArgoCD) and the Machine Config Operator to ensure zero-downtime node evacuations and patching.
- Efficiency & Capacity Architecture: Optimize rack density for high-performance GPU clusters while managing thermal loads and power distribution (PDU) to prevent circuit-trip outages.
- Hardware Lifecycle: Perform precision physical installation and replacement of critical components (CPUs, GPUs, NVMe storage) in a live production environment without impacting cluster quorum.
- Monitor and maintain data center systems with a focus on "Zero Single Point of Failure" (ZSPoF) architecture for OpenShift control planes and worker nodes.
- Implement and manage OpenShift 4.x clusters across multiple power and cooling zones to ensure 99.99% uptime.
Requirements
- Platform Expertise: 5+ years of experience specifically operating Red Hat OpenShift (OCP) in a production environment.
- Hardware Fluency: Deep experience racking/stacking and cabling high-density GPU systems (e.g., NVIDIA DGX or similar) and specialized AI/ML hardware.
- Infrastructure as Code (IaC): Advanced proficiency in Ansible or Pulumi for automating bare-metal provisioning and cluster configuration.
- Virtualization & Storage: Experience with vSphere or KVM, and persistent storage solutions like OpenShift Data Foundation (ODF) or Ceph.
- Education: Bachelor's degree in Computer Science, IT, or equivalent experience.
- Scripting: Strong Python and Bash skills for developing custom health-check scripts and API integrations.
- Linux Mastery: Expert-level CoreOS and RHEL administration, including kernel tuning and systemd management.
- Networking: Solid understanding of BGP, VLAN tagging, LACP, and Load Balancing (F5/NGINX) essential for cluster ingress.
- Tooling: Familiarity with DCIM tools (Netbox) and monitoring stacks ( ELK/Lok..etci).
- Physical Requirements
- Lifting: Ability to lift and move equipment up to 50 pounds (e.g., high-density 2U/4U servers).
- Environment: Comfortable working in high-decibel, climate-controlled data center aisles.
- Dexterity: Capable of standing, walking, and performing precision cabling in tight rack spaces for extended periods.
- Travel: May require occasional travel to remote data center sites or edge locations.
Compensation
- The compensation offered for this position will depend on qualifications, experience, and work location.
- For candidates who receive an offer at the posted level, the starting base salary (for non-sales roles) or base salary + commission target (for sales/com-missioned roles) is expected to be the annual range listed below.
- The offered compensation may also include restricted stock units and a bonus.
- $147,000.00 - $237,500.00/yr
- We're trailblazers that dream big, take risks, and challenge cybersecurity's status quo.
- Competitive salary commensurate with high-availability expertise.
Benefits
- Comprehensive health, dental, and vision insurance.
- 401(k) retirement plan with company match.
- Maintain accurate documentation and integrate hardware health metrics (IPMI/SNMP) into Prometheus/Grafana for proactive alerting.
- Strong Python and Bash skills for developing custom health-check scripts and API integrations.
- A description of our employee benefits may be found here.
Company info
- In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
- We believe collaboration thrives in person. That's why most of our teams work from the office full time, with flexibility when it's needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.
- Your mission is to ensure 99.99% availability by architecting resilient physical layouts and automating the deployment, scaling, and self-healing capabilities of our production clusters.
- At Palo Alto Networks®, we're united by a shared mission-to protect our digital way of life.
- We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking.
- Here, everyone has a voice, and every idea counts.
- If you're ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you're in the right place.
- In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry.
- This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion.
- We weave AI into the fabric of everything we do and use it to augment the impact every individual can have.
- If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
- We believe collaboration thrives in person.
- That's why most of our teams work from the office full time, with flexibility when it's needed.
- This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.
- Disruption, Collaboration, Execution, Integrity, and Inclusion.
- Position Overview
Equal opportunity
- If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.
- equal opportunity employer.
- All your information will be kept confidential according to EEO guidelines.
Visa & Work Authorization
- Is role eligible for Immigration Sponsorship?
- Please note that we will not sponsor applicants for work visas for this position.
Apply directly at CyberArk →Create a free account for alerts like thisView CyberArk immigration profile
This listing is sourced directly from CyberArk's careers page and normalized into a canonical job model.