TechniPros
DevOps Engineer-Dallas, TX
Dallas, Texas, USA · Contract
Sponsorship not specifiedDetected 34 days ago
PythonGoBashDistributed SystemsAWSGCPAzureCloud PlatformsDockerKubernetesCI/CDLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringStakeholder ManagementBudgetingCommunicationCollaborationProblem Solving
About the role
- We are seeking an experienced Site Reliability Engineer (SRE) / DevOps Engineer with a strong background in Incident Management, Change Control, Error Budgeting, Remediation, and Production Operations.
- The ideal candidate will be responsible for ensuring the reliability, scalability, performance, and operational excellence of cloud-native platforms and distributed systems.
- This role requires deep expertise in cloud infrastructure, automation, observability, incident response, and operational governance.
Responsibilities
- Manage and improve platform reliability, availability, and performance across production environments.
- Lead and participate in incident management, root cause analysis, remediation planning, and post-incident reviews.
- Drive change control processes and ensure operational governance standards are followed.
- Monitor and manage error budgets while implementing reliability improvements.
- Design, build, and maintain scalable cloud infrastructure and automation frameworks.
- Deploy and manage containerized applications using Kubernetes and Docker.
- Develop and maintain CI/CD pipelines to support efficient software delivery.
- Implement Infrastructure as Code (IaC) solutions for automated provisioning and configuration management.
- Collaborate with development, infrastructure, security, and business teams to ensure platform stability.
Nice to have
- Experience implementing reliability engineering best practices and SRE methodologies.
- Experience supporting large-scale enterprise production environments.
- Familiarity with high-availability and disaster recovery architectures.
- Experience automating operational workflows and infrastructure management.
- Knowledge of security best practices within cloud environments.
- Experience working in Agile and DevOps-driven organizations.
Skills
- 7+ years of experience in Site Reliability Engineering (SRE), DevOps, Cloud Infrastructure, or Production Operations.
- Excellent verbal and written communication skills.
- Strong analytical and problem-solving capabilities.
- Ability to perform effectively during high-severity production incidents.
- Strong stakeholder management and cross-functional collaboration skills.
- Ability to prioritize multiple tasks in a fast-paced environment.
- Proactive mindset focused on continuous improvement and operational excellence.
- Microsoft Azure Amazon Web Services (AWS)
Apply directly at TechniPros →Create a free account for alerts like thisView TechniPros immigration profile
This listing is sourced directly from TechniPros's careers page and normalized into a canonical job model.