Praxy
← Back to jobs

Senior Site Reliability Engineer – Compute Platforms

Five9

Location
Remote · United States (Remote), US
Posted
1mo ago
EngineeringSoftware SystemsOperations Execution

About this role

<div class="content-intro"><p><img src="https://www.five9.com/sites/default/files/2025-02/five9-logo.svg" alt="" width="100" style="max-width: 100%;"></p> <p>Join us in bringing joy to customer experience.  Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide.   </p> <p>Living our values everyday results in our team-first culture and enables us to innovate, grow, and thrive while enjoying the journey together. We celebrate diversity and foster an inclusive environment, empowering our employees to be their authentic selves. </p></div><p>We are seeking a highly experienced <strong>Senior Site Reliability Engineer – Compute Platforms </strong>to design, implement, and support <strong>Kubernetes on baremetal and hypervisor platforms</strong> in a private cloud environment. This role is responsible for the <strong>architecture, design, and standardization</strong> of enterprise compute and hypervisor environments spanning bare metal infrastructure, operating systems, hypervisors, private cloud orchestration, and Kubernetes using Infrastructure-as-Code and GitOps practices.</p> <p>This is a deeply technical role requiring <strong>expert-level understanding of compute hardware management, Kubernetes, OpenStack, hypervisors and extensive working knowledge on Linux Operating systems. </strong>You will also collaborate with platform and SRE teams to maintain secure, performant, and multi-tenant-isolated services that serve high-throughput, mission-critical applications. </p> <p><strong>Key Responsibilities</strong></p> <ul> <li>Lead the <strong>architecture and design</strong> of enterprise compute and hypervisor platform solutions across hardware, OS, virtualization, cloud orchestration, and container orchestration layers</li> <li>Define standards and automation frameworks for <strong>bare metal provisioning and lifecycle management</strong></li> <li>Design and implement <strong>Bare Metal as a Service (BMaaS)</strong> capabilities for scalable infrastructure consumption</li> <li>Architect and design <strong>Kubernetes platforms</strong> on bare metal with QoS and Affinity (ArgoCD)</li> <li>Architect and validate automated deployments of operating systems and hypervisors including <strong>Ubuntu</strong> and <strong>Harvester</strong></li> <li>Design and maintain <strong>PXE-based provisioning environments</strong> leveraging <strong>Redfish APIs</strong> for large-scale server deployments</li> <li>Develop <strong>Infrastructure-as-Code</strong> using Ansible, Terraform, Helm and Git, with Python/Bash automation.</li> <li>Implement <strong>CI/CD pipelines</strong> for infrastructure updates, patching, upgrades, testing, and rollback.</li> <li>Design automated workflows for server build, firmware lifecycle management, patching, and hardware validation</li> <li>Evaluate and standardize enterprise hardware platforms to meet performance, scalability, and reliability requirements</li> <li>Produce detailed <strong>high-level and low-level design documentation</strong>, build guides, and operational handoff materials</li> <li>Perform deep troubleshooting across <strong>storage, Kubernetes, hypervisors, networking, and Linux systems</strong></li> <li>Partner with operations, network, storage, and platform teams to ensure designs are supportable and production-ready</li> <li>Participate in <strong>on-call escalation support</strong> for complex platform-related issues</li> <li>Collaborate globally on <strong>change management</strong>, documentation, and operational best practices</li> </ul> <p><strong>Minimum Qualifications</strong></p> <ul> <li><strong>6</strong><strong>+ years of experience</strong> in infrastructure engineering, platform engineering, or DevOps with a strong focus on Compute system design</li> <li>Proven experience designing and automating <strong>bare metal compute environments</strong> at scale</li> <li>Strong hands-on experience with <strong>PXE boot, network-based OS provisioning, and automated server imaging</strong></li> <li>Experience implementing or supporting <strong>Bare Metal as a Service (BMaaS)</strong> platforms</li> <li>Practical experience using <strong>Redfish APIs</strong> for hardware provisioning, power management, and remote lifecycle operations</li> <li>Deep expertise with <strong>Ubuntu Linux</strong> in enterprise environments</li> <li>Strong Hands-on experience with KVM hypervisors (Suse Harvester, OpenStack).</li> <li>Experience designing and deploying <strong>production-grade Kubernetes clusters</strong></li> <li>Strong background with <strong>enterprise compute hardware platforms</strong>, including Cisco UCS, Dell PowerEdge, Supermicro systems & HPE</li> <li>Proficiency with <strong>Infrastructure as Code tools</strong> (e.g., Terraform, Ansible, or similar)</li> <li>Experience building or supporting <strong>CI/CD pipelines</strong> for infrastructure and platform automation</li> <li>Strong scripting skills in <strong>Python, Bash, or similar languages</strong></li> <li>Demonstrated ability to produce clear, structured <strong>technical design documentation</strong></li> <li>Excellent written and verbal <strong>communication skills</strong></li> <li>Bachelor’s degree in computer science or equivalent professional experience</li> </ul> <p><strong>Preferred Qualifications</strong> </p> <ul> <li>OpenStack, Ubuntu KVM administration.</li> <li>BareMetal as a Service (PXE, Redfish).</li> <li>Kubernetes on BareMetal</li> <li><strong>CIS/NIST</strong> security and infrastructure lifecycle management.</li> <li>ITIL Foundation/advanced certifications in support of ITSM standard methodology.</li> <li>Background in telco, edge cloud, or large enterprise environments.</li> <li>Ubuntu Certifications, CNCF Certified Kubernetes Administrator (CKA), Certified Kubernetes Security Specialist (CKS)</li> <li>Master’s degree in computer science, IT, Engineering, or a related field preferred; equivalent experience and relevant industry certifications will also be considered</li> </ul> <p><strong>What You’ll Get</strong> </p> <ul> <li>A collaborative team that’s deeply invested in infrastructure excellence.</li> <li>Complex technical challenges that require creative, scalable solutions.</li> <li>The opportunity to shape a next-generation private cloud platform-built reliability</li> <li>Access to the latest tools, frameworks, and upstream project developments</li> </ul> <p><strong>Skills and Attributes:</strong> </p> <ul> <li><strong>Analytical Thinking & Problem Solving:</strong> Demonstrated ability to translate complex, cross-domain requirements into scalable and resilient clou

If this role is no longer available, it will disappear from Praxy automatically.