Cl
Staff Site Reliability Engineer (Kubernetes & AWS)
Austin, TX & Remote
•
$170,000 - $220,000 /yr
•
Full-time
Remote
About the Role
CloudScale is seeking a Staff SRE to own platform reliability, observability, and disaster recovery across thousands of multi-region Kubernetes clusters. You will write code to eliminate toil.
Key Responsibilities
-
Define and enforce Service Level Objectives (SLOs) and Error Budgets across engineering.
-
Lead chaos engineering experiments and incident retrospectives.
-
Build automated self-healing tooling in Go or Python.
Requirements & Qualifications
-
Expertise in managing production Kubernetes clusters (EKS/GKE) with GitOps (ArgoCD).
-
Extensive infrastructure-as-code mastery with Terraform.
-
Deep knowledge of Linux networking, TCP/IP, eBPF, and container runtimes.
-
Experience managing high-cardinality monitoring stacks (Prometheus, Cortex, Datadog).
Perks & Compensation Benefits
-
Top-tier 401(k) with 6% company match.
-
Full health coverage (medical, dental, vision).
-
Generous remote work budget and annual company retreats.
Job Overview
Date Posted
Oct 01, 2026
Application Deadline
Oct 21, 2026
Experience Level
Lead
Job Type
Full-time
Workplace Model
Remote / Hybrid
Views
296 views
Cl
CloudScale Systems
Cloud Infrastructure & Security
CloudScale is an industry leader in multi-cloud cost optimization and zero-trust perimeter networking tools serving over 4,000 enterprise clients globally.
Related Roles
Senior Full-Stack Laravel & Vue Developer
TechPulse Solutions
•
$135,000 - $170,000 /yr
AI / Machine Learning Engineer (LLM Applications)
TechPulse Solutions
•
$160,000 - $210,000 /yr
Lead UI/UX & Design Systems Designer
DesignCraft Studio
•
$125,000 - $160,000 /yr