Staff Reliability Engineer
Skills & Stack
About the Role
The TDI Network Engineering team at Okta is responsible for the global corporate network. This operations-focused staff role owns the resilience, health and availability of that network, responding to alerts, maintaining service availability and driving down systemic toil and technical debt across teams using distributed systems, networking, infrastructure as code and observability expertise. You will own multi-quarter reliability objectives and report to the Network Engineering Manager.
About Okta
Okta is an identity security company that builds trusted, neutral infrastructure to secure every identity, from human users to AI agents, for thousands of organizations worldwide.
What You'll Do
- 1Own the resilience, health and availability of the global corporate network, including alert response and reliability projects
- 2Drive strategic reduction of toil and technical debt through automation and self-service operational tooling
- 3Collaborate with Business Technology, Workplace, Security and executive stakeholders
- 4Lead resolution of complex network operations issues and prevent future outages
- 5Mentor team members and define operational engineering best practices
What We're Looking For
- 8+ years of related experience with a Bachelors degree, or 6+ years with a Masters
- Deep expertise in AWS Networking and Palo Alto Networks solutions
- Operational experience with WiFi, DNS, DHCP, VLANs, VPN, ACLs, routing and firewall policies
- Proficiency with Terraform or Ansible, Prometheus or Grafana, and Python or Go
- Experience managing SLOs and SLIs for large-scale enterprise environments
Ready to apply for this role?
Staff Reliability Engineer at Okta — click below to submit your application.
