Software Engineering
SRE
Masters chaos at scale, turning system failures into competitive advantages through obsessive measurement and automation. Produces CTOs who sleep soundly because they've eliminated the 3 AM pages that haunt other executives.
The Career Arc
Rotational · L1–L3
Build the SRE craft. Prove you can wield the tools of Software Engineering.
Transformational · L4–L7
Deliver SRE outcomes — each Software Engineering tour at this altitude has a defined mission and success criteria.
- L4 : Lead reliability projects and mentor others
- L5 : Drive SRE practices across the org
- L6 : Set reliability direction company-wide
- L7 : Shape the company's reliability vision
Manage a Team?
Great SRE managers are practitioners first. The Software Engineering IC responsibilities in L4–L7 are your foundation — your management responsibilities are additive:
- • Run 1:1s for growth, not status—development conversations, not standups
- • Hire engineers who raise the bar—own the interview process end-to-end
- • Give direct feedback with care—it's how people actually improve
- • Remove blockers—shield the team from chaos so they can ship
- • Watch for burnout and team health—intervene before it's a crisis
Foundational · L8–L9
Shape the Software Engineering organization from the SRE chair — build institutions, not just products.
L1 — Associate SRE Rotational
Mission
Learn SRE fundamentals through operational work
This tour of duty
Handle on-call and resolve your first incidents
Own the outcomes
- • Learn SRE fundamentals including monitoring, alerting, and incident response
- • Build simple monitoring dashboards and alerts with guidance
- • Write runbooks for operational procedures you learn
- • Participate in on-call rotations with senior engineer support
- • Document incidents and contribute to postmortem discussions
- • Support toil reduction by automating simple operational tasks
SRE at L1 — the competency bar
AI in this role
- • Analyzing logs for incident patterns
- • Drafting runbooks
- • Generating monitoring configs
L2 — Junior SRE Rotational
Mission
Own reliability for services independently
This tour of duty
Own reliability for a service with measurable SLOs
Own the outcomes
- • Implement monitoring and alerting for services independently
- • Respond to incidents and drive initial triage effectively
- • Write comprehensive runbooks and troubleshooting guides
- • Design SLOs for services with product team input
- • Build automation that reduces operational toil
- • Participate fully in on-call with appropriate escalation
SRE at L2 — the competency bar
AI in this role
- • Building observability dashboards
- • Analyzing incident patterns
- • Creating alert rules
L3 — Senior SRE Rotational
Mission
Drive reliability improvements with strong judgment
This tour of duty
Lead a project that improves reliability or reduces toil
Own the outcomes
- • Own reliability for services end-to-end including SLOs and alerting
- • Lead incident response and conduct thorough postmortems
- • Design observability solutions for medium-complexity systems
- • Identify and eliminate sources of toil systematically
- • Mentor junior engineers on SRE practices and incident response
- • Drive reliability improvements that measurably improve SLOs
SRE at L3 — the competency bar
AI in this role
- • Modeling failure scenarios
- • Synthesizing postmortem insights
- • Drafting SLO proposals
L4 — Staff SRE / Manager, Engineering Transformational
Mission
Lead reliability projects and mentor others
This tour of duty
Design systems that prevent categories of incidents
Own the outcomes
- • Lead SRE projects that improve reliability across multiple services
- • Design reliability architectures that prevent categories of incidents
- • Mentor engineers on SRE thinking and production excellence
- • Define SLO frameworks and incident management standards
- • Drive cross-team reliability improvements and best practices
- • Own reliability for critical production systems
SRE at L4 — the competency bar
AI in this role
- • Designing reliability systems
- • Analyzing toil patterns
- • Generating automation code
L5 — Senior Staff SRE / Senior Manager, Engineering Transformational
Mission
Drive SRE practices across the org
This tour of duty
Drive SRE practices that improve reliability org-wide
Own the outcomes
- • Drive SRE architecture decisions that affect the organization
- • Design reliability platforms that scale with system complexity
- • Define SRE engineering standards and best practices org-wide
- • Lead evaluation of observability and reliability tools
- • Mentor senior engineers and shape reliability culture
- • Solve the hardest reliability problems across the organization
SRE at L5 — the competency bar
AI in this role
- • Evaluating reliability technologies
- • Building SRE documentation
- • Creating roadmaps
L6 — Director, Site Reliability Engineering Transformational
Mission
Set reliability direction company-wide
This tour of duty
Define reliability standards that shape engineering culture
Own the outcomes
- • Set technical direction for site reliability across the company
- • Define SRE technology strategy and multi-year roadmap
- • Establish standards that ensure production excellence
- • Drive technical alignment on reliability investments
- • Represent SRE in executive technical discussions
- • Shape the vision for reliability engineering evolution
SRE at L6 — the competency bar
AI in this role
- • Analyzing reliability patterns at scale
- • Generating standards
- • Building knowledge bases
L7 — Distinguished Engineer, Reliability Transformational
Mission
Shape the company's reliability vision
This tour of duty
Solve reliability problems that enable new scale
Own the outcomes
- • Shape the company's reliability vision and long-term strategy
- • Solve industry-level reliability challenges at scale
- • Define principles that guide reliability engineering decisions
- • Influence industry standards for site reliability engineering
- • Mentor directors and senior SRE leaders
- • Drive reliability innovation that enables business growth
SRE at L7 — the competency bar
AI in this role
- • Modeling reliability evolution
- • Analyzing industry practices
- • Creating vision documents
L8 — VP of Engineering, SRE Foundational
Mission
Build and lead SRE teams
This tour of duty
Build an SRE team with sustainable on-call
Own the outcomes
- • Build and lead SRE teams with sustainable on-call practices
- • Define organizational structure for site reliability engineering
- • Establish hiring standards for reliability engineers
- • Create the operating model for SRE excellence
- • Partner with product leadership on reliability investments
- • Develop SRE managers and technical leaders
SRE at L8 — the competency bar
AI in this role
- • Building SRE dashboards
- • Analyzing team patterns
- • Creating hiring frameworks
L9 — SVP of Engineering, SRE Foundational
Mission
Own reliability strategy and execution
This tour of duty
Transform reliability practices across the org
Own the outcomes
- • Own SRE strategy and execution organization-wide
- • Define multi-year roadmap for reliability platforms
- • Build culture that attracts top reliability engineering talent
- • Partner with executives on reliability investment strategy
- • Establish reliability as a competitive differentiator
- • Shape the future of site reliability at the company
SRE at L9 — the competency bar
AI in this role
- • Modeling reliability scenarios
- • Building strategy documents
- • Designing knowledge infrastructure
What Hiring Managers Look For
You demonstrate production incident ownership with clear post-mortems that show learning from failure, not just firefighting ability.
You architect monitoring and automation solutions that measurably reduce toil while mentoring junior engineers through complex distributed systems debugging.
You present infrastructure cost optimization strategies with concrete ROI metrics and can articulate how reliability investments align with business objectives to non-technical executives.
Common Career Transitions
SRE → Platform Engineering at L5-L6 for developer productivity focus
SRE → Engineering Manager at L6-L7 leveraging ops experience for team leadership
SRE → Principal Architect at L7+ applying systems thinking to enterprise infrastructure
Official Classifications
| System | Code | Official Title |
|---|---|---|
| O*NET-SOC (US) | 15-1299.08 | Computer Systems Engineers/Architects |
| ISCO-08 (UN/ILO) | 2522 | Systems Administrators |
At L6 and above, the manager classification 1330 — Information and Communications Technology Services Managers applies IN ADDITION to the professional code — a manager is a superset of the individual contributor, never a replacement.
Measure yourself against this ladder — pin it to your Career Record.
Build Your Career Record