Software Engineering

SRE

Masters chaos at scale, turning system failures into competitive advantages through obsessive measurement and automation. Produces CTOs who sleep soundly because they've eliminated the 3 AM pages that haunt other executives.

L1 – L9 · 9 tours Leads to: CTO → What's a Reference DRS?

The Career Arc

Rotational · L1–L3

Build the SRE craft. Prove you can wield the tools of Software Engineering.

  • L1 : Learn SRE fundamentals through operational work
  • L2 : Own reliability for services independently
  • L3 : Drive reliability improvements with strong judgment

Transformational · L4–L7

Deliver SRE outcomes — each Software Engineering tour at this altitude has a defined mission and success criteria.

  • L4 : Lead reliability projects and mentor others
  • L5 : Drive SRE practices across the org
  • L6 : Set reliability direction company-wide
  • L7 : Shape the company's reliability vision

Manage a Team?

Great SRE managers are practitioners first. The Software Engineering IC responsibilities in L4–L7 are your foundation — your management responsibilities are additive:

  • Run 1:1s for growth, not status—development conversations, not standups
  • Hire engineers who raise the bar—own the interview process end-to-end
  • Give direct feedback with care—it's how people actually improve
  • Remove blockers—shield the team from chaos so they can ship
  • Watch for burnout and team health—intervene before it's a crisis

Foundational · L8–L9

Shape the Software Engineering organization from the SRE chair — build institutions, not just products.

  • L8 : Build and lead SRE teams
  • L9 : Own reliability strategy and execution
→ C-Suite: L10 is the CTO path — a distinct page, not duplicated here.

L1 — Associate SRE Rotational

Mission

Learn SRE fundamentals through operational work

This tour of duty

Handle on-call and resolve your first incidents

Own the outcomes

  • Learn SRE fundamentals including monitoring, alerting, and incident response
  • Build simple monitoring dashboards and alerts with guidance
  • Write runbooks for operational procedures you learn
  • Participate in on-call rotations with senior engineer support
  • Document incidents and contribute to postmortem discussions
  • Support toil reduction by automating simple operational tasks

SRE at L1 — the competency bar

Software Engineering
2
Information Security
2
IT Operations
1

AI in this role

  • Analyzing logs for incident patterns
  • Drafting runbooks
  • Generating monitoring configs

L2 — Junior SRE Rotational

Mission

Own reliability for services independently

This tour of duty

Own reliability for a service with measurable SLOs

Own the outcomes

  • Implement monitoring and alerting for services independently
  • Respond to incidents and drive initial triage effectively
  • Write comprehensive runbooks and troubleshooting guides
  • Design SLOs for services with product team input
  • Build automation that reduces operational toil
  • Participate fully in on-call with appropriate escalation

SRE at L2 — the competency bar

Software Engineering
2
IT Operations
2
Information Security
1
Quality Engineering
1
Operational Excellence
1

AI in this role

  • Building observability dashboards
  • Analyzing incident patterns
  • Creating alert rules

L3 — Senior SRE Rotational

Mission

Drive reliability improvements with strong judgment

This tour of duty

Lead a project that improves reliability or reduces toil

Own the outcomes

  • Own reliability for services end-to-end including SLOs and alerting
  • Lead incident response and conduct thorough postmortems
  • Design observability solutions for medium-complexity systems
  • Identify and eliminate sources of toil systematically
  • Mentor junior engineers on SRE practices and incident response
  • Drive reliability improvements that measurably improve SLOs

SRE at L3 — the competency bar

Software Engineering
3
IT Operations
2
Information Security
1
Quality Engineering
1
Operational Excellence
1

AI in this role

  • Modeling failure scenarios
  • Synthesizing postmortem insights
  • Drafting SLO proposals

L4 — Staff SRE / Manager, Engineering Transformational

Mission

Lead reliability projects and mentor others

This tour of duty

Design systems that prevent categories of incidents

Own the outcomes

  • Lead SRE projects that improve reliability across multiple services
  • Design reliability architectures that prevent categories of incidents
  • Mentor engineers on SRE thinking and production excellence
  • Define SLO frameworks and incident management standards
  • Drive cross-team reliability improvements and best practices
  • Own reliability for critical production systems

SRE at L4 — the competency bar

Software Engineering
3
IT Operations
2
Information Security
2
Quality Engineering
1
Operational Excellence
1

AI in this role

  • Designing reliability systems
  • Analyzing toil patterns
  • Generating automation code

L5 — Senior Staff SRE / Senior Manager, Engineering Transformational

Mission

Drive SRE practices across the org

This tour of duty

Drive SRE practices that improve reliability org-wide

Own the outcomes

  • Drive SRE architecture decisions that affect the organization
  • Design reliability platforms that scale with system complexity
  • Define SRE engineering standards and best practices org-wide
  • Lead evaluation of observability and reliability tools
  • Mentor senior engineers and shape reliability culture
  • Solve the hardest reliability problems across the organization

SRE at L5 — the competency bar

Software Engineering
4
IT Operations
3
Information Security
2
Quality Engineering
1
Operational Excellence
1

AI in this role

  • Evaluating reliability technologies
  • Building SRE documentation
  • Creating roadmaps

L6 — Director, Site Reliability Engineering Transformational

Mission

Set reliability direction company-wide

This tour of duty

Define reliability standards that shape engineering culture

Own the outcomes

  • Set technical direction for site reliability across the company
  • Define SRE technology strategy and multi-year roadmap
  • Establish standards that ensure production excellence
  • Drive technical alignment on reliability investments
  • Represent SRE in executive technical discussions
  • Shape the vision for reliability engineering evolution

SRE at L6 — the competency bar

Software Engineering
4
IT Operations
3
Information Security
3
Strategy
2
Quality Engineering
1
Operational Excellence
1

AI in this role

  • Analyzing reliability patterns at scale
  • Generating standards
  • Building knowledge bases

L7 — Distinguished Engineer, Reliability Transformational

Mission

Shape the company's reliability vision

This tour of duty

Solve reliability problems that enable new scale

Own the outcomes

  • Shape the company's reliability vision and long-term strategy
  • Solve industry-level reliability challenges at scale
  • Define principles that guide reliability engineering decisions
  • Influence industry standards for site reliability engineering
  • Mentor directors and senior SRE leaders
  • Drive reliability innovation that enables business growth

SRE at L7 — the competency bar

Software Engineering
2
IT Operations
2
Strategy
2
Information Security
1
Quality Engineering
1
Operational Excellence
1

AI in this role

  • Modeling reliability evolution
  • Analyzing industry practices
  • Creating vision documents

L8 — VP of Engineering, SRE Foundational

Mission

Build and lead SRE teams

This tour of duty

Build an SRE team with sustainable on-call

Own the outcomes

  • Build and lead SRE teams with sustainable on-call practices
  • Define organizational structure for site reliability engineering
  • Establish hiring standards for reliability engineers
  • Create the operating model for SRE excellence
  • Partner with product leadership on reliability investments
  • Develop SRE managers and technical leaders

SRE at L8 — the competency bar

Software Engineering
2
Information Security
2
IT Operations
1
Strategy
1

AI in this role

  • Building SRE dashboards
  • Analyzing team patterns
  • Creating hiring frameworks

L9 — SVP of Engineering, SRE Foundational

Mission

Own reliability strategy and execution

This tour of duty

Transform reliability practices across the org

Own the outcomes

  • Own SRE strategy and execution organization-wide
  • Define multi-year roadmap for reliability platforms
  • Build culture that attracts top reliability engineering talent
  • Partner with executives on reliability investment strategy
  • Establish reliability as a competitive differentiator
  • Shape the future of site reliability at the company

SRE at L9 — the competency bar

Information Security
2
Software Engineering
1
IT Operations
1
Strategy
1

AI in this role

  • Modeling reliability scenarios
  • Building strategy documents
  • Designing knowledge infrastructure

What Hiring Managers Look For

You demonstrate production incident ownership with clear post-mortems that show learning from failure, not just firefighting ability.

You architect monitoring and automation solutions that measurably reduce toil while mentoring junior engineers through complex distributed systems debugging.

You present infrastructure cost optimization strategies with concrete ROI metrics and can articulate how reliability investments align with business objectives to non-technical executives.

Common Career Transitions

SRE → Platform Engineering at L5-L6 for developer productivity focus

SRE → Engineering Manager at L6-L7 leveraging ops experience for team leadership

SRE → Principal Architect at L7+ applying systems thinking to enterprise infrastructure

Official Classifications

System Code Official Title
O*NET-SOC (US) 15-1299.08 Computer Systems Engineers/Architects
ISCO-08 (UN/ILO) 2522 Systems Administrators

At L6 and above, the manager classification 1330 — Information and Communications Technology Services Managers applies IN ADDITION to the professional code — a manager is a superset of the individual contributor, never a replacement.

Measure yourself against this ladder — pin it to your Career Record.

Build Your Career Record