Customer service design and support leadership
In 2020 I led the customer support operation of a global-scale cloud provider — the year work, school and family traffic all moved home at once. This is what running support at scale taught me about how service fails.
Context
I had joined as a cloud engineer in 2019. A year later the pandemic moved everything onto the platform at the same time, and I was leading the team that answered when it broke — customer-facing support and network operations in one function.
The problem
Support fails two ways. Technically: an incident outruns our understanding. As service: the right information exists somewhere in the company but never reaches the person waiting for it — the customer, or the engineer who could fix it. The second failure is more common and more damaging, and it's a design problem, not a staffing one.
My role
Owning the service end to end: how customers get helped, escalation paths, shift handovers, runbooks, and the people awake at 3 a.m.
Constraints
At that scale no single person holds the system in their head, so context has to survive handovers intact. The customer on the other end of a broken service doesn't care whose layer it is. And the constraint I cared about most: the youngest engineer on the 4 a.m. shift still had to make good calls with incomplete information.
Discovery
Patterns across incidents and tickets taught the durable lessons: which alerts predicted trouble and which were noise, which escalation paths worked and which just relocated anxiety, where runbooks had been written for their author instead of their reader, and what customers actually needed to hear during an outage — honest status, not reassurance. Postmortems were the curriculum.
The decision
Three decisions that stuck. One: a fixed handover format — what we know, what we've ruled out, what we're watching, who owns the next step. Two: runbooks written to be executed by a tired stranger, not their author. Three: an explicit norm that 'we don't know yet' is an acceptable status. Stated ambiguity beats false confidence every time.
Trade-offs
Structure costs speed in quiet moments to buy correctness in bad ones. A fixed handover format feels bureaucratic on a quiet Tuesday and is priceless during a multi-region event. Choosing clarity over heroics also means the brilliant-improvisation path is deliberately closed.
Execution
Working with engineering on alert quality, with shift leads on handover discipline, with every incident review on feeding lessons back into the runbooks — and keeping what we told customers during incidents honest. Half the job was protocol; the other half was trust.
Outcome
A support function that got through the platform's hardest traffic year, and incident reports engineering could act on without re-deriving them.
What I learned
Most 'technical' failures at scale are interface failures between humans. That conclusion is what later moved me toward product.