Every summer, I load up my RV and head out to a dry lakebed in the Nevada desert, but it’s not exactly for a vacation. For two weeks, I’m part of the Emergency Services Department at Burning Man, working in emergency dispatch for a temporary city of more than 60,000 people. I started there as a dispatcher in 2010, and I’ve been a Dispatch Supervisor since 2012.
The ESD dispatch center looks more like a high-tech operations center than you might expect. It runs a modern computer-aided dispatch (CAD) system and a municipal-grade two-way radio system, and it operates around the clock. Cell coverage is spotty when that many people descend on an empty stretch of desert, so nearly all of our emergency calls come in by radio, from event staff, participants, and partner agencies. We coordinate emergency response across ESD and its partners: the Black Rock Rangers, the on-site advanced medical care facility, and federal, state, and local law enforcement. The stakes are immediate and physical in a way that a software outage rarely is.
As a Dispatch Supervisor, my job is to keep track of the overall picture: every active incident, conditions throughout the city (from white-out dust storms to roads made impassable by rain), and the status of the emergency response system as a whole (from the function of the radio system to the availability and location of emergency units of various types). I direct the dispatchers’ work, handle the escalations that don’t fit the routine, and serve as the liaison to our partner agencies.
I take a radio channel myself only in a few situations: when we’re short-staffed, while a dispatcher takes a break, or when I urgently need to relay critical information directly to a responding crew, such as something I’ve just learned from our law enforcement partners. Those are deliberate, temporary exceptions. It’s similar to the division of labor I recommend between an incident commander and the engineers doing the technical work: the coordinator can step in briefly when it helps, but by default stays hands-off so that someone is always watching the whole picture.
Back in the tech world, I help engineering leaders strengthen their companies’ incident response. Many of the practices the public safety world takes for granted are unfamiliar to their teams, even though those practices adapt readily to software.
I’ve been working across both of these worlds for more than thirty years. I started as a volunteer search and rescue pilot while building a tech career in Silicon Valley, and the two tracks have run in parallel ever since. On the public safety side, that’s meant search and rescue missions, CERT instruction, and emergency dispatch. On the tech side, it’s meant building Google’s incident management system, leading incident response at Slack, helping companies from startups to large enterprises strengthen their incident management capabilities as a consultant, and now writing a book, Incident Management for DevOps and SRE. Each side makes me better at the other.
What you see from both sides
When you’ve worked across both fields long enough, certain things stand out that are hard to see from within either one alone.
Incidents are normal work, not failures. For a fire department, emergencies are routine business. Nobody blames the department when it has a busy week; the questions are how well the crews handled the calls, and how to prevent similar ones in the future. Software is different in one important way: the people who respond to incidents are usually the same people who build and run the systems, and incident response competes with everything else on their plate. That makes it easy to treat each incident as an interruption, or as evidence that someone failed. But anyone operating complex systems will have incidents, and responding to them is part of the work of building and running those systems. Many tech companies’ cultures haven’t yet embraced that. They still treat incidents as something between an embarrassment and an accusation, which makes people reluctant to declare them, reluctant to escalate, and reluctant to be forthcoming afterward about what happened.
The hard part isn’t technical. In both worlds, the technical work matters enormously, whether that’s debugging a cascading failure or treating a patient with a heat emergency. But whether the response goes well or badly usually comes down to coordination. How do you organize people who arrive at different times with different information, different skills, and different concerns? How do you communicate clearly when the situation keeps changing? How do you make decisions when you can’t wait for all the facts? These questions are the same in both fields. The Incident Command System, developed by fire departments in the 1970s and refined by emergency agencies ever since, exists to answer them. It was built to solve coordination problems in emergencies, which is why its principles transfer to software incidents even though the technical work underneath is completely different.
The initial response is decided before the call comes in. Public safety agencies think through in advance what kinds of calls they’re likely to get and what the initial response to each will be. At Burning Man, for example, a small fire gets one engine, while a large fire gets three engines and a water tender (a large tanker truck). A minor medical call such as a twisted ankle gets a QRV (Quick Response Vehicle: a small, nimble utility vehicle with a medical crew and minimal equipment, which can more easily move through the crowded roads and camps than a full-size ambulance), while a major medical call such as a heart attack gets a QRV, an advanced life support ambulance (which carries more equipment and emergency drugs for patient treatment), and a supervisor. With that dispatch matrix, nobody has to work out, in the middle of an emergency, how much help to send: the dispatcher classifies the call, and the matrix provides a starting point that responders can adjust once they’re on scene, scaling up, scaling down, or adding a different kind of help, such as a crisis intervention team when a call turns out to involve a mental health issue. The tech equivalent is the severity matrix, which many companies aren’t using to its full potential.
Good incident response looks boring from the outside. When a fire department arrives at a structure fire, the coordination is almost invisible. The incident commander is tracking the situation, allocating resources, and communicating with incoming units, while the crews fight the fire. Everyone knows their role and what information they need to share, and outsiders often marvel at how calm and focused the responders are. The smooth responses are unremarkable; the dramatic ones are the exceptions. Contrast that with many tech companies, where every incident is handled ad hoc, the most senior engineer is simultaneously debugging and trying to coordinate, and nobody quite knows who’s doing what.
What this means for tech companies
That contrast between structured coordination and improvising from scratch is one of the most common gaps I see when I assess companies’ incident management programs. The teams aren’t short on technical skills or motivation. They’ve never seen what emergency coordination looks like when it’s a routine capability rather than something that gets reinvented every time.
I don’t want to oversell the analogy. Software incidents and physical emergencies differ in important ways: reversibility, speed of communication, the ability to respond from anywhere in the world, and the physical danger responders face. The Incident Command System as practiced by a fire department doesn’t transfer directly to a software team, and anyone who tells you it does hasn’t spent enough time in both worlds to see where the parallels break down. What transfers are the underlying principles. For example: coordination needs explicit structure; the person coordinating generally shouldn’t also be doing the technical work; communication patterns should be designed in advance; and incidents are a normal part of operating complex systems.
Those principles are robust across domains. They’ve been tested under pressures that make a software outage look gentle by comparison. And the companies I’ve worked with that have adopted them, adapted for their own context, handle incidents with noticeably less chaos, less stress, and less customer impact than the ones still improvising from scratch.
That’s the perspective I carry back and forth between the dispatch center and the consulting engagement: the conviction that the coordination problems companies struggle with during incidents are neither new nor unsolvable, and that the path forward doesn’t require inventing anything from scratch. It requires studying what’s already been learned in fields where the stakes have always been high, and asking which of those lessons fit.