, , ,

Zero-Downtime Citizen Services: The Leadership Mandate We Can’t Ignore 

When a streaming service goes down, people complain about it online. When a government service goes down, somebody can get hurt.

I don’t say that for effect. I spent more than three decades in state government IT, across Pennsylvania, Washington and North Carolina, and I’ve been in the room on the nights when a system residents depended on wasn’t working and nobody could tell me yet why. Those nights taught me something I didn’t fully appreciate earlier in my career: availability isn’t a technical subject. It gets discussed as one, usually near the end of a briefing, in percentages and maintenance windows. But what’s actually being decided in that conversation is whether the public can rely on their government to function.

That’s a leadership question. It belongs to governors, mayors, agency heads and budget directors, not just to the CIO.

No Resident Has Ever Asked Me About Six Nines

For most of my career, uptime lived in a dashboard. We reported it monthly, we negotiated it into service-level agreements, and we talked about it in language nobody outside IT uses. I’ve sat in front of legislators and agency secretaries with slides full of availability figures, and I can tell you those numbers landed as reassurance, not as a decision point. Nobody pushed back. Nobody asked what would happen on the worst day.

Meanwhile the person on the other end of the system has a much simpler standard. They want the thing to work when they need it. Someone calling 911 during a medical emergency, a parent logging into a court portal at ten at night to check a hearing date, a small business owner trying to file before a deadline, a teacher waiting on a paycheck that funds the mortgage payment: None of them are consuming technology. They’re exercising a right, meeting an obligation or asking their government for help.

When the system isn’t there, what they take away isn’t “the vendor had an issue.” What they take away is that their government couldn’t be counted on. I think we’ve badly underestimated how much that accumulates. One outage is a bad day. A pattern of outages is a story people tell about whether government works.

The Systems That Can’t Afford a Bad Day

Some services are simply too consequential to fail, even briefly, and in my experience those are frequently the systems running on the oldest infrastructure with the most fragile integrations and the most manual workarounds holding them together. That’s not an accident. Critical systems get hardened against change precisely because they’re critical, and over time that caution turns into technical debt nobody wants to touch.

Emergency dispatch is the clearest case. A routing problem or a database hiccup in a 911 environment isn’t an inconvenience, and it isn’t something you get to explain away afterward. That kind of resilience has to be governed as a public safety matter, with architecture and testing that assume surges, failures and attacks are coming, because they are.

Courts are the case I’d urge leaders to look at harder, because the harm is less visible. Case management systems and public portals are now how most people actually interact with the justice system. When a portal is down and someone can’t confirm a hearing time, file a document or make a payment, the consequences don’t fall evenly. The resident with a lawyer and flexible work hours absorbs it. The resident without either one misses a date and picks up a fine or a warrant. That’s not an IT problem with an equity side effect. That’s an equity problem that happens to be caused by IT.

Payroll is the one that taught me the most. A payroll system serving teachers, first responders and social workers has a hard deadline that no amount of explanation moves. A routine upgrade goes sideways and thousands of people either don’t get paid on schedule, get paid less than expected or don’t get paid at all. The effects are immediate and personal, and word travels through an agency faster than any status page. I’ve watched a single missed cycle do lasting damage to trust between a workforce and its leadership. Calling that a technical mishap lets too many people off the hook.

Tax and revenue systems carry a different kind of weight. They’re how residents comply with the law and how the state funds everything else it does, and their peak demand is entirely predictable. So is the damage when they’re unavailable during it.

And then there are the public portals, the front doors. For a lot of residents, particularly in the rural parts of every state I’ve worked in, the website isn’t one channel among several. It’s the only practical one. If that door is locked often enough, people stop assuming it will be open.

The Tradeoff I Stopped Accepting

Somewhere along the way, our field trained leaders to accept downtime as the cost of doing business. We have to bring the system down for the upgrade. We didn’t anticipate this attack. The platform was never designed for this much load. I’ve said versions of all three, and I no longer think they hold up, because the technology has moved and the expectations have moved with it.

On the security side, the honest framing has changed. The question isn’t whether an incident will happen, and I’d be skeptical of anyone who tells you they can prevent all of them. The question worth asking your team is whether the state can keep delivering essential services while responding to one. That’s a very different design conversation. Segmented architectures, immutable backups, tested rapid recovery and active-active designs turn a binary outcome into a graduated one, where the critical functions stay up, possibly in a reduced mode, while remediation happens out of the public’s view. Residents are going to judge us less on whether we got hit and more on whether their day was disrupted.

On the upgrade side, the old norm was that the system would be unavailable over a weekend and everyone would live with it. That norm made sense when the alternative was a phone call or a field office visit. It makes much less sense now that the portal is the service. Taking a tax system offline near a filing deadline, or a court scheduling system during a heavy hearing period, or a payroll system close to a pay date, is a decision with real cost. A commercial company that did that repeatedly would lose customers. We don’t lose residents. We lose their confidence, which is harder to get back and much harder to measure. Rolling upgrades and non-disruptive change are available, they work, and I’d treat them as a baseline requirement rather than a premium feature.

What I’d Ask if I Were Still in the Chair

Making this a genuine mandate means changing how decisions get made, not just how systems get built. Looking back, the questions I wish I’d forced onto the agenda earlier are fairly plain. 

  • Which services in our portfolio must never go down, and have we ever written that list down and agreed to it?
  • Does funding and oversight actually follow availability, or only features and compliance?
  • Is our cyber response plan built around staying up, or does it default to shutting things off? 
  • When an agency brings forward a new system, do we scrutinize its upgrade and availability strategy as seriously as we scrutinize its price?

Technology teams can build resilient systems. I’ve worked with people who are very good at it. What they can’t do on their own is make continuity non-negotiable when the budget gets tight or the timeline slips. That part only comes from the top.

Uptime isn’t a number buried in a dashboard. From where I sit, it’s one of the most visible statements a government makes about how seriously it takes its obligations to the people it serves.


For state, local, and education leaders thinking through what that looks like in practice, see additional resources on modernization, cyber resilience, and building a more predictable data foundation.

Jim Weaver is a former state CIO and nationally recognized public-sector technology leader with deep experience guiding government IT strategy, modernization, and cybersecurity initiatives. He served as Secretary and Chief Information Officer for the North Carolina Department of Information Technology, where he oversaw statewide IT strategy, procurement, cybersecurity, and broadband expansion. Prior to that, he served as CIO for the state of Washington, helping strengthen the state’s IT infrastructure and advance technology adoption across government. Jim also served as president of the National Association of State Chief Information Officers (NASCIO), contributing to IT policy and collaboration nationwide.

Photo Credit: August de Richelieu, Pexels

Leave a Comment

Leave a comment

Leave a Reply