Home / Blog / Uptime monitoring
Uptime monitoring

How much detail should a public postmortem include after a customer-facing outage?

Too little detail reads as evasive, too much reads as an excuse. Here is a structure for the public version of a postmortem that earns trust without exposing what should stay internal.

A small group of people sitting in a circle on chairs in a bright loft, talking seriously, paper notebooks on their laps

Who the public postmortem is for

An internal postmortem is for the team. It exists to find contributing causes, decide on fixes, and improve the system without blame. The public version has a different reader: a customer who lost something during the outage, or a prospect deciding whether you can be trusted with their business. They are not reading to learn how your infrastructure works. They are reading to answer three questions: do these people understand what happened, have they fixed it, and will it happen again? Every paragraph should serve one of those questions. Related: How do you run an on-call rotation when your whole team is three people?

This changes what detail means. The customer needs the timeline, the impact, the cause in plain terms, and the concrete changes. They do not need the name of the internal service, the stack trace, or the debate about which fix to choose. Detail that helps the reader judge your competence belongs in. Detail that only makes sense to your own engineers, or that exposes attack surface, belongs in the internal document. The public version is a translation, not a redaction.

Keep reading: How to Choose an Uptime Check Interval That Actually Catches Outages, Status Page Best Practices That Reduce Support Tickets, Uptime vs Response Time: What to Monitor and Why Both Matter. See how PingCrumb helps you uptime monitoring and hosted status pages.

What must always be in it

Start with impact in the customer's terms: what was unavailable, for whom, for how long, and whether any data was lost or delayed. Be specific about time, using a single time zone and stating it. If some customers were affected and others were not, say how to tell which group you were in. Then give a timeline with the key moments: when the problem started, when you detected it, when you understood it, when it was mitigated, and when it was fully resolved. The gap between start and detection is the most honest number in the document, and readers notice when it is missing.

Then explain the cause in a way a technical customer can follow without inside knowledge. Saying that a deploy introduced a configuration change that exhausted the database connection pool under normal load is a complete public explanation. The precise setting, the module name, and who approved the change are not needed. Finish with the changes: what has already been done, what is scheduled, and roughly when. Vague promises to improve processes read as nothing. A specific list of three or four actions, even small ones, reads as a team that took it seriously. Related: When should a small SaaS start monitoring DNS resolution separately from HTTP checks?

What to leave out, and why

Leave out anything that helps an attacker: internal hostnames, software versions, the exact nature of a vulnerability before it is patched everywhere, and details of your security tooling. Leave out names of individuals. A postmortem that blames a person teaches customers that you fire people instead of fixing systems, and teaches your own team to hide mistakes. Leave out the raw metrics dumps and log excerpts, and replace them with one or two clear sentences about what the data showed.

Be careful with third parties. If a vendor outage contributed, you can say that a provider you depend on had an incident, and you can explain what you are changing so their next incident affects you less. Assigning blame to a named vendor rarely helps and often reads as deflection, because the customer bought from you, not from them. The same applies to root cause claims that are still under investigation. If you do not yet know, say what you know, say what you are still looking into, and commit to an update by a date. Related: How to Write a Clear Incident Update Your Customers Trust

Tone, timing, and where to publish it

Publish the postmortem within a few business days of the incident. Sooner than that and it is usually incomplete; later and customers assume you have moved on. Post it as a follow-up on the status page incident itself so that anyone who subscribed to updates sees it, and link it from any email you send to affected customers. If the outage was significant, a short note in your changelog or blog is reasonable, but the status page is the canonical home. Keep a permanent archive, since prospects doing diligence will look for it. Related: Status Page Best Practices That Reduce Support Tickets

Write it in first person plural, in plain language, without marketing. Apologize once, clearly, at the start, and do not repeat it in every paragraph. Avoid words that minimize, like minor or brief, unless the customer would agree. Avoid words that dramatize, too. A calm, specific, complete account is what convinces people that the same team will be calm, specific, and complete during the next incident. That is the actual product of a public postmortem: not closure on this outage, but credibility for the next one.

Key takeaways
  • The public postmortem answers three questions: do you understand what happened, is it fixed, and will it recur.
  • Always include customer-facing impact, an honest timeline with detection time, a plain-language cause, and a specific list of changes.
  • Leave out internal hostnames, versions, individual names, raw logs, and blame aimed at named vendors.
  • Publish within a few business days as a status page follow-up, in plain first-person language with a single clear apology.
Julien Jimenez
Written by

Julien Jimenez

Julien Jimenez is an independent software builder based in Paris. He designs, ships, and operates focused SaaS products for small businesses and independent professionals. Read the full author page.

Know before your customers do

Uptime monitoring and hosted status pages. PingCrumb is built to help you put this into practice.

Start monitoring

Get the PingCrumb playbook

Practical guides on uptime monitoring, straight to your inbox as we publish them. No spam, unsubscribe any time.

By subscribing you agree to our privacy policy.