Uptime Monitoring and Incident Response for Small Teams: Know Before Your Customers Do
The worst way to learn that your application is down is an email from a customer. It is common in small companies, because nobody was assigned to watch the system and nothing was set up to watch it for them.
Uptime monitoring and incident response do not require an operations department. They require a handful of automated checks, an alert that reaches a named person, and a short written routine for what happens next. If a small team runs your custom software or web application, this guide covers what to set up and what to ask of your developers and your host.
Start by checking from the outside
The first check to set up is an external one: a service outside your own infrastructure that requests an endpoint every minute or so and alerts you when it fails. Google's site reliability engineering book calls this black-box monitoring, meaning testing externally visible behavior as a user would see it. A monitor that lives on the same server as the application goes down with it.
Point the check at something meaningful. A static home page can keep loading while the database is unreachable and nobody can log in. Many frameworks provide a health endpoint for this purpose. Laravel, for example, ships a health route at /up that returns a 200 response if the application boots without exceptions and a 500 otherwise, and it can be extended to check the database or cache.
What to monitor beyond up or down
Most real incidents are partial: the site responds, but something inside it is failing. Five more things are worth watching.
- Errors: the rate of failed requests and unhandled exceptions, so a spike after a release is visible within minutes.
- Response time: a page that takes twenty seconds is down as far as the user is concerned.
- Background jobs: queues, scheduled tasks and backups fail silently. Have each job report in when it completes, and alert when the report does not arrive.
- Expiry dates: TLS certificates, domain names and the payment cards that renew them.
- Capacity: disk space, memory, database connections and database storage.
Certificates deserve a note. Automatic renewal is now normal, but automation breaks quietly. Let's Encrypt, a widely used free certificate authority, stopped sending expiration warning emails on June 4, 2025, so if those emails were your safety net, you need your own check that warns you well before expiry.
Alerts that reach a person
An alert sent to a shared inbox or a busy chat channel is a record, not an alert. When customers cannot use the product, the notification has to reach one specific person on a channel that wakes them up, such as a phone call or SMS from an on-call tool, and escalate to a second person if the first does not acknowledge it.
Decide who that person is outside working hours, and write it down. The honest options are a rota among the developers, a paid on-call arrangement with your development partner, or an explicit decision that out-of-hours outages wait until morning. The third is reasonable for an internal tool used in one time zone. It is not reasonable if you sell to both US and UK customers, because one country's evening is the other's working day.
Keep the number of waking alerts small. The same Google chapter argues that every page should be actionable. Separate the alerts that interrupt someone now from those that create a ticket for the next working day, or the team will learn to ignore both.
A four-step incident process for a small team
You do not need a formal framework. You need four steps that everyone knows.
- Acknowledge. One person says "I have this" in a known place, so that two people do not make conflicting changes and nobody assumes someone else is on it.
- Communicate. Tell customers and colleagues there is a problem before you know the cause.
- Fix. Restore service first and understand it later. Rolling back the last release is usually faster than diagnosing in production.
- Review. Within a few days, write down what happened, the impact, and what will change.
The review is the step small teams skip, and it is the one that prevents a repeat. Google describes a blameless postmortem as one that identifies contributing causes without indicting any individual or team. In practice, ask why the system let the mistake through, not who made it. One page is enough if every action has an owner and a date.
Status pages and customer communication
A status page is a public page showing whether your service is working and the history of recent incidents. Host it separately from your application, or it will be unavailable at the exact moment people look for it.
During an incident, post early and plainly: what is affected, what is not, and when the next update will come. Then meet that time, even if the update is "still investigating". Avoid guessing at a cause or promising a fix time you cannot control. Afterward, a short honest summary of what happened and what you changed does more for trust than silence.
What 99.9 percent uptime means in hours
Uptime percentages sound similar and are not. A year has 365 x 24 = 8,760 hours. At 99.9 percent uptime, the permitted downtime is 0.1 percent of that: 8,760 x 0.001 = 8.76 hours a year. A 30-day month has 720 hours, so the monthly allowance is 720 x 0.001 = 0.72 hours, or about 43 minutes.
The same arithmetic gives 87.6 hours a year (3.65 days) at 99 percent and about 53 minutes a year at 99.99 percent, which matches Google's published availability table. Each extra nine cuts the allowance by a factor of ten and generally demands more redundancy and on-call cover. Read the definition too: whether planned maintenance counts, whether the figure is measured per month or per year, and who does the measuring.
Service level agreements with your host and your developers
A hosting SLA is a financial term, not a guarantee that the service stays up. Amazon's compute SLA, for example, commits to commercially reasonable efforts toward a monthly uptime percentage, and the remedy for a miss is a service credit against future bills that the customer has to request. It covers the host's infrastructure, not your code or configuration.
The agreement with your developers or agency matters more day to day. Ask for response and resolution targets by severity, the hours and time zone covered, how you raise an urgent issue, and who owns monitoring. These terms normally sit in a support contract, covered in software maintenance and support explained. A response time is a promise to start work, not to finish it.
Common mistakes
A few errors account for many painful outages in small teams. Monitoring runs on the server it is meant to watch. Alerts go to a channel everyone has muted. Only the home page is checked. The domain renewal is tied to the email address or card of someone who has left. Backups have never been test-restored. One person holds all the access and is on a flight.
The opposite mistake also exists: an elaborate monitoring platform and a 24-hour rota for an internal tool with thirty users. Match the effort to what an hour of downtime actually costs you.
What to do next
This week, set up one external check on a real health endpoint, route its alert to a phone, and write down who answers it at night and at weekends. Add expiry checks for certificates and domains, and confirm that more than one person can log in to the host, the registrar and the DNS provider. Then draft a one-page incident routine and rehearse it once.
Conclusion
Knowing about an outage before your customers do comes down to a few habits: check from outside, watch for the quiet failures, make sure an alert reaches a named person, follow the same four steps every time, and read what an uptime figure or an SLA actually promises.
If you are planning a new application or want the monitoring and support terms for an existing one reviewed, you can contact Entrant Technologies to talk it through.