<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>OnlineOrNot</title>
        <link>https://onlineornot.com</link>
        <description>OnlineOrNot monitors your websites and APIs so you get notified instantly when they go down.</description>
        <lastBuildDate>Fri, 04 Sep 2026 10:35:05 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <image>
            <title>OnlineOrNot</title>
            <url>https://onlineornot.com/onlineornot-brand-og.png</url>
            <link>https://onlineornot.com</link>
        </image>
        <copyright>All rights reserved 2026, Max Rozen</copyright>
        <item>
            <title><![CDATA[7 best Atlassian Statuspage alternatives for 2026]]></title>
            <link>https://onlineornot.com/atlassian-statuspage-alternative</link>
            <guid>https://onlineornot.com/atlassian-statuspage-alternative</guid>
            <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Atlassian Statuspage is one of the best-known tools for communicating outages and scheduled maintenance. It gives customers a place to check service health instead of asking your support team whether something is down.</p>
<p>But it is not the right fit for every team.</p>
<p>You may want an Atlassian Statuspage alternative that includes monitoring, costs less as your audience grows, is easier to configure, or can be self-hosted. The difficult part is comparing complete incident workflows rather than screenshots and starting prices.</p>
<p>This guide compares seven alternatives, including hosted, free, and open-source options.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#atlassian-statuspage-alternative-the-short-answer">Atlassian Statuspage alternative: the short answer</a></li>
<li><a href="#quick-comparison">Quick comparison</a></li>
<li><a href="#what-does-atlassian-statuspage-do">What does Atlassian Statuspage do?</a></li>
<li><a href="#why-consider-an-alternative">Why consider an alternative?</a></li>
<li><a href="#1-onlineornot">1. OnlineOrNot</a></li>
<li><a href="#2-better-stack">2. Better Stack</a></li>
<li><a href="#3-instatus">3. Instatus</a></li>
<li><a href="#4-openstatus">4. OpenStatus</a></li>
<li><a href="#5-statuspal">5. StatusPal</a></li>
<li><a href="#6-uptimerobot">6. UptimeRobot</a></li>
<li><a href="#7-uptime-kuma">7. Uptime Kuma</a></li>
<li><a href="#free-and-open-source-atlassian-statuspage-alternatives">Free and open-source alternatives</a></li>
<li><a href="#how-to-choose-an-atlassian-statuspage-alternative">How to choose</a></li>
<li><a href="#how-to-migrate-from-atlassian-statuspage">How to migrate</a></li>
<li><a href="#faq">FAQ</a></li>
</ul>
<h2>Atlassian Statuspage alternative: the short answer</h2>
<p><a href="/status-pages">OnlineOrNot</a> is a strong Atlassian Statuspage alternative for SaaS companies and small engineering teams that want uptime monitoring and status pages in one product.</p>
<p>OnlineOrNot provides:</p>
<ul>
<li>Public and private status pages</li>
<li>Uptime, API, heartbeat, DNS, TCP, and browser monitoring</li>
<li>Automatic incidents connected to monitoring</li>
<li>Manual incidents and scheduled maintenance</li>
<li>Custom domains with automatic HTTPS</li>
<li>Unlimited email subscribers on paid plans</li>
<li>A limited free status page and a 14-day trial of paid features</li>
</ul>
<p>Paid plans start at $12/month with annual billing or $15/month billed monthly. The first public status page is included, without per-subscriber fees.</p>
<p><a href="/get-started">Create your status page</a> or <a href="/pricing">compare OnlineOrNot pricing</a>.</p>
<h2>Quick comparison</h2>
<p>| Tool | Monitoring included | Free option | Private pages | Best for |
|------|---------------------|-------------|---------------|----------|
| <strong>OnlineOrNot</strong> | Yes | Yes | Yes | SaaS teams wanting monitoring and status pages together |
| <strong>Better Stack</strong> | Yes | Yes | Yes | Teams wanting a broader incident-response platform |
| <strong>Instatus</strong> | Yes | Yes | Plan-dependent | Teams prioritizing a status-page-first workflow |
| <strong>OpenStatus</strong> | Yes | Yes / open source | Hosted add-on or DIY | Developers wanting open source with managed hosting |
| <strong>StatusPal</strong> | Via integrations | Trial | Yes | Teams wanting dedicated incident communication |
| <strong>UptimeRobot</strong> | Yes | Yes | Plan-dependent | Existing UptimeRobot users and basic monitoring |
| <strong>Uptime Kuma</strong> | Yes | Self-hosted | DIY | Home labs and teams comfortable operating infrastructure |</p>
<p>Features and limits change. Use this table to build a shortlist, then confirm that the current plan includes the page count, subscriber limit, access controls, and notification channels you need.</p>
<h2>What does Atlassian Statuspage do?</h2>
<p><a href="https://www.atlassian.com/software/statuspage">Atlassian Statuspage</a> is incident communication software. It publishes the health of services and components, active incidents, scheduled maintenance, metrics, and previous incidents.</p>
<p>Atlassian offers three broad page types:</p>
<ul>
<li><strong>Public pages</strong> for customer-facing services</li>
<li><strong>Private pages</strong> for employees or selected users</li>
<li><strong>Audience-specific pages</strong> that show different components to different customers</li>
</ul>
<p>Statuspage is primarily a communication product. Teams typically connect monitoring and incident-management systems through integrations or APIs to automate component states and incident updates.</p>
<p>That model works well for large organizations with established Atlassian workflows. Smaller teams may prefer a product where monitoring, alerts, and customer communication share the same data from the start.</p>
<h2>Why consider an alternative?</h2>
<p>Teams usually compare Atlassian Statuspage alternatives for four reasons.</p>
<h3>Monitoring and communication live in separate products</h3>
<p>A status page is only useful when it reflects what is actually happening.</p>
<p>If monitoring lives somewhere else, your team must maintain an integration or update incidents manually. A combined platform can detect an outage, alert the team, change affected components, and publish an incident from the same source of truth.</p>
<p>Automation should not remove human control. Monitoring can report that an API is unavailable, but a person still needs to explain customer impact, available workarounds, and when the next update will arrive.</p>
<h3>The full price is higher than the starting price</h3>
<p>The cheapest plan rarely represents a growing team's actual configuration.</p>
<p>Compare costs after adding:</p>
<ul>
<li>Public and private pages</li>
<li>Team members or responders</li>
<li>Email subscribers</li>
<li>Custom domains and branding</li>
<li>SSO, IP allowlists, or passwords</li>
<li>SMS and phone notifications</li>
<li>Monitoring and incident-management tools</li>
</ul>
<p>A tool with a higher starting price can still cost less if it replaces separate monitoring and status-page subscriptions.</p>
<h3>The workflow is more complex than the team needs</h3>
<p>Large incident-management platforms support approval processes, audience segmentation, and enterprise controls. Those features are valuable when you need them and extra administration when you do not.</p>
<p>A small SaaS team may only need reliable monitoring, automatic incidents, manual updates, and subscriber notifications.</p>
<h3>You want control over hosting</h3>
<p>Open-source tools let you control infrastructure and data. The tradeoff is responsibility for deployment, upgrades, security, backups, email delivery, and availability.</p>
<p>A self-hosted status page should run independently from the application it reports on. Otherwise, the page may disappear during the exact outage it is supposed to explain.</p>
<h2>1. OnlineOrNot</h2>
<p><a href="/status-pages">OnlineOrNot</a> combines hosted status pages with uptime, API, heartbeat, DNS, TCP, and Playwright browser monitoring.</p>
<p>Connect a monitor to a status-page component and OnlineOrNot can automatically create and resolve incidents when its state changes. Your team can still post manual incidents, progress updates, retrospective incidents, and scheduled maintenance.</p>
<p>The status page is served separately from your application on Cloudflare's global edge network, so your incident communication does not depend on your production application remaining available.</p>
<p><strong>What OnlineOrNot does well:</strong></p>
<ul>
<li><strong>Monitoring and status pages share the same data</strong> - Reduce integration work and avoid conflicting service states.</li>
<li><strong>Public and private pages</strong> - Publish a customer-facing page or protect an internal page with a password or IP allowlist.</li>
<li><strong>Custom domains with automatic HTTPS</strong> - Use a domain such as <code>status.example.com</code> without managing certificates yourself.</li>
<li><strong>Unlimited subscribers on paid plans</strong> - Customers receive incident updates without per-subscriber fees.</li>
<li><strong>Automatic and manual incidents</strong> - Automate clear outages while keeping people in control of customer communication.</li>
<li><strong>Third-party components</strong> - Show the health of important providers alongside services you monitor directly.</li>
<li><strong>API and webhooks</strong> - Connect status-page workflows to your existing tooling.</li>
</ul>
<p><strong>Where it may not fit:</strong></p>
<ul>
<li>OnlineOrNot does not provide log aggregation or distributed tracing.</li>
<li>Atlassian may suit large organizations better when they require complex audience-specific pages or deeply integrated Atlassian approval workflows.</li>
</ul>
<p><strong>Pricing:</strong> A limited public status page is available on the free plan. Paid plans start at $12/month annually or $15/month monthly and include the first public status page. Additional public and private pages are priced separately.</p>
<p><strong>Best for:</strong> SaaS companies, developers, startups, and small teams that want monitoring and customer communication in one product.</p>
<h2>2. Better Stack</h2>
<p><a href="https://betterstack.com/status-page">Better Stack</a> combines status pages with uptime monitoring, on-call scheduling, incident management, logs, and other observability products.</p>
<p>It is a strong option when a status page is one part of a broader response workflow. The same vendor can detect an incident, alert the on-call engineer, and communicate with customers.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>Built-in monitoring and incident response</li>
<li>On-call schedules and escalation policies</li>
<li>Public and private status pages</li>
<li>Integrations with existing monitoring tools</li>
<li>A broader logs and observability platform</li>
</ul>
<p><strong>Tradeoffs:</strong> The wider platform may introduce more product and pricing complexity than a team needs for straightforward monitoring and customer communication.</p>
<p><strong>Best for:</strong> Teams that want monitoring, on-call response, logs, and status pages from one vendor.</p>
<h2>3. Instatus</h2>
<p><a href="https://instatus.com/">Instatus</a> is a status-page-first product that also offers uptime monitoring and incident-response features.</p>
<p>Its focus is creating polished customer-facing pages quickly. Teams can start with a free public page and add custom domains, faster monitoring, private pages, and more notification options on paid plans.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>Fast status-page setup</li>
<li>Polished customer-facing design</li>
<li>Built-in monitoring</li>
<li>Incident and maintenance communication</li>
<li>Free entry point</li>
</ul>
<p><strong>Tradeoffs:</strong> Custom domains, subscriber limits, private pages, monitoring frequency, and SSO depend on the selected plan. Teams requiring advanced browser or API monitoring may still need another product.</p>
<p><strong>Best for:</strong> Startups and software teams that prioritize a simple, status-page-first workflow.</p>
<h2>4. OpenStatus</h2>
<p><a href="https://www.openstatus.dev/">OpenStatus</a> is an open-source uptime monitoring and status page platform with a managed cloud service.</p>
<p>It gives developers a choice: use hosted plans to avoid operational work or self-host for more control. The product includes global monitoring, incidents, subscribers, themes, and status pages.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>Open-source codebase</li>
<li>Managed and self-hosted options</li>
<li>Monitoring and status pages together</li>
<li>Developer-focused workflow</li>
</ul>
<p><strong>Tradeoffs:</strong> Monitor intervals, regions, page counts, white-labeling, and access controls vary by hosted tier or add-on. Self-hosting makes your team responsible for availability and maintenance.</p>
<p><strong>Best for:</strong> Developers who want open-source software without giving up the option of managed hosting.</p>
<h2>5. StatusPal</h2>
<p><a href="https://www.statuspal.io/">StatusPal</a> is a dedicated status page and incident communication product. It supports public and private pages, custom domains, subscriber notifications, and integrations with monitoring and incident-management tools.</p>
<p>It is closer to the standalone Statuspage model than a monitoring-first product. That can be an advantage when your team already has detection and response tools it wants to keep.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>Dedicated incident communication workflow</li>
<li>Public and private pages</li>
<li>Custom branding and domains</li>
<li>Monitoring and incident-management integrations</li>
<li>Subscriber notifications</li>
</ul>
<p><strong>Tradeoffs:</strong> Teams buying both monitoring and a status page should compare the combined cost and integration work with an all-in-one option.</p>
<p><strong>Best for:</strong> Organizations that want a dedicated communication layer connected to their existing monitoring stack.</p>
<h2>6. UptimeRobot</h2>
<p><a href="https://uptimerobot.com/status-page/">UptimeRobot</a> is best known for uptime monitoring and also provides hosted status pages.</p>
<p>It is a practical starting point for personal projects and small teams that already use UptimeRobot monitors. Monitor state and uptime history can appear on a public page without connecting another vendor.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>Free way to start</li>
<li>Monitoring and status pages together</li>
<li>Familiar, straightforward interface</li>
<li>Public incident and maintenance communication</li>
</ul>
<p><strong>Tradeoffs:</strong> Check the current plan limits for custom domains, branding, subscribers, team access, and private pages. The free monitoring interval is more appropriate for low-risk projects than critical production services.</p>
<p><strong>Best for:</strong> Personal projects, small websites, and existing UptimeRobot customers who need a basic page.</p>
<h2>7. Uptime Kuma</h2>
<p><a href="https://github.com/louislam/uptime-kuma">Uptime Kuma</a> is a popular open-source monitoring tool with built-in status pages.</p>
<p>There is no SaaS subscription or per-monitor charge. You run the software on your own infrastructure and control its data and configuration.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>Free and open source</li>
<li>Many monitor types</li>
<li>Built-in public status pages</li>
<li>Large community</li>
<li>Full control over hosting and data</li>
</ul>
<p><strong>Tradeoffs:</strong> Your team owns hosting, upgrades, backups, security, TLS, access control, and notification delivery. A single self-hosted instance also provides less geographic verification than a managed multi-region monitoring service.</p>
<p><strong>Best for:</strong> Home labs, internal services, and technical teams comfortable operating their own monitoring stack.</p>
<h2>Free and open-source Atlassian Statuspage alternatives</h2>
<p>Searches for a "free Atlassian Statuspage alternative" often mix two different categories.</p>
<h3>Free hosted plans</h3>
<p>OnlineOrNot, Better Stack, Instatus, OpenStatus, and UptimeRobot offer ways to start without paying. These plans normally limit some combination of subscribers, monitors, team members, check frequency, private access, or custom domains.</p>
<p>A free hosted plan is useful when you want to test the incident workflow without operating infrastructure.</p>
<h3>Free open-source software</h3>
<p>Uptime Kuma and OpenStatus can be self-hosted. Other projects can be found in GitHub collections such as <a href="https://github.com/ivbeg/awesome-status-pages">awesome-status-pages</a>.</p>
<p>The software may be free, but operating it is not. Include server costs and the time required for:</p>
<ul>
<li>Upgrades and security patches</li>
<li>Backups and recovery testing</li>
<li>Email or SMS delivery</li>
<li>DNS and TLS configuration</li>
<li>Monitoring the monitoring system</li>
<li>Keeping the page available during primary infrastructure failures</li>
</ul>
<p>Reddit discussions about Statuspage alternatives often recommend self-hosting because the license costs nothing. That can be the right decision for a home lab or infrastructure-focused team. For a small SaaS, a managed page may cost less than the engineering time needed to run one reliably.</p>
<h2>How to choose an Atlassian Statuspage alternative</h2>
<p>Choose based on what happens during an outage, not only what the finished page looks like.</p>
<h3>1. Decide whether monitoring should update the page</h3>
<p>If you already have reliable monitoring, a standalone status page with strong integrations may be enough.</p>
<p>If you are buying both, a product with built-in monitoring can reduce setup and eliminate conflicting incident states. Look for the ability to automate component changes while preserving manual control over customer-facing messages.</p>
<h3>2. Keep the status page independent</h3>
<p>Your status page should not share a likely failure mode with the application it reports on.</p>
<p>For self-hosted products, consider the cloud account, region, DNS provider, deployment pipeline, and network dependencies. For hosted products, understand which dependencies remain under your control, especially custom-domain DNS.</p>
<h3>3. Compare the complete price</h3>
<p>Calculate the price for the configuration you expect to use, including:</p>
<ul>
<li>Status pages</li>
<li>Monitors and check frequency</li>
<li>Team members</li>
<li>Subscribers</li>
<li>Private access</li>
<li>Custom domains</li>
<li>Notification channels</li>
<li>SSO and audit logs</li>
<li>Data retention</li>
</ul>
<h3>4. Test the incident workflow</h3>
<p>Use the trial to rehearse a real incident:</p>
<ol>
<li>A monitor fails</li>
<li>The team receives an alert</li>
<li>An incident opens</li>
<li>Affected components change state</li>
<li>Subscribers receive an update</li>
<li>An operator adds context</li>
<li>The service recovers</li>
<li>The incident resolves and remains in history</li>
</ol>
<p>Also test scheduled maintenance and a manual incident. The interface your team uses under pressure matters more than the marketing page.</p>
<h3>5. Check access and subscriber requirements</h3>
<p>Public pages work for most customer-facing services. Internal systems and enterprise customers may require a password, IP allowlist, email authentication, or SSO.</p>
<p>Do not assume every product uses the same definition of "private page." Confirm that the access method and subscriber limits fit your actual audience.</p>
<h2>How to migrate from Atlassian Statuspage</h2>
<p>A migration does not need to interrupt customer communication.</p>
<ol>
<li><strong>Audit the existing page.</strong> List components, groups, incidents, maintenance windows, metrics, subscribers, integrations, and access controls.</li>
<li><strong>Create the replacement page.</strong> Match the current component structure before trying to improve it.</li>
<li><strong>Connect monitoring.</strong> Map monitors to components and decide which incidents can be automated safely.</li>
<li><strong>Configure branding and access.</strong> Set the logo, colors, custom domain, password, IP allowlist, or SSO requirements.</li>
<li><strong>Test notifications.</strong> Confirm that incidents, updates, maintenance, and resolutions reach the intended subscribers.</li>
<li><strong>Rehearse an outage.</strong> Test both automatic and manual publishing with the people who will operate the page.</li>
<li><strong>Move the custom domain.</strong> Lower DNS TTL in advance, change the record, and verify HTTPS before announcing the migration.</li>
<li><strong>Retire the old page carefully.</strong> Keep a record of historical incidents and redirect old links where possible.</li>
</ol>
<p>For OnlineOrNot, start with the guide to <a href="/docs/how-to/status-pages/create">create a status page</a>, then <a href="/docs/how-to/status-pages/link-uptime-monitoring">link uptime monitoring to the page</a>.</p>
<h2>FAQ</h2>
<h3>What are the best alternatives to Atlassian Statuspage?</h3>
<p>OnlineOrNot is a strong choice for teams that want monitoring and status pages together. Better Stack suits broader incident-response requirements, Instatus and StatusPal focus on status-page workflows, OpenStatus offers open-source and hosted options, UptimeRobot provides a simple free starting point, and Uptime Kuma is popular for self-hosting.</p>
<h3>Is Atlassian Statuspage free?</h3>
<p>Atlassian offers a limited free plan for public status pages. Paid public plans increase limits and customization, while private and audience-specific pages use separate pricing. Confirm current features and prices on Atlassian's official site before purchasing.</p>
<h3>Who competes with Atlassian Statuspage?</h3>
<p>Competitors include OnlineOrNot, Better Stack, Instatus, OpenStatus, StatusPal, UptimeRobot, Status.io, Hyperping, and open-source projects such as Uptime Kuma and Cachet.</p>
<h3>What is the best free Atlassian Statuspage alternative?</h3>
<p>OnlineOrNot, Better Stack, Instatus, OpenStatus, and UptimeRobot all provide hosted ways to start for free, with different limits. Uptime Kuma is a strong free self-hosted option when your team is prepared to operate it.</p>
<p>The best choice depends on whether you need monitoring, a custom domain, private access, faster checks, multiple team members, or more subscribers.</p>
<h3>What is the best open-source Statuspage alternative?</h3>
<p>Uptime Kuma is a popular option when you want monitoring and a public page in one self-hosted product. OpenStatus combines an open-source codebase with managed hosting. Cachet is another option when you want a dedicated status page and already have monitoring.</p>
<h3>Should monitoring and status pages use the same platform?</h3>
<p>They do not have to, but combining them reduces integration work and makes automatic component and incident updates easier.</p>
<p>The status page itself should still be delivered independently from the application being monitored. A shared product workflow is useful; shared failure infrastructure is not.</p>
<h3>How much does an OnlineOrNot status page cost?</h3>
<p>OnlineOrNot includes a limited public status page on its free Hobby plan. Paid plans start at $12/month with annual billing or $15/month billed monthly and include the first public status page. Paid plans include unlimited email subscribers.</p>
<h3>When should I keep Atlassian Statuspage?</h3>
<p>Keep Atlassian Statuspage when it already meets your requirements, your team depends on Atlassian integrations, or you need mature audience-specific communication and enterprise governance.</p>
<p>Switching tools introduces migration and training work. An alternative should solve a concrete pricing, monitoring, access, or workflow problem rather than simply offering a different interface.</p>
<hr>
<h2>Create a status page connected to monitoring</h2>
<p>OnlineOrNot gives customers one place to check service health while helping your team detect outages and publish updates from the same product.</p>
<p><a href="/get-started">Start a 14-day trial</a> or learn more about <a href="/status-pages">OnlineOrNot status pages</a>.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[10 best status page tools for 2026 (hosted and open source)]]></title>
            <link>https://onlineornot.com/status-page-tools</link>
            <guid>https://onlineornot.com/status-page-tools</guid>
            <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Status page tools give customers one place to check whether your service is working, what went wrong, and when they can expect another update. The best status page software also connects that communication to monitoring, incident response, and subscriber notifications.</p>
<p>The page itself is the easy part. The differences show up during a real incident: whether monitoring can update the page automatically, how quickly your team can publish an explanation, who gets notified, and whether the page stays available when your own infrastructure does not.</p>
<p>This guide compares the status page tools worth considering in 2026, with a focus on practical tradeoffs: built-in monitoring, incident workflows, custom domains, subscribers, private pages, pricing, and who each tool is best for.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#quick-comparison-table">Quick comparison table</a></li>
<li><a href="#what-is-a-status-page-tool">What is a status page tool?</a></li>
<li><a href="#best-for-most-teams-onlineornot">Best for most teams: OnlineOrNot</a></li>
<li><a href="#best-for-large-enterprises-atlassian-statuspage">Best for large enterprises: Atlassian Statuspage</a></li>
<li><a href="#best-for-incident-response-better-stack">Best for incident response: Better Stack</a></li>
<li><a href="#best-status-page-first-option-instatus">Best status-page-first option: Instatus</a></li>
<li><a href="#best-free-starting-point-uptimerobot">Best free starting point: UptimeRobot</a></li>
<li><a href="#best-open-source-and-self-hosted-status-page-tools">Best open-source and self-hosted status page tools</a></li>
<li><a href="#other-hosted-status-page-tools-to-consider">Other hosted status page tools to consider</a></li>
<li><a href="#how-to-choose-a-status-page-tool">How to choose a status page tool</a></li>
<li><a href="#faq">FAQ</a></li>
</ul>
<h2>Quick comparison table</h2>
<p>| Tool | Free option | Monitoring included | Custom domain | Private pages | Best for |
|------|-------------|---------------------|---------------|---------------|----------|
| <strong>OnlineOrNot</strong> | Yes | Yes | Yes | Yes | SaaS teams wanting monitoring and status pages together |
| <strong>Atlassian Statuspage</strong> | Yes | Via integrations | Paid plans | Yes | Large organizations in the Atlassian ecosystem |
| <strong>Better Stack</strong> | Yes | Yes | Yes | Yes | Teams wanting monitoring, on-call, and incident response |
| <strong>Instatus</strong> | Yes | Yes | Paid plans | Higher tiers | Teams prioritizing a polished status-page workflow |
| <strong>UptimeRobot</strong> | Yes | Yes | Paid plans | Plan-dependent | Personal projects and basic monitoring |
| <strong>Uptime Kuma</strong> | Self-hosted | Yes | DIY | DIY | Home labs and teams comfortable self-hosting |
| <strong>OpenStatus</strong> | Yes / open source | Yes | Paid hosted plans or DIY | Paid add-on or DIY | Developers wanting open source with hosted plans |
| <strong>Cachet</strong> | Self-hosted | Via API | DIY | DIY | Teams wanting a dedicated open-source status page |
| <strong>Status.io</strong> | Trial | Via integrations | Yes | Higher tiers | Organizations with complex communication needs |
| <strong>Hyperping</strong> | Yes | Yes | Paid plans | Paid plans | Small teams wanting monitoring, on-call, and status pages |</p>
<p>Features and plan limits change. Use this table as a shortlist, then confirm that the current plan includes the subscriber count, access controls, and notification channels you need.</p>
<h2>What is a status page tool?</h2>
<p>A status page tool publishes the current health of your website, API, application, or infrastructure. It normally displays components, active incidents, planned maintenance, uptime history, and a way for users to subscribe to updates.</p>
<p>A useful status page answers four questions:</p>
<ol>
<li><strong>Is the service working?</strong></li>
<li><strong>Which components are affected?</strong></li>
<li><strong>Does the team know about the problem?</strong></li>
<li><strong>When will the next update arrive?</strong></li>
</ol>
<p>There are three common approaches.</p>
<h3>Standalone status page software</h3>
<p>A standalone tool focuses on incident communication. Your team posts updates manually or connects monitoring and incident-management tools through integrations.</p>
<p>This works well when the people detecting incidents and the people communicating with customers use different systems.</p>
<h3>Status pages with built-in monitoring</h3>
<p>These tools monitor your services and can automatically open or resolve incidents when a check changes state.</p>
<p>This reduces the chance that your team fixes an outage but forgets to update the status page. It also means one product can own uptime checks, alerts, incident history, and customer communication.</p>
<h3>Self-hosted status pages</h3>
<p>Open-source tools give you control over hosting, data, and customization. The tradeoff is operational responsibility: deployment, upgrades, backups, email delivery, security, and availability are now your problem.</p>
<p>A status page should be independent from the application it reports on. If both fail together, customers lose the place they were supposed to check for updates.</p>
<h2>Best for most teams: OnlineOrNot</h2>
<p><a href="/status-pages">OnlineOrNot</a> combines hosted status pages with uptime, API, browser, and heartbeat monitoring.</p>
<p>It is built for developers and small software teams that want monitoring to drive incident communication without adopting a larger enterprise incident-management suite.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Monitoring and status pages share the same data</strong> - Connect an uptime check or heartbeat to a component, then automatically create and resolve incidents when its state changes.</li>
<li><strong>Public and private pages</strong> - Publish a customer-facing page or protect an internal page with a shared password or IP allowlist.</li>
<li><strong>Custom domains and automatic HTTPS</strong> - Host the page on your own domain without managing certificates yourself.</li>
<li><strong>Unlimited subscribers</strong> - Customers can subscribe to incident updates by email without creating an OnlineOrNot account. RSS feeds are also available.</li>
<li><strong>Third-party components</strong> - Display the health of dependencies alongside services you monitor directly, so customers get more context during provider outages.</li>
<li><strong>Manual control when it matters</strong> - Publish manual incidents, updates, retrospective incidents, and scheduled maintenance alongside automated incidents.</li>
<li><strong>API and webhooks</strong> - Manage pages and incident workflows programmatically instead of relying only on the dashboard.</li>
</ul>
<p><strong>What it does not do:</strong></p>
<ul>
<li>No log aggregation or distributed tracing</li>
<li>Less suitable than a dedicated enterprise platform if you need complex audience-specific pages or heavily customized approval workflows</li>
</ul>
<p><strong>Pricing:</strong> A limited public status page is available on the free plan. Paid plans start at $12/month with annual billing or $15/month billed monthly and include the first public status page. Additional public and private pages are priced separately.</p>
<p><strong>Best for:</strong> SaaS companies, developers, startups, and small teams that want uptime monitoring and customer communication in one product.</p>
<h2>Best for large enterprises: Atlassian Statuspage</h2>
<p><a href="https://www.atlassian.com/software/statuspage">Atlassian Statuspage</a> is one of the best-known dedicated status page products. It offers public, private, and audience-specific pages, with integrations across the Atlassian ecosystem.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Mature incident communication</strong> - Components, incidents, scheduled maintenance, metrics, and subscriber notifications are established parts of the product.</li>
<li><strong>Several page types</strong> - Public, private, and audience-specific pages cover customer-facing and internal use cases.</li>
<li><strong>Atlassian integrations</strong> - A natural fit for teams already coordinating incidents in Jira Service Management, Opsgenie, and other Atlassian products.</li>
<li><strong>Enterprise controls</strong> - Higher tiers support the governance and access requirements larger organizations expect.</li>
<li><strong>Free public-page entry point</strong> - The free plan is enough to test the basic workflow with limited subscribers and team members.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>Monitoring is not the main product</strong> - Most teams connect separate monitoring or incident-management systems to automate component status.</li>
<li><strong>Costs rise quickly</strong> - Subscriber limits, team members, private pages, and enterprise requirements can make it expensive.</li>
<li><strong>More workflow than small teams need</strong> - A two-person SaaS team may not benefit from the extra administration.</li>
</ul>
<p><strong>Pricing:</strong> Free public plan available. Paid public plans start at $29/month, with private and audience-specific pages using separate pricing.</p>
<p><strong>Best for:</strong> Larger organizations, especially teams already using Atlassian for service management and incident response.</p>
<p>For a closer look at migration tradeoffs, compare the <a href="/atlassian-statuspage-alternative">best Atlassian Statuspage alternatives</a>.</p>
<h2>Best for incident response: Better Stack</h2>
<p><a href="https://betterstack.com/status-page">Better Stack</a> combines status pages with uptime monitoring, on-call scheduling, incident management, logs, and observability products.</p>
<p>This makes it a strong choice when a status page is one part of a broader response workflow rather than a standalone communication page.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Incident response built in</strong> - Monitoring alerts, escalation policies, on-call schedules, and status updates can live in one platform.</li>
<li><strong>Built-in uptime monitoring</strong> - The same system that detects a problem can update component status and start the incident workflow.</li>
<li><strong>Custom domains on the free plan</strong> - Small teams can publish a branded page without immediately paying for a status-page-only plan.</li>
<li><strong>Private status pages</strong> - SSO, password, and IP protection support internal teams and enterprise customers.</li>
<li><strong>Broad integrations</strong> - Connect existing monitoring and observability tools if Better Stack is not your only source of incident data.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>A broad platform can add complexity</strong> - If you only want a simple status page, logs, tracing, and on-call features may be more than you need.</li>
<li><strong>Pricing depends on the wider stack</strong> - Compare the combined cost of monitors, responders, data, and status-page features you will use.</li>
</ul>
<p><strong>Pricing:</strong> Free plan available. Paid pricing depends on monitoring, incident management, team, and observability usage.</p>
<p><strong>Best for:</strong> Teams that want status pages, monitoring, and on-call incident response under one vendor.</p>
<h2>Best status-page-first option: Instatus</h2>
<p><a href="https://instatus.com/">Instatus</a> is a status-page-first product that has expanded into uptime monitoring, alerts, and on-call workflows.</p>
<p>It suits teams that care most about quickly publishing a polished page and want basic monitoring included rather than connected from another vendor.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Fast setup</strong> - The product is centered on creating and maintaining status pages rather than configuring a broad observability stack.</li>
<li><strong>Useful free plan</strong> - The free tier includes a public status page, monitoring, team members, and a limited subscriber count.</li>
<li><strong>Built-in monitoring</strong> - Paid tiers provide faster checks and more monitors, reducing the need for a separate uptime product.</li>
<li><strong>Incident notifications</strong> - Email, SMS, calls, and on-call features are available as you move through the plans.</li>
<li><strong>Polished customer-facing pages</strong> - A strong choice when design and simple communication workflows are the priority.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>Important features are plan-dependent</strong> - Custom domains, private pages, SAML SSO, subscriber limits, and monitor counts depend on the tier.</li>
<li><strong>Monitoring depth is lighter than a monitoring-first platform</strong> - Teams needing browser checks, richer API assertions, or broader observability may still need another tool.</li>
</ul>
<p><strong>Pricing:</strong> Free plan available. Paid plans add faster monitoring, custom domains, more subscribers, and private pages.</p>
<p><strong>Best for:</strong> Startups and software teams wanting a polished, status-page-first product with monitoring included.</p>
<h2>Best free starting point: UptimeRobot</h2>
<p><a href="https://uptimerobot.com/status-page/">UptimeRobot</a> combines its well-known uptime monitoring with hosted status pages for sharing service health, incidents, and maintenance updates.</p>
<p>It is a practical starting point for personal projects and small teams that already use UptimeRobot monitors and want a basic public page without adopting a separate incident communication product.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Free way to start</strong> - Create a basic status page alongside free uptime monitors.</li>
<li><strong>Monitoring is already connected</strong> - Display monitor state and uptime history without wiring together separate products.</li>
<li><strong>Simple public communication</strong> - Share incidents and planned maintenance from a familiar dashboard.</li>
<li><strong>Customizable pages</strong> - Paid plans add more branding and status page features.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>Less incident workflow depth</strong> - Dedicated status page and incident-response products offer richer subscriber, access-control, and collaboration features.</li>
<li><strong>Important features depend on the plan</strong> - Check the current limits for custom domains, branding, subscribers, and private access.</li>
<li><strong>Free monitoring is slower</strong> - The free monitoring interval is better suited to personal and low-risk projects than critical production services.</li>
</ul>
<p><strong>Pricing:</strong> Free plan available. Paid plans add faster monitoring and more advanced status page features.</p>
<p><strong>Best for:</strong> Personal projects, small websites, and existing UptimeRobot users who need a straightforward public page.</p>
<h2>Best open-source and self-hosted status page tools</h2>
<p>Open-source status page software gives you control over deployment, data, and customization. It also makes your team responsible for availability, upgrades, security, backups, and notification delivery.</p>
<h3>Uptime Kuma</h3>
<p><a href="https://github.com/louislam/uptime-kuma">Uptime Kuma</a> is a popular open-source monitoring tool with built-in status pages.</p>
<p>It has a friendly interface, supports many monitor types, and does not charge per monitor. The tradeoff is that your team runs the infrastructure behind both monitoring and incident communication.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Free and open source</strong> - There is no SaaS subscription or per-monitor fee.</li>
<li><strong>Monitoring included</strong> - Supports HTTP, TCP, DNS, Docker, ping, keyword, and other monitor types.</li>
<li><strong>Simple public status pages</strong> - Group monitors and show their current state and history without adopting another tool.</li>
<li><strong>Data control</strong> - Useful for home labs, private networks, and teams that prefer to host their own operational data.</li>
<li><strong>Large community</strong> - It is widely used and actively maintained.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>You own reliability</strong> - Hosting, upgrades, backups, security, and notification delivery are your responsibility.</li>
<li><strong>Independence takes planning</strong> - Hosting Uptime Kuma beside the application it monitors creates a shared failure point.</li>
<li><strong>Advanced communication workflows are limited</strong> - Dedicated products usually offer richer subscriber management, private access, incident editing, and enterprise controls.</li>
<li><strong>Custom domains and access control are DIY</strong> - Expect to configure your own DNS, reverse proxy, TLS, and authentication where required.</li>
</ul>
<p><strong>Pricing:</strong> Free software. You pay for infrastructure and the time required to operate it.</p>
<p><strong>Best for:</strong> Self-hosters, home labs, internal services, and technical teams comfortable maintaining their own monitoring stack.</p>
<h3>OpenStatus</h3>
<p><a href="https://www.openstatus.dev/">OpenStatus</a> is an open-source uptime monitoring and status page platform with a managed cloud service.</p>
<p>The hosted product combines global monitoring, incidents, subscribers, custom themes, and status pages. Self-hosting gives technical teams more control, while paid hosted plans remove much of the operational work.</p>
<p>The main tradeoff is plan structure: monitor intervals, region coverage, page counts, white-labeling, and access controls vary by tier or add-on.</p>
<p><strong>Best for:</strong> Developers who value open source but still want the option of a managed service.</p>
<h3>Cachet</h3>
<p><a href="https://cachethq.io/">Cachet</a> is a dedicated open-source status page system. It supports components, incidents, planned maintenance, metrics, multiple languages, and API-based automation.</p>
<p>Unlike Uptime Kuma, Cachet is status-page-first rather than monitoring-first. You normally connect your monitoring through its API or other automation.</p>
<p>You are responsible for hosting, upgrades, email delivery, backups, and keeping the page available during incidents.</p>
<p><strong>Best for:</strong> Technical teams that want a self-hosted, dedicated status page and already have monitoring.</p>
<h2>Other hosted status page tools to consider</h2>
<h3>Status.io</h3>
<p><a href="https://status.io/">Status.io</a> is a hosted incident communication platform aimed at organizations that need detailed components, subscriber notifications, metrics, automation, and private-page controls.</p>
<p>It supports integrations with existing monitoring tools rather than trying to replace the rest of your observability stack. White-labeling, access controls, audit trails, and subscriber compliance tools make it more relevant to larger organizations.</p>
<p>There is no permanent free plan, and entry pricing is higher than lightweight status page products.</p>
<p><strong>Best for:</strong> Organizations with complex public or private incident communication requirements.</p>
<h3>Hyperping</h3>
<p><a href="https://hyperping.com/">Hyperping</a> combines uptime monitoring, status pages, incident management, and on-call scheduling.</p>
<p>Its free plan is useful for basic monitoring and a simple page. Paid plans add faster checks, custom domains, more subscribers, browser checks, and escalation policies.</p>
<p>Pricing grows with monitors, seats, and page requirements, so compare the full configuration rather than only the starting plan.</p>
<p><strong>Best for:</strong> Small teams wanting a modern monitoring and status page product with on-call features.</p>
<h2>How to choose a status page tool</h2>
<p>The best tool is not necessarily the one with the longest feature list. It is the one your team can update accurately while an incident is unfolding.</p>
<h3>1. Decide whether monitoring should update the page</h3>
<p>Manual updates give your team complete control, but someone has to remember to publish them.</p>
<p>Automatic updates work well for clear failures such as an unreachable API or a missed heartbeat. They are less useful when an incident needs context before customers see it.</p>
<p>A practical setup combines both:</p>
<ul>
<li>Monitoring changes the affected component state</li>
<li>Automation creates or resolves straightforward incidents</li>
<li>A human writes context, scope, workarounds, and the next update time</li>
<li>Your team can pause or override automation when needed</li>
</ul>
<p>If you already have reliable monitoring, a standalone status page with good integrations may be enough. If you are buying both, a monitoring-first tool usually requires less setup.</p>
<h3>2. Keep the status page outside your own infrastructure</h3>
<p>A status page is most valuable when your application is unavailable. Hosting both in the same account, region, cluster, or deployment pipeline creates a shared failure mode.</p>
<p>For self-hosted pages, use separate infrastructure and test what happens when your primary systems fail. For hosted products, consider which infrastructure and DNS dependencies the page shares with your application.</p>
<h3>3. Compare subscriber limits and notification channels</h3>
<p>Some tools look inexpensive until you add customers.</p>
<p>Check the limits and extra costs for:</p>
<ul>
<li>Email subscribers</li>
<li>SMS or phone notifications</li>
<li>Webhooks</li>
<li>Slack and Microsoft Teams</li>
<li>RSS or Atom feeds</li>
<li>Per-component subscriptions</li>
<li>Subscriber imports and exports</li>
</ul>
<p>If your status page is public but nobody sees updates, it only solves half the problem.</p>
<h3>4. Choose the right access model</h3>
<p>Public pages work for most customer-facing products. Internal services and enterprise accounts may need more control.</p>
<p>Common access options include:</p>
<ul>
<li>Shared password</li>
<li>IP allowlist</li>
<li>Email authentication</li>
<li>SAML or OIDC SSO</li>
<li>Separate pages for specific customers or audiences</li>
</ul>
<p>Do not assume that a tool offering “private pages” supports your preferred authentication method. These features are also frequently limited to expensive plans.</p>
<h3>5. Test the incident workflow, not just the page design</h3>
<p>A beautiful page does not help if publishing an update takes ten steps.</p>
<p>During a trial, rehearse an incident from start to finish:</p>
<ol>
<li>A monitor fails</li>
<li>The on-call person is alerted</li>
<li>An incident is opened</li>
<li>Affected components change state</li>
<li>Subscribers receive an update</li>
<li>Your team publishes progress</li>
<li>The service recovers</li>
<li>The incident is resolved and kept in history</li>
</ol>
<p>Test scheduled maintenance and retrospective incidents too. The dashboard used under pressure matters more than the marketing screenshot.</p>
<h3>6. Calculate the full cost</h3>
<p>Status page pricing can depend on more than the number of pages.</p>
<p>Include:</p>
<ul>
<li>Monitors and check frequency</li>
<li>Team members or responders</li>
<li>Subscribers</li>
<li>Public, private, and audience-specific pages</li>
<li>SMS and phone credits</li>
<li>Custom domains and white-labeling</li>
<li>SSO and audit logs</li>
<li>Data retention</li>
</ul>
<p>A free page can be ideal for a small project. For a growing SaaS, predictable subscriber and team pricing often matters more than the cheapest entry tier.</p>
<h2>FAQ</h2>
<h3>What is the best status page tool?</h3>
<p>For most small SaaS teams, the best status page tool combines reliable monitoring, automatic incident updates, manual control, custom domains, and subscriber notifications without enterprise complexity.</p>
<p>OnlineOrNot is a good fit for that use case. Atlassian Statuspage is strong for large organizations, Better Stack is strong for incident response, Instatus is a polished status-page-first option, and Uptime Kuma is a good self-hosted choice.</p>
<h3>Is Atlassian Statuspage free?</h3>
<p>Yes. Atlassian Statuspage has a free public-page plan with limited subscribers, team members, components, and metrics. Paid public plans add higher limits and more customization, while private and audience-specific pages use separate pricing.</p>
<h3>What should a status page include?</h3>
<p>A useful status page should include current service health, affected components, active incidents, scheduled maintenance, incident history, uptime history, and a way to subscribe to updates.</p>
<p>During an incident, every update should explain what is affected, what the team is doing, and when customers should expect the next update.</p>
<h3>How do you manage a status page?</h3>
<p>Connect the page to monitoring for component health, assign someone to own customer communication, and use a repeatable incident update cadence. During an outage, publish the affected components, customer impact, current response, and next update time. Resolve the incident when service recovers, then preserve the timeline for future reference.</p>
<p>Schedule maintenance in advance, review subscriber and access settings regularly, and rehearse the full publishing workflow before a real incident.</p>
<h3>Are free status page tools good enough?</h3>
<p>Free status page tools are often good enough for personal projects, open-source projects, and small products with a limited audience.</p>
<p>A paid plan becomes worthwhile when you need a custom domain, more subscribers, faster monitoring, private access, multiple team members, SMS notifications, SSO, audit logs, or guaranteed support.</p>
<h3>Should I self-host a status page?</h3>
<p>Self-host if you need full control, have strict data requirements, or already operate the infrastructure needed to keep it reliable.</p>
<p>Use a hosted service if you want the page to remain independent from your application and do not want to manage upgrades, email delivery, TLS, backups, and availability during an outage.</p>
<h3>What is the difference between a public and private status page?</h3>
<p>A public status page is available to anyone and is normally used for customer-facing services. A private status page restricts access using a password, IP allowlist, email authentication, or SSO and is normally used for internal systems or specific enterprise customers.</p>
<h3>Should status pages update automatically?</h3>
<p>Automate clear component changes and routine incidents, but keep a human in control of customer communication.</p>
<p>Monitoring can say that an API is unavailable. It cannot always explain which customers are affected, whether data is safe, what workaround exists, or when the next meaningful update will arrive.</p>
<h3>How is a status page different from uptime monitoring?</h3>
<p><a href="/uptime-monitoring-tools">Uptime monitoring</a> detects whether a website, API, or service is working and alerts your team when it fails.</p>
<p>A status page communicates service health and incident updates to customers or internal stakeholders. Many modern tools combine both, but they solve different parts of the incident: detection and communication.</p>
<hr>
<h2>Related guides</h2>
<ul>
<li>Learn how OnlineOrNot's <a href="/status-pages">status page software</a> connects incidents to monitoring.</li>
<li>Follow the guide to <a href="/docs/how-to/status-pages/create">create a status page</a>.</li>
<li>Compare the <a href="/uptime-monitoring-tools">best uptime monitoring tools</a> if detection is your main priority.</li>
<li>Compare <a href="/synthetic-monitoring-tools">synthetic monitoring tools</a> for browser journeys and advanced API checks.</li>
</ul>
<p>A status page earns trust when it is accurate, available, and updated before customers need to ask what is happening.</p>
<p>If you want monitoring and incident communication in one place, <a href="/get-started">OnlineOrNot</a> includes uptime checks, automatic incident updates, custom domains, public and private pages, and unlimited subscribers without requiring a full enterprise observability platform.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from May 2026]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2026-may</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2026-may</guid>
            <pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Last month I focused on adding Telegram alerts to OnlineOrNot, making scripted browser checks generally available, and making it clearer to understand why checks fail.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-alerts">Features for Alerts</a>
<ul>
<li><a href="#telegram-alerting">Telegram alerting</a></li>
</ul>
</li>
<li><a href="#features-for-browser-checks">Features for Browser Checks</a>
<ul>
<li><a href="#scripted-browser-checks">Scripted browser checks</a></li>
<li><a href="#better-failure-details">Better failure details</a></li>
<li><a href="#more-reliable-browser-checks">More reliable browser checks</a></li>
</ul>
</li>
<li><a href="#features-for-the-dashboard">Features for the Dashboard</a>
<ul>
<li><a href="#better-check-results-explorer">Better check results explorer</a></li>
</ul>
</li>
<li><a href="#features-for-status-pages">Features for Status Pages</a>
<ul>
<li><a href="#better-custom-domain-handling">Better custom domain handling</a></li>
</ul>
</li>
</ul>
</li>
</ul>
<h2>What's new</h2>
<h3>Features for Alerts</h3>
<h4>Telegram alerting</h4>
<p>OnlineOrNot now supports Telegram alerts.</p>
<p>You can connect Telegram from the dashboard and receive alerts for uptime checks and heartbeat checks, alongside the existing Slack, Discord, email, and webhook options.</p>
<p>If your team already coordinates incidents in Telegram, OnlineOrNot can now send downtime and recovery alerts there too. It supports private chats, groups/supergroups, and channels. See the <a href="/docs/how-to/alerts/telegram">docs</a> to get started.</p>
<h3>Features for Browser Checks</h3>
<h4>Scripted browser checks</h4>
<p>OnlineOrNot now supports scripted browser checks via Playwright.</p>
<p>Regular uptime checks are still the right default for APIs, health endpoints, and simple pages. Browser checks are for the flows where "the URL returned HTTP 200" isn't enough.</p>
<p>They're useful for monitoring things like:</p>
<ul>
<li>Login flows</li>
<li>Checkout and payment pages</li>
<li>Dashboards behind JavaScript-heavy frontends</li>
<li>Forms and customer portals</li>
<li>"The app loads, but the important action is broken" failures</li>
</ul>
<p>Scripted browser checks are powered by Playwright scripts, so they're best suited for technical teams right now. If you already know the flow you care about, you can script it directly, draft it with an LLM, or use this as a reason to finally turn your most painful manual smoke test into monitoring.</p>
<p>You can <a href="https://onlineornot.com/app/browser-checks/add">create a scripted browser check here</a>.</p>
<h4>Better failure details</h4>
<p>Scripted browser checks also include clearer failure details when something goes wrong.</p>
<p>Instead of only seeing that a browser check failed, alerts can now include more useful context about the failure, helping you understand what broke faster.</p>
<h3>Features for the Dashboard</h3>
<h4>Better check results explorer</h4>
<p>The check results explorer has been improved to make recent check activity easier to inspect.</p>
<p>This should make it simpler to understand what happened during a failure, retry, or recovery.</p>
<h3>Features for Status Pages</h3>
<h4>Better custom domain handling</h4>
<p>Status page custom domains are now easier to manage and more reliable to visit.</p>
<p>OnlineOrNot now handles custom domains consistently whether you enter <code>example.com</code>, <code>https://example.com</code>, or include a trailing slash. Dashboard links now open the right public status page URL, and visitors see clearer messages if a custom-domain status page is unavailable.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Best Pingdom alternatives for 2026 (free and paid)]]></title>
            <link>https://onlineornot.com/best-pingdom-alternatives</link>
            <guid>https://onlineornot.com/best-pingdom-alternatives</guid>
            <pubDate>Fri, 29 May 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Pingdom is one of the best-known website monitoring products. It has been around for years, is now part of SolarWinds, and remains a familiar choice for teams that want uptime monitoring, page speed monitoring, and transaction checks.</p>
<p>But it is not the only option. If you are comparing Pingdom alternatives in 2026, you are probably looking for one of three things: simpler pricing, faster uptime checks, or monitoring that covers more than public URLs.</p>
<p>This guide compares practical Pingdom alternatives for engineering teams, startups, agencies, and developers.</p>
<p>If you specifically want to compare OnlineOrNot and Pingdom side by side, read the <a href="/pingdom-alternative">Pingdom alternative comparison</a>.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#quick-comparison-table">Quick comparison table</a></li>
<li><a href="#best-overall-pingdom-alternative-onlineornot">Best overall Pingdom alternative: OnlineOrNot</a></li>
<li><a href="#best-free-starting-point-uptimerobot">Best free starting point: UptimeRobot</a></li>
<li><a href="#best-for-incident-response-better-stack">Best for incident response: Better Stack</a></li>
<li><a href="#best-self-hosted-option-uptime-kuma">Best self-hosted option: Uptime Kuma</a></li>
<li><a href="#best-enterprise-option-datadog-synthetic-monitoring">Best enterprise option: Datadog Synthetic Monitoring</a></li>
<li><a href="#other-pingdom-alternatives-to-consider">Other Pingdom alternatives to consider</a></li>
<li><a href="#how-to-choose-a-pingdom-alternative">How to choose a Pingdom alternative</a></li>
<li><a href="#faq">FAQ</a></li>
</ul>
<h2>Quick comparison table</h2>
<p>| Tool | Check frequency | Free plan | Browser checks | Cron monitoring | Status pages | Best for |
|------|----------------|-----------|----------------|-----------------|--------------|----------|
| <strong>OnlineOrNot</strong> | 30 seconds | Yes | Yes | Yes | Yes | Engineering teams that want uptime, browser checks, cron monitoring, and status pages together |
| <strong>UptimeRobot</strong> | 5 min free, 1 min paid | Yes | Limited | Heartbeats | Yes | Simple uptime checks and hobby projects |
| <strong>Better Stack</strong> | 30 seconds | Yes | Yes | Heartbeats | Yes | Teams that want monitoring plus incident response |
| <strong>Uptime Kuma</strong> | Configurable | Self-hosted | Limited | Push monitors | Yes | Self-hosters and private infrastructure |
| <strong>Datadog Synthetic Monitoring</strong> | 1 minute+ | No | Yes | No | No | Enterprises already using Datadog |
| <strong>StatusCake</strong> | 5 min free, 30s on higher tiers | Yes | Yes | No | Yes | Website monitoring plus page speed and domains |
| <strong>Uptime.com</strong> | 1 minute | Trial | Yes | Yes | Yes | Larger teams with broad monitoring needs |
| <strong>Checkly</strong> | Usage-based | Trial/free developer usage | Yes | No | No | Code-first synthetic monitoring |
| <strong>Hyperping</strong> | Plan-dependent | Yes | No | Heartbeats | Yes | Lightweight uptime monitoring and status pages |</p>
<h2>Best overall Pingdom alternative: OnlineOrNot</h2>
<p><a href="/">OnlineOrNot</a> is built for developers and small teams that want practical production monitoring without buying a full enterprise observability suite.</p>
<p>It covers the core workflows teams usually assemble across several tools: uptime checks, API monitoring, Playwright browser checks, cron job monitoring, alerts, and hosted status pages.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>
<p><strong>30-second uptime checks</strong> - OnlineOrNot checks production services quickly enough for SaaS apps, APIs, and customer-facing workflows where several minutes of downtime matters.</p>
</li>
<li>
<p><strong>Multi-location verification</strong> - Failures are verified from multiple regions before alerting, which helps reduce false positives caused by regional network issues.</p>
</li>
<li>
<p><strong>Playwright browser checks</strong> - If your login, signup, checkout, or dashboard flow breaks while the homepage still returns <code>200 OK</code>, browser checks can catch it.</p>
</li>
<li>
<p><strong>Cron job monitoring</strong> - Backups, billing jobs, queue workers, and data pipelines can check in on a schedule. If they miss their window, OnlineOrNot alerts you.</p>
</li>
<li>
<p><strong>Status pages included</strong> - Hosted status pages are included with paid plans, so you can communicate incidents without adding another vendor.</p>
</li>
<li>
<p><strong>Team-friendly pricing</strong> - Pro starts at $15/month for 10 monitors and includes unlimited team members.</p>
</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>No APM or distributed tracing</li>
<li>No log aggregation</li>
<li>Not designed to replace Datadog, New Relic, or a full observability platform</li>
</ul>
<p><strong>Best for:</strong> SaaS teams, startups, agencies, and engineering teams that want uptime checks, browser checks, cron monitoring, and status pages in one lightweight product.</p>
<h2>Best free starting point: UptimeRobot</h2>
<p><a href="https://uptimerobot.com/">UptimeRobot</a> is a popular Pingdom alternative for teams that want simple uptime monitoring at a low starting cost.</p>
<p>It is especially common for personal projects, side projects, and small sites where the monitoring setup needs to be quick and inexpensive.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Generous free plan</strong> - Useful for personal projects and low-risk sites.</li>
<li><strong>Fast setup</strong> - Add a URL, choose alert channels, and start monitoring.</li>
<li><strong>Low-cost paid plans</strong> - A reasonable fit when you mostly need basic HTTP checks.</li>
<li><strong>Familiar product</strong> - Many developers have used UptimeRobot before.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>The free tier checks less frequently than many paid production setups need.</li>
<li>It is less focused on browser-based user journeys.</li>
<li>Teams that need cron monitoring, status pages, alert routing, and browser checks may outgrow it.</li>
</ul>
<p><strong>Best for:</strong> Hobby projects, simple marketing sites, and teams that mainly need basic uptime checks.</p>
<p>If you are comparing OnlineOrNot directly with UptimeRobot, read the <a href="/uptimerobot-alternative">UptimeRobot alternative comparison</a>.</p>
<h2>Best for incident response: Better Stack</h2>
<p><a href="https://betterstack.com/">Better Stack</a> combines uptime monitoring, incident management, on-call scheduling, logs, and status pages.</p>
<p>That makes it useful if your team wants more than “send an alert when a check fails.” Better Stack can help assign incidents, escalate alerts, and communicate with customers from the same broader platform.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Monitoring plus incident response</strong> - Good fit for teams that want on-call workflows with monitoring.</li>
<li><strong>Status pages</strong> - Public incident communication is part of the product.</li>
<li><strong>Logs and broader platform features</strong> - Useful if you want more than uptime monitoring.</li>
<li><strong>Team workflows</strong> - Better fit for operational teams than simple personal monitoring tools.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>More product than some teams need if they only want lightweight uptime monitoring.</li>
<li>Pricing can increase as you adopt more of the suite.</li>
<li>Teams that already use another incident-management workflow may not need the full platform.</li>
</ul>
<p><strong>Best for:</strong> Teams that want uptime monitoring and incident response in one place.</p>
<h2>Best self-hosted option: Uptime Kuma</h2>
<p><a href="https://github.com/louislam/uptime-kuma">Uptime Kuma</a> is an open-source monitoring tool you run yourself.</p>
<p>It is popular because it has a friendly interface, supports many monitor types, and does not charge per monitor. The tradeoff is that you are responsible for keeping the monitoring system itself online.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Free and open source</strong> - No SaaS subscription.</li>
<li><strong>Flexible monitor types</strong> - Supports HTTP, TCP, DNS, Docker containers, and more.</li>
<li><strong>Good UI</strong> - More approachable than many self-hosted monitoring tools.</li>
<li><strong>Private network support</strong> - Useful when you need monitoring inside your own infrastructure.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>You maintain the server, backups, updates, and alert delivery.</li>
<li>Single-location monitoring is the default unless you design around it.</li>
<li>If the server running Uptime Kuma fails, your monitoring can fail with it.</li>
</ul>
<p><strong>Best for:</strong> Self-hosters, home labs, internal services, and teams comfortable maintaining their own monitoring stack.</p>
<h2>Best enterprise option: Datadog Synthetic Monitoring</h2>
<p><a href="https://www.datadoghq.com/product/synthetic-monitoring/">Datadog Synthetic Monitoring</a> is strongest when your company already uses Datadog for logs, metrics, traces, dashboards, and incident investigation.</p>
<p>For simple Pingdom replacement use cases, Datadog is usually more than you need. For larger teams that want synthetic checks connected to the rest of their observability data, it can be powerful.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Deep observability integration</strong> - Failed checks can connect to traces, logs, services, and infrastructure metrics.</li>
<li><strong>API and browser tests</strong> - Useful for complex synthetic monitoring.</li>
<li><strong>Enterprise controls</strong> - Good fit for larger organizations.</li>
<li><strong>Global locations</strong> - Checks can run from many locations around the world.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>Pricing can be hard to forecast for small teams.</li>
<li>It is not lightweight if all you need is uptime monitoring.</li>
<li>Public status pages are not the main product.</li>
</ul>
<p><strong>Best for:</strong> Enterprises already standardized on Datadog.</p>
<h2>Other Pingdom alternatives to consider</h2>
<h3>StatusCake</h3>
<p><a href="https://www.statuscake.com/">StatusCake</a> offers uptime monitoring, SSL monitoring, domain monitoring, page speed checks, and status pages.</p>
<p>It is a practical Pingdom alternative if you want a broader website monitoring suite and do not mind feature access changing by plan.</p>
<p><strong>Best for:</strong> Teams that want uptime monitoring plus page speed, SSL, and domain checks.</p>
<h3>Uptime.com</h3>
<p><a href="https://uptime.com/">Uptime.com</a> is a broader website monitoring platform with uptime checks, transaction checks, API monitoring, status pages, and reporting.</p>
<p>It is positioned more toward businesses with larger monitoring requirements than individual developers.</p>
<p><strong>Best for:</strong> Larger teams that need multiple monitoring types and reporting.</p>
<h3>Checkly</h3>
<p><a href="https://www.checklyhq.com/">Checkly</a> focuses on code-first synthetic monitoring. It is strongest when your team wants to write checks as code and integrate monitoring into development workflows.</p>
<p><strong>Best for:</strong> Engineering teams that want programmable API and browser checks.</p>
<h3>Hyperping</h3>
<p><a href="https://hyperping.com/">Hyperping</a> is a lightweight uptime monitoring and status page product.</p>
<p>It is worth considering if you want a simple Pingdom alternative focused on uptime checks, incident communication, and a clean interface.</p>
<p><strong>Best for:</strong> Small teams that want lightweight uptime monitoring and status pages.</p>
<h2>How to choose a Pingdom alternative</h2>
<p>The best Pingdom alternative depends on why you are leaving Pingdom.</p>
<h3>1. If you need faster alerts</h3>
<p>Look for 30-second or 1-minute checks on production plans. A 5-minute check interval may be fine for side projects, but it can be too slow for production SaaS apps and APIs.</p>
<h3>2. If you get too many false positives</h3>
<p>Prioritize multi-location verification. A single regional network issue should not wake up your on-call team if customers are not affected.</p>
<h3>3. If public URLs are not enough</h3>
<p>Choose a tool that supports browser checks, API checks, and cron job monitoring. Your homepage can be online while signup, billing, or a background job is broken.</p>
<h3>4. If customers need outage updates</h3>
<p>Choose a monitoring tool with built-in status pages or a clean status page integration. During incidents, communication matters as much as detection.</p>
<h3>5. If your team is growing</h3>
<p>Check how pricing changes with users, monitors, integrations, and status pages. A tool that is cheap for one person can become awkward for a team.</p>
<h2>FAQ</h2>
<h3>What is the best Pingdom alternative?</h3>
<p>For engineering teams that want uptime checks, browser checks, cron job monitoring, alerting, and status pages in one product, OnlineOrNot is a strong Pingdom alternative. UptimeRobot is a good free starting point, Better Stack is good for incident response, and Datadog is best for enterprises already using Datadog.</p>
<h3>Is there a free Pingdom alternative?</h3>
<p>Yes. UptimeRobot, Uptime Kuma, Better Stack, Hyperping, and OnlineOrNot all have free or low-cost ways to start, depending on whether you want hosted monitoring or self-hosted monitoring.</p>
<h3>Why do teams switch away from Pingdom?</h3>
<p>Common reasons include pricing, needing faster checks, wanting fewer false positives, needing browser checks, needing cron job monitoring, or wanting status pages and team workflows in one product.</p>
<h3>Is UptimeRobot better than Pingdom?</h3>
<p>UptimeRobot can be a better fit for simple uptime monitoring and low-cost monitoring. Pingdom can be a better fit for teams that already use SolarWinds products or want a mature website monitoring vendor. For teams that need browser checks and cron monitoring, OnlineOrNot may be a better fit than either.</p>
<h3>Should I use a self-hosted Pingdom alternative?</h3>
<p>Use a self-hosted tool like Uptime Kuma if you want control and are comfortable maintaining the monitoring server. Use a hosted tool if you want monitoring to remain independent from your own infrastructure.</p>
<hr>
<h2>Related guides</h2>
<ul>
<li>For a direct side-by-side comparison, read the <a href="/pingdom-alternative">Pingdom alternative</a> guide.</li>
<li>If you are still choosing a category, compare the broader <a href="/uptime-monitoring-tools">uptime monitoring tools</a> market.</li>
<li>If you need browser journeys or scripted checks, compare <a href="/synthetic-monitoring-tools">synthetic monitoring tools</a>.</li>
</ul>
<p>The best Pingdom alternative depends on what you need next. If your team wants fast uptime checks, multi-location verification, browser checks, cron monitoring, and status pages without enterprise complexity, <a href="/get-started">try OnlineOrNot</a>.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Best UptimeRobot alternatives for 2026 (free and paid)]]></title>
            <link>https://onlineornot.com/best-uptimerobot-alternatives</link>
            <guid>https://onlineornot.com/best-uptimerobot-alternatives</guid>
            <pubDate>Fri, 29 May 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>UptimeRobot is one of the most popular uptime monitoring tools because it is simple, recognizable, and easy to start with.</p>
<p>UptimeRobot now covers multi-location, API, and heartbeat monitoring on paid plans. Teams still compare alternatives when they need Playwright scripts, different check-interval packaging, more team access, or a different incident workflow.</p>
<p>This guide compares the UptimeRobot alternatives worth considering in 2026.</p>
<p>If you specifically want to compare OnlineOrNot and UptimeRobot side by side, read the <a href="/uptimerobot-alternative">UptimeRobot alternative comparison</a>.</p>
<p>UptimeRobot details were verified on July 27, 2026 against its public <a href="https://uptimerobot.com/pricing/">pricing and feature table</a>, <a href="https://uptimerobot.com/location-specific-monitoring/">multi-location documentation</a>, and <a href="https://uptimerobot.com/cron-job-monitoring/">heartbeat monitoring documentation</a>. Plans can change, so verify current limits before purchasing.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#quick-comparison-table">Quick comparison table</a></li>
<li><a href="#best-overall-uptimerobot-alternative-onlineornot">Best overall UptimeRobot alternative: OnlineOrNot</a></li>
<li><a href="#best-for-incident-response-better-stack">Best for incident response: Better Stack</a></li>
<li><a href="#best-self-hosted-option-uptime-kuma">Best self-hosted option: Uptime Kuma</a></li>
<li><a href="#best-for-enterprise-monitoring-datadog-synthetic-monitoring">Best for enterprise monitoring: Datadog Synthetic Monitoring</a></li>
<li><a href="#best-pingdom-style-enterprise-option-pingdom">Best Pingdom-style enterprise option: Pingdom</a></li>
<li><a href="#other-uptimerobot-alternatives-to-consider">Other UptimeRobot alternatives to consider</a></li>
<li><a href="#how-to-choose-an-uptimerobot-alternative">How to choose an UptimeRobot alternative</a></li>
<li><a href="#faq">FAQ</a></li>
</ul>
<h2>Quick comparison table</h2>
<p>| Tool | Check frequency | Free plan | Browser checks | Cron monitoring | Status pages | Best for |
|------|----------------|-----------|----------------|-----------------|--------------|----------|
| <strong>OnlineOrNot</strong> | 30 seconds | Yes | Yes | Yes | Yes | Production teams that want uptime, browser checks, cron monitoring, and status pages together |
| <strong>Better Stack</strong> | 30 seconds | Yes | Yes | Heartbeats | Yes | Monitoring plus incident response |
| <strong>Uptime Kuma</strong> | Configurable | Self-hosted | Limited | Push monitors | Yes | Self-hosters and internal services |
| <strong>Datadog Synthetic Monitoring</strong> | 1 minute+ | No | Yes | No | No | Enterprises already using Datadog |
| <strong>Pingdom</strong> | 1 minute | No | Transaction monitoring | No | Via related offering | Teams that want a mature SolarWinds monitoring product |
| <strong>StatusCake</strong> | 5 min free, 30s on higher tiers | Yes | Yes | No | Yes | Website monitoring plus SSL, page speed, and domains |
| <strong>Uptime.com</strong> | 1 minute | Trial | Yes | Yes | Yes | Larger teams with broad monitoring needs |
| <strong>Hyperping</strong> | Plan-dependent | Yes | No | Heartbeats | Yes | Lightweight uptime monitoring and status pages |
| <strong>Checkly</strong> | Usage-based | Trial/free developer usage | Yes | No | No | Code-first synthetic monitoring |</p>
<h2>Best overall UptimeRobot alternative: OnlineOrNot</h2>
<p><a href="/">OnlineOrNot</a> is built for developers and small teams that need production monitoring without a heavy enterprise observability stack.</p>
<p>It covers uptime checks, API monitoring, Playwright browser checks, cron job monitoring, alerts, and hosted status pages in one product.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>
<p><strong>30-second uptime checks</strong> - Faster checks are useful when downtime affects revenue, customer trust, or support volume.</p>
</li>
<li>
<p><strong>Multi-location verification</strong> - OnlineOrNot verifies failures from multiple regions before alerting, which helps reduce false positives from regional network problems.</p>
</li>
<li>
<p><strong>Playwright browser checks</strong> - Uptime monitoring can tell you a URL responds. Browser checks can tell you whether login, signup, checkout, or a dashboard flow still works.</p>
</li>
<li>
<p><strong>Cron job monitoring</strong> - Scheduled jobs, queue workers, backups, billing jobs, and data pipelines can check in on a schedule. If a job misses its window, OnlineOrNot alerts you.</p>
</li>
<li>
<p><strong>Status pages included</strong> - Hosted status pages help communicate incidents without adding a separate status page vendor.</p>
</li>
<li>
<p><strong>Unlimited team members on paid plans</strong> - Helpful when monitoring is owned by a team, not one person.</p>
</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>No log aggregation</li>
<li>No APM or distributed tracing</li>
<li>Not designed to replace a full observability platform</li>
</ul>
<p><strong>Best for:</strong> SaaS companies, agencies, startups, and production teams that have outgrown basic uptime checks.</p>
<h2>Best for incident response: Better Stack</h2>
<p><a href="https://betterstack.com/">Better Stack</a> combines uptime monitoring with incident management, on-call scheduling, logs, and status pages.</p>
<p>It is a strong UptimeRobot alternative if your team wants alerting and incident response in the same product.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>30-second checks</strong> - Fast enough for most production websites and APIs.</li>
<li><strong>Incident response workflows</strong> - On-call scheduling, escalation, and incident timelines are built in.</li>
<li><strong>Status pages</strong> - Public communication is part of the platform.</li>
<li><strong>Broader suite</strong> - Useful if logs and incident workflows are part of the buying decision.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>More product than some small teams need.</li>
<li>Pricing can increase as you adopt more of the suite.</li>
<li>If you already have incident management elsewhere, the overlap may be unnecessary.</li>
</ul>
<p><strong>Best for:</strong> Teams that want monitoring and incident response together.</p>
<h2>Best self-hosted option: Uptime Kuma</h2>
<p><a href="https://github.com/louislam/uptime-kuma">Uptime Kuma</a> is an open-source uptime monitoring tool you run yourself.</p>
<p>It is one of the most common UptimeRobot alternatives for self-hosters because it has a friendly UI and supports many monitor types.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Free and open source</strong> - No SaaS subscription.</li>
<li><strong>Flexible monitor types</strong> - HTTP, TCP, DNS, Docker containers, and more.</li>
<li><strong>Good UI</strong> - Approachable for self-hosted monitoring.</li>
<li><strong>Private infrastructure</strong> - Useful when checks need to run inside your network.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>You maintain updates, backups, hosting, and alert delivery.</li>
<li>Monitoring can fail with your infrastructure if you host it in the same environment.</li>
<li>Multi-region monitoring requires extra work.</li>
</ul>
<p><strong>Best for:</strong> Self-hosters, home labs, private services, and teams comfortable maintaining their own monitoring infrastructure.</p>
<h2>Best for enterprise monitoring: Datadog Synthetic Monitoring</h2>
<p><a href="https://www.datadoghq.com/product/synthetic-monitoring/">Datadog Synthetic Monitoring</a> is a good fit for organizations already using Datadog for metrics, logs, traces, and dashboards.</p>
<p>It is usually too much if you only need simple uptime monitoring, but it is powerful when synthetic checks need to connect to the rest of your observability stack.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Deep observability integration</strong> - Failed checks can connect to traces, logs, services, and infrastructure data.</li>
<li><strong>API and browser tests</strong> - Useful for complex synthetic monitoring.</li>
<li><strong>Enterprise controls</strong> - Better suited to large organizations than lightweight uptime-only products.</li>
<li><strong>Global test locations</strong> - Checks can run from many locations.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>Pricing can be hard to forecast.</li>
<li>It is not lightweight for teams that only need uptime alerts.</li>
<li>Public status pages are not the core product.</li>
</ul>
<p><strong>Best for:</strong> Enterprises already standardized on Datadog.</p>
<h2>Best Pingdom-style enterprise option: Pingdom</h2>
<p><a href="https://www.pingdom.com/">Pingdom</a> is a long-running website monitoring product from SolarWinds.</p>
<p>It offers uptime monitoring, page speed monitoring, transaction monitoring, and alerting. If your team wants to move from UptimeRobot to a more established website monitoring vendor, Pingdom is one option to consider.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Mature product</strong> - Pingdom has been in the category for a long time.</li>
<li><strong>Website monitoring focus</strong> - Uptime, page speed, and transaction monitoring are core use cases.</li>
<li><strong>Recognizable vendor</strong> - Useful for teams that prefer established vendors.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li>It may be more product and vendor complexity than smaller teams want.</li>
<li>Cron job monitoring and modern developer workflows may require another tool.</li>
<li>Pricing and plan details should be checked carefully against current public pricing.</li>
</ul>
<p><strong>Best for:</strong> Teams that want a mature SolarWinds website monitoring product.</p>
<p>If you are comparing OnlineOrNot directly with Pingdom, read the <a href="/pingdom-alternative">Pingdom alternative comparison</a>.</p>
<h2>Other UptimeRobot alternatives to consider</h2>
<h3>StatusCake</h3>
<p><a href="https://www.statuscake.com/">StatusCake</a> offers uptime monitoring, SSL monitoring, domain monitoring, page speed checks, and status pages.</p>
<p>It is a practical alternative if you want a broader website monitoring suite and do not mind plan-based feature access.</p>
<p><strong>Best for:</strong> Teams that want uptime monitoring plus SSL, domain, and page speed checks.</p>
<h3>Uptime.com</h3>
<p><a href="https://uptime.com/">Uptime.com</a> is a broader monitoring platform with uptime checks, transaction checks, API monitoring, status pages, and reporting.</p>
<p>It is positioned more toward businesses with larger monitoring needs than individual developers.</p>
<p><strong>Best for:</strong> Larger teams that need multiple monitoring types and reporting.</p>
<h3>Hyperping</h3>
<p><a href="https://hyperping.com/">Hyperping</a> is a lightweight uptime monitoring and status page product.</p>
<p>It is worth considering if you want a simple hosted alternative focused on uptime checks and customer communication.</p>
<p><strong>Best for:</strong> Small teams that want lightweight uptime monitoring and status pages.</p>
<h3>Checkly</h3>
<p><a href="https://www.checklyhq.com/">Checkly</a> focuses on code-first synthetic monitoring.</p>
<p>It is strongest when your team wants checks defined in code and integrated into engineering workflows.</p>
<p><strong>Best for:</strong> Engineering teams that want programmable API and browser monitoring.</p>
<h2>How to choose an UptimeRobot alternative</h2>
<p>The best UptimeRobot alternative depends on what made you start looking.</p>
<h3>1. If check frequency is the issue</h3>
<p>Choose a tool with 30-second or 1-minute checks on the plan you will actually use. Free tiers often check less frequently than production teams need.</p>
<h3>2. If false positives are the issue</h3>
<p>Look for multi-region verification. A single monitoring location can create noisy alerts when the problem is regional routing rather than a real outage.</p>
<h3>3. If basic uptime checks are not enough</h3>
<p>Pick a tool that supports browser checks, API checks, content assertions, and cron job monitoring. A service can respond with <code>200 OK</code> while the customer experience is broken.</p>
<h3>4. If incident communication matters</h3>
<p>Choose a tool with hosted status pages or a clean status page integration. During outages, customers want to know whether you know about the problem.</p>
<h3>5. If your team is growing</h3>
<p>Check how pricing changes with users, monitors, alert channels, status pages, and synthetic checks. Team pricing matters once monitoring is no longer owned by one person.</p>
<h2>FAQ</h2>
<h3>What is the best UptimeRobot alternative?</h3>
<p>For engineering teams that need uptime checks, browser checks, cron monitoring, status pages, and multi-location verification, OnlineOrNot is a strong UptimeRobot alternative. Better Stack is strong for incident response, Uptime Kuma is strong for self-hosting, and Datadog is best for enterprises already using Datadog.</p>
<h3>Is there a free UptimeRobot alternative?</h3>
<p>Yes. OnlineOrNot, Better Stack, Hyperping, and StatusCake have free or low-cost ways to start. Uptime Kuma is free and open source if you are comfortable hosting it yourself.</p>
<h3>Why do teams switch away from UptimeRobot?</h3>
<p>Common reasons include needing Playwright scripts, 30-second checks outside an Enterprise tier, different team-seat pricing, more flexible alerting, or a different incident workflow. UptimeRobot already offers paid multi-location and heartbeat monitoring, so those features alone are not reasons to switch.</p>
<h3>Is Uptime Kuma better than UptimeRobot?</h3>
<p>Uptime Kuma can be better if you want self-hosted monitoring and control over your infrastructure. UptimeRobot is easier if you want hosted monitoring without maintaining a server. For production teams that want hosted monitoring plus browser checks and cron monitoring, OnlineOrNot is usually a better fit than self-hosting.</p>
<h3>Is Pingdom better than UptimeRobot?</h3>
<p>Pingdom can be a better fit for teams that want a mature website monitoring product from SolarWinds. UptimeRobot can be a better fit for a large free monitor allowance or its current paid feature set. OnlineOrNot is a better fit when the comparison is about Playwright scripts, 30-second Pro checks, and unlimited team access.</p>
<hr>
<h2>Related guides</h2>
<ul>
<li>For a direct side-by-side comparison, read the <a href="/uptimerobot-alternative">UptimeRobot alternative</a> guide.</li>
<li>If you are still choosing a category, compare the broader <a href="/uptime-monitoring-tools">uptime monitoring tools</a> market.</li>
<li>If you need browser journeys or scripted checks, compare <a href="/synthetic-monitoring-tools">synthetic monitoring tools</a>.</li>
</ul>
<p>UptimeRobot has grown beyond basic uptime checks. If you want Playwright scripts, 30-second Pro checks, cron monitoring, status pages, and unlimited team members in one product, <a href="/get-started">try OnlineOrNot</a>.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[10 Best Uptime Monitoring Tools for 2026]]></title>
            <link>https://onlineornot.com/uptime-monitoring-tools</link>
            <guid>https://onlineornot.com/uptime-monitoring-tools</guid>
            <pubDate>Thu, 21 May 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Uptime monitoring tools check your website, API, or app on a schedule and alert you when something stops working.</p>
<p>That sounds simple, but the details matter. A tool that checks every 5 minutes from one region will catch different problems than a tool checking every 30 seconds from multiple regions. A tool built for hobby projects will feel very different from an enterprise observability platform.</p>
<p>This guide compares the uptime monitoring tools worth considering in 2026, with a focus on practical tradeoffs: check frequency, false positives, alerting, status pages, pricing, and who each tool is actually best for.</p>
<h2>Best uptime monitoring tools: short answer</h2>
<ul>
<li><strong>OnlineOrNot</strong> is the best fit for small software teams that want 30-second checks, multi-region verification, browser and cron monitoring, and status pages in one product.</li>
<li><strong>UptimeRobot</strong> is a practical free starting point for basic website checks.</li>
<li><strong>Better Stack</strong> is the strongest fit when uptime monitoring needs to share a workflow with incident management and on-call scheduling.</li>
<li><strong>Uptime Kuma</strong> is the best fit for teams that want to self-host and are prepared to operate their own monitoring infrastructure.</li>
<li><strong>Datadog Synthetic Monitoring</strong> is the best fit for enterprises that already use Datadog logs, metrics, and traces.</li>
</ul>
<p>There is no universal winner. Shortlist tools based on the check interval, verification method, alert delivery, monitor types, status-page requirements, and total price for your production setup.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#best-uptime-monitoring-tools-short-answer">Best uptime monitoring tools: short answer</a></li>
<li><a href="#quick-comparison-table">Quick comparison table</a></li>
<li><a href="#what-is-uptime-monitoring">What is uptime monitoring?</a></li>
<li><a href="#best-for-most-teams-onlineornot">Best for most teams: OnlineOrNot</a></li>
<li><a href="#best-free-starting-point-uptimerobot">Best free starting point: UptimeRobot</a></li>
<li><a href="#best-for-incident-management-better-stack">Best for incident management: Better Stack</a></li>
<li><a href="#best-self-hosted-option-uptime-kuma">Best self-hosted option: Uptime Kuma</a></li>
<li><a href="#best-for-enterprises-datadog">Best for enterprises: Datadog</a></li>
<li><a href="#other-uptime-monitoring-tools-to-consider">Other uptime monitoring tools to consider</a></li>
<li><a href="#how-to-choose-an-uptime-monitoring-tool">How to choose an uptime monitoring tool</a></li>
<li><a href="#faq">FAQ</a></li>
</ul>
<h2>Quick comparison table</h2>
<p>| Tool | Check frequency | Free plan | Multi-region checks | Status pages | Best for |
|------|----------------|-----------|---------------------|--------------|----------|
| <strong>OnlineOrNot</strong> | 30 seconds | Yes | Yes | Yes | Developers, startups, SaaS teams |
| <strong>UptimeRobot</strong> | 5 min free, 1 min paid | Yes | Limited | Yes | Hobby projects, simple sites |
| <strong>Better Stack</strong> | 30 seconds | Yes | Yes | Yes | Teams wanting monitoring + incident response |
| <strong>Uptime Kuma</strong> | Configurable | Self-hosted | DIY | Yes | Self-hosters, home labs |
| <strong>Pingdom</strong> | 1 minute | No | Yes | No | Legacy enterprise environments |
| <strong>StatusCake</strong> | 5 min free, 30s on higher tiers | Yes | Yes | Yes | Teams needing uptime + page speed |
| <strong>Uptime.com</strong> | 1 minute | Trial | Yes | Yes | Larger teams with broader monitoring needs |
| <strong>HetrixTools</strong> | 1 minute | Yes | Yes | Yes | Server and blacklist monitoring |
| <strong>Datadog Synthetic Monitoring</strong> | 1 minute+ | No | Yes | No | Enterprises already using Datadog |
| <strong>Sentry Uptime Monitoring</strong> | 1 minute | Included with Sentry plans | Yes | No | Teams already using Sentry |</p>
<h2>What is uptime monitoring?</h2>
<p>Uptime monitoring is the practice of regularly checking whether a website, API, server, or application is reachable and responding correctly.</p>
<p>At its simplest, an uptime monitor sends a request to your URL every few minutes. If the request fails, times out, or returns the wrong status code, the tool sends an alert.</p>
<p>Good uptime monitoring goes further:</p>
<ul>
<li><strong>Checks from multiple regions</strong> so one network blip does not create a false alarm</li>
<li><strong>Verifies response content</strong> so a broken page returning <code>200 OK</code> is still caught</li>
<li><strong>Tracks response time</strong> so you can spot slowdowns before a full outage</li>
<li><strong>Routes alerts correctly</strong> so the right person hears about the right incident</li>
<li><strong>Publishes status pages</strong> so customers can see what is happening during an outage</li>
</ul>
<p>If you are new to the category, start with the broader <a href="/website-monitoring-guide">website monitoring guide</a>. If you need browser journeys or scripted API tests, compare <a href="/synthetic-monitoring-tools">synthetic monitoring tools</a> instead.</p>
<h2>Best for most teams: OnlineOrNot</h2>
<p><a href="/">OnlineOrNot</a> is built for developers and small teams that want fast, practical uptime monitoring without buying a full enterprise observability stack.</p>
<p>It covers the basics most teams actually need: uptime checks, API monitoring, cron job monitoring, browser checks, alerts, and hosted status pages.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>
<p><strong>30-second checks on paid plans</strong> - Fast checks are useful when downtime costs money. Waiting 5 minutes to find out your checkout or API is down is too long for most production apps.</p>
</li>
<li>
<p><strong>Multi-location verification</strong> - OnlineOrNot checks from multiple regions before alerting, which helps reduce false positives caused by regional network issues.</p>
</li>
<li>
<p><strong>Simple pricing</strong> - Pro starts at $15/month and includes 10 uptime monitors, 3,000 browser check runs, one public status page, and unlimited team members. Annual billing reduces the base price to $12/month.</p>
</li>
<li>
<p><strong>Status pages included</strong> - Public status pages are included, so you can communicate incidents without paying for a separate status page product.</p>
</li>
<li>
<p><strong>More than basic ping checks</strong> - You can monitor websites, APIs, cron jobs, and browser checks in the same product.</p>
</li>
</ul>
<p><strong>What it does not do:</strong></p>
<ul>
<li>No APM or distributed tracing</li>
<li>No log aggregation</li>
<li>Not designed to replace Datadog, New Relic, or a full observability platform</li>
</ul>
<p><strong>Pricing:</strong> The free Hobby plan includes 3 monitors. Pro starts at $15/month, or $12/month billed annually, with 10 uptime monitors included. See the <a href="/pricing">current OnlineOrNot pricing</a>.</p>
<p><strong>Best for:</strong> SaaS companies, developers, startups, and small teams that need reliable uptime monitoring without enterprise complexity.</p>
<h2>Best free starting point: UptimeRobot</h2>
<p><a href="https://uptimerobot.com/">UptimeRobot</a> is one of the best-known uptime monitoring tools. Many developers start there because it is free, simple, and quick to set up.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Generous free plan</strong> - The free tier is useful for personal projects and low-stakes sites.</li>
<li><strong>Very easy setup</strong> - Add a URL, choose alert channels, and you are done.</li>
<li><strong>Cheap paid plans</strong> - Paid plans are inexpensive compared with enterprise tools.</li>
<li><strong>Recognizable product</strong> - It has been around for a long time, so many teams already know how it works.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>5-minute checks on the free tier</strong> - A 5-minute interval can mean several minutes of downtime before you know anything happened.</li>
<li><strong>Limited depth</strong> - It is good for basic uptime checks, but less useful if you need richer API assertions, browser monitoring, or more advanced team workflows.</li>
<li><strong>Free plan restrictions</strong> - If you are using monitoring commercially, check the current terms before relying on the free tier.</li>
</ul>
<p><strong>Pricing:</strong> Free plan available. Paid plans start at low monthly pricing depending on monitor count and features.</p>
<p><strong>Best for:</strong> Hobby projects, side projects, personal websites, and teams that only need basic uptime checks.</p>
<h2>Best for incident management: Better Stack</h2>
<p><a href="https://betterstack.com/">Better Stack</a> combines uptime monitoring with incident management, status pages, logs, and on-call scheduling.</p>
<p>This is useful if you do not just want to know that something broke. You also want to assign incidents, escalate alerts, and keep a public status page updated from the same workflow.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>30-second checks</strong> - Fast enough for most production websites and APIs.</li>
<li><strong>Incident response built in</strong> - On-call scheduling, escalation, incident timelines, and status pages are part of the product.</li>
<li><strong>Good team workflows</strong> - Better fit for teams than tools aimed at individual developers.</li>
<li><strong>Broad monitoring suite</strong> - Useful if you also want logs and incident management under one brand.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>More product than some teams need</strong> - If you just want a few uptime checks, the incident management layer may be unnecessary.</li>
<li><strong>Pricing rises as you adopt more of the suite</strong> - It can be good value if you use the whole platform, but less compelling if you only need uptime monitoring.</li>
</ul>
<p><strong>Pricing:</strong> Free plan available. Paid monitoring plans are typically more expensive than lightweight uptime-only tools.</p>
<p><strong>Best for:</strong> Teams that want uptime monitoring and incident response in one place.</p>
<h2>Best self-hosted option: Uptime Kuma</h2>
<p><a href="https://github.com/louislam/uptime-kuma">Uptime Kuma</a> is an open-source uptime monitoring tool you run yourself.</p>
<p>It is popular because it has a friendly UI, supports many monitor types, and does not charge per monitor. The tradeoff is that you are responsible for keeping your monitoring system online.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Free and open source</strong> - No monthly SaaS subscription.</li>
<li><strong>Nice interface</strong> - Much more polished than many self-hosted monitoring tools.</li>
<li><strong>Flexible monitor types</strong> - Supports HTTP, TCP, DNS, Docker containers, and more.</li>
<li><strong>Good for private infrastructure</strong> - Useful when you need monitoring inside a private network.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>You run the infrastructure</strong> - Updates, backups, security, uptime, and alert delivery are your responsibility.</li>
<li><strong>Single-location by default</strong> - Unless you deploy multiple instances, you only monitor from wherever your server runs.</li>
<li><strong>Monitoring can fail silently</strong> - If the machine running Uptime Kuma dies, your monitoring dies with it.</li>
</ul>
<p><strong>Pricing:</strong> Free, but you pay for hosting and maintenance time.</p>
<p><strong>Best for:</strong> Self-hosters, home labs, internal services, and teams comfortable maintaining their own monitoring infrastructure.</p>
<h2>Best for enterprises: Datadog</h2>
<p><a href="https://www.datadoghq.com/product/synthetic-monitoring/">Datadog Synthetic Monitoring</a> is a strong option for companies already using Datadog for logs, metrics, traces, and dashboards.</p>
<p>For simple uptime monitoring, Datadog is usually overkill. For large teams that need to connect external checks to backend traces and infrastructure telemetry, it can be powerful.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Deep observability integration</strong> - A failing uptime check can connect to traces, logs, services, and infrastructure metrics.</li>
<li><strong>API and browser tests</strong> - Useful for complex synthetic monitoring, not just simple uptime checks.</li>
<li><strong>Enterprise controls</strong> - Strong fit for larger organizations with existing Datadog workflows.</li>
<li><strong>Global locations</strong> - Checks can run from many locations around the world.</li>
</ul>
<p><strong>Where it falls short:</strong></p>
<ul>
<li><strong>Pricing can be hard to predict</strong> - Usage-based synthetic test pricing requires more planning than per-monitor pricing.</li>
<li><strong>Not lightweight</strong> - If all you need is “tell me when my site is down,” Datadog is more tool than you need.</li>
<li><strong>No built-in public status page</strong> - You will likely use another product for customer-facing incident communication.</li>
</ul>
<p><strong>Pricing:</strong> Usage-based synthetic monitoring pricing. API and browser tests are priced differently.</p>
<p><strong>Best for:</strong> Enterprises already standardized on Datadog.</p>
<h2>Other uptime monitoring tools to consider</h2>
<h3>Pingdom</h3>
<p><a href="https://www.pingdom.com/">Pingdom</a> is a long-running website monitoring product now owned by SolarWinds.</p>
<p>It offers uptime monitoring, page speed monitoring, transaction monitoring, and real user monitoring. It is reliable and recognizable, but the product can feel more legacy than newer developer-focused tools.</p>
<p><strong>Best for:</strong> Teams already using SolarWinds or companies that want a familiar enterprise vendor.</p>
<h3>StatusCake</h3>
<p><a href="https://www.statuscake.com/">StatusCake</a> offers uptime monitoring, SSL monitoring, domain monitoring, page speed checks, and status pages.</p>
<p>It is a good middle-ground option if you want a broader website monitoring suite and do not mind feature access changing by tier.</p>
<p><strong>Best for:</strong> Teams that want uptime monitoring plus page speed and domain checks.</p>
<h3>Uptime.com</h3>
<p><a href="https://uptime.com/">Uptime.com</a> is a broader website monitoring platform with uptime checks, transaction checks, API monitoring, status pages, and reporting.</p>
<p>It is positioned more toward businesses with larger monitoring needs than individual developers.</p>
<p><strong>Best for:</strong> Larger teams that need multiple monitoring types and reporting features.</p>
<h3>HetrixTools</h3>
<p><a href="https://hetrixtools.com/uptime-monitor/">HetrixTools</a> combines uptime monitoring with server monitoring and blacklist monitoring.</p>
<p>The blacklist monitoring angle makes it especially useful for teams that care about email deliverability or IP reputation.</p>
<p><strong>Best for:</strong> Teams monitoring servers, IP reputation, and uptime together.</p>
<h3>Sentry Uptime Monitoring</h3>
<p><a href="https://sentry.io/product/uptime-monitoring/">Sentry</a> now offers uptime monitoring alongside its error monitoring and performance products.</p>
<p>This can be convenient if your team already uses Sentry every day. The value is less about replacing dedicated uptime monitoring tools and more about connecting downtime to errors and traces.</p>
<p><strong>Best for:</strong> Engineering teams already using Sentry for errors and performance.</p>
<h2>How to choose an uptime monitoring tool</h2>
<p>The best uptime monitoring tool depends less on the logo and more on how you operate.</p>
<h3>1. Pick the right check frequency</h3>
<p>Check frequency determines how quickly you find out something is wrong.</p>
<ul>
<li><strong>30 seconds:</strong> Good for production SaaS apps, APIs, checkout flows, and customer-facing services</li>
<li><strong>1 minute:</strong> Good default for most business websites and APIs</li>
<li><strong>5 minutes:</strong> Acceptable for hobby projects or low-risk pages</li>
</ul>
<p>If downtime costs money or breaks customer trust, avoid relying on 5-minute checks for critical services.</p>
<h3>2. Decide whether you need multi-region verification</h3>
<p>Single-region monitoring is simple, but it can create false positives. A network issue between the monitoring server and your website can look like an outage even if customers are fine.</p>
<p>Multi-region verification reduces noise by checking from more than one location before alerting.</p>
<p>This matters most if:</p>
<ul>
<li>Your customers are global</li>
<li>You use a CDN</li>
<li>Your team ignores alerts after a few false alarms</li>
<li>You have had “site down” alerts that turned out to be regional routing problems</li>
</ul>
<h3>3. Check more than the homepage</h3>
<p>Monitoring only your homepage is better than nothing, but it misses many real failures.</p>
<p>For a production SaaS, consider monitoring:</p>
<ul>
<li>Homepage</li>
<li>Login page</li>
<li>Signup page</li>
<li>API health endpoint</li>
<li>Billing or checkout flow</li>
<li>Public status page</li>
<li>Cron jobs and background jobs</li>
<li>SSL certificate expiry</li>
</ul>
<p>The goal is not to monitor every URL. It is to monitor the paths that matter when customers are trying to use or buy your product.</p>
<h3>4. Verify correctness, not just reachability</h3>
<p>A page can return <code>200 OK</code> and still be broken.</p>
<p>Common examples:</p>
<ul>
<li>Your app serves a blank page because JavaScript failed to load</li>
<li>Your backend returns a branded error page with a <code>200</code> status code</li>
<li>Your CDN serves stale maintenance content</li>
<li>Your API returns a valid status code with invalid data</li>
</ul>
<p>At minimum, use content checks for key pages. For APIs, use assertions against status codes, response bodies, and response time thresholds. For JavaScript-heavy apps, consider browser checks.</p>
<h3>5. Make sure alerting matches severity</h3>
<p>The tool matters less if alerts go to the wrong place.</p>
<p>Use different alerting rules for different services:</p>
<ul>
<li><strong>Critical production services:</strong> SMS, phone call, PagerDuty, or on-call escalation</li>
<li><strong>Important but non-critical pages:</strong> Slack or Teams</li>
<li><strong>Low-priority marketing pages:</strong> Email or low-urgency Slack channels</li>
</ul>
<p>If every alert is urgent, none of them are urgent. Alert fatigue is one of the most common ways monitoring setups fail.</p>
<h3>6. Consider status pages early</h3>
<p>During an outage, customers want to know two things: whether you know about the problem, and when they should expect an update.</p>
<p>A public status page helps you communicate without replying to every support ticket manually.</p>
<p>If you already know you need a status page, choose a monitoring tool that includes one or integrates cleanly with the one you use.</p>
<h2>FAQ</h2>
<h3>What is the best uptime monitoring tool?</h3>
<p>For most small teams and SaaS companies, the best uptime monitoring tool is one that offers fast checks, reliable alerting, multi-region verification, and status pages without enterprise complexity.</p>
<p>OnlineOrNot is a good fit for that use case. UptimeRobot is a good free starting point. Better Stack is a good choice if you also want incident management. Datadog is best if your company already uses Datadog heavily.</p>
<h3>What is uptime monitoring software?</h3>
<p>Uptime monitoring software checks whether your website, API, or service is reachable. It runs checks on a schedule, records availability and response time, and alerts you when a check fails.</p>
<h3>What is the difference between uptime monitoring and website monitoring?</h3>
<p>Uptime monitoring focuses on whether a service is up or down. <a href="/website-monitoring-guide">Website monitoring</a> is broader and can include uptime, page speed, SSL certificates, broken pages, real user monitoring, and synthetic browser tests.</p>
<p>In practice, people often use the terms interchangeably.</p>
<h3>What is the difference between uptime monitoring and synthetic monitoring?</h3>
<p>Uptime monitoring usually means simple scheduled checks: is this URL reachable, did it return the expected status code, and did it respond fast enough?</p>
<p><a href="/synthetic-monitoring-tools">Synthetic monitoring</a> is broader. It can include scripted API checks, browser-based user journeys, login flows, checkout tests, and multi-step workflows.</p>
<h3>How often should I check my website uptime?</h3>
<p>For production websites and APIs, every 30 seconds to 1 minute is a good default. For low-risk personal sites, every 5 minutes may be enough.</p>
<p>Use faster intervals for services where downtime affects revenue, customer trust, or incident response.</p>
<h3>Are free uptime monitoring tools good enough?</h3>
<p>Free uptime monitoring tools are good enough for personal projects, early-stage side projects, and low-risk websites.</p>
<p>For business-critical websites and APIs, paid tools are usually worth it because they offer faster checks, better alerting, multi-region verification, team features, and more reliable incident communication.</p>
<h3>Should I use a self-hosted uptime monitoring tool?</h3>
<p>Use a self-hosted tool like Uptime Kuma if you want control, need to monitor private infrastructure, or enjoy maintaining your own systems.</p>
<p>Use a managed tool if you want your monitoring to remain independent from your own infrastructure. If your self-hosted monitoring server goes down at the same time as your app, you may not get alerted.</p>
<hr>
<h2>Related comparisons</h2>
<ul>
<li>If you are replacing UptimeRobot, read the <a href="/uptimerobot-alternative">UptimeRobot alternative</a> comparison or the broader <a href="/best-uptimerobot-alternatives">best UptimeRobot alternatives</a> guide.</li>
<li>If you are replacing Pingdom, read the <a href="/pingdom-alternative">Pingdom alternative</a> comparison or the broader <a href="/best-pingdom-alternatives">best Pingdom alternatives</a> guide.</li>
<li>If simple uptime checks are not enough, compare <a href="/synthetic-monitoring-tools">synthetic monitoring tools</a>.</li>
</ul>
<p>Uptime monitoring is one of the simplest reliability improvements you can make. Start with your most important URL, add alerts you will actually notice, then expand to APIs, signup flows, cron jobs, and status pages.</p>
<p>If you want a practical starting point, <a href="/pricing">compare OnlineOrNot plans</a> or <a href="/get-started">start a 14-day trial</a> with 30-second uptime checks, multi-region verification, and status pages without adopting a full observability platform.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Best synthetic monitoring tools for 2026 (tested and compared)]]></title>
            <link>https://onlineornot.com/synthetic-monitoring-tools</link>
            <guid>https://onlineornot.com/synthetic-monitoring-tools</guid>
            <pubDate>Tue, 10 Mar 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Synthetic monitoring tools check your websites and APIs before your users do. They run automated tests from multiple locations, catch outages, and alert you when something breaks.</p>
<p>The problem is that there are dozens of options, ranging from free open-source tools to enterprise platforms costing thousands per month.</p>
<p>This guide compares the ones worth considering in 2026, based on actual usage - not vendor marketing.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#what-is-synthetic-monitoring">What is synthetic monitoring?</a></li>
<li><a href="#quick-comparison-table">Quick comparison table</a></li>
<li><a href="#best-for-most-teams-onlineornot">Best for most teams: OnlineOrNot</a></li>
<li><a href="#best-for-enterprise-datadog-synthetic-monitoring">Best for enterprise: Datadog Synthetic Monitoring</a></li>
<li><a href="#best-budget-option-uptimerobot">Best budget option: UptimeRobot</a></li>
<li><a href="#best-self-hosted-uptime-kuma">Best self-hosted: Uptime Kuma</a></li>
<li><a href="#other-options-worth-considering">Other options worth considering</a></li>
<li><a href="#how-to-choose-the-right-tool">How to choose the right tool</a></li>
<li><a href="#faq">FAQ</a></li>
</ul>
<h2>What is synthetic monitoring?</h2>
<p>Synthetic monitoring means running automated tests against your websites and APIs on a schedule. Unlike real user monitoring (RUM), which tracks actual user sessions, synthetic monitoring proactively checks your services from the outside.</p>
<p>Think of it as a robot that visits your site every 30 seconds and tells you if something breaks.</p>
<p>The basic flow:</p>
<ol>
<li>You configure a check (URL, expected response, acceptable response time)</li>
<li>The tool runs that check at regular intervals from multiple locations</li>
<li>If something fails, you get an alert</li>
</ol>
<p>Most tools also track response times over time, so you can spot performance degradation before it becomes an outage.</p>
<p><strong>When you need synthetic monitoring:</strong></p>
<ul>
<li>You want to know about outages before users report them</li>
<li>You need to verify that deployments didn't break production</li>
<li>You're monitoring APIs that external services depend on</li>
<li>You want to track uptime for SLA reporting</li>
</ul>
<p><strong>When RUM is better:</strong></p>
<ul>
<li>You need to understand real user experience across different devices</li>
<li>You're debugging performance issues specific to certain browsers or regions</li>
<li>You want to correlate frontend errors with user sessions</li>
</ul>
<p>Most teams benefit from both, but synthetic monitoring is the baseline everyone should have.</p>
<h2>Quick comparison table</h2>
<p>| Tool | Check frequency | Pricing | Multi-region | Browser tests | Best for |
|------|----------------|---------|--------------|---------------|----------|
| <strong>OnlineOrNot</strong> | 30 seconds | From $12/mo | Yes (17 regions) | Yes | Developers, startups, growing teams |
| <strong>Datadog Synthetics</strong> | 1 minute+ | ~$5+ per 10k tests | Yes | Yes | Enterprise with existing Datadog stack |
| <strong>Better Stack</strong> | 30 seconds | From $24/mo | Yes | No | Teams wanting incident management included |
| <strong>UptimeRobot</strong> | 5 min (free), 1 min (paid) | Free / $7+/mo | Limited | No | Hobbyists, basic monitoring |
| <strong>Pingdom</strong> | 1 minute | From $15/mo | Yes | Yes | Legacy enterprise environments |
| <strong>StatusCake</strong> | 5 min (free), 30s (Business) | From €20/mo | Yes | No | Teams needing page speed monitoring |
| <strong>Uptime Kuma</strong> | Configurable | Free (self-hosted) | DIY | No | Self-hosters, home labs |
| <strong>Checkly</strong> | 1 minute | From $30/mo | Yes | Yes | Teams with complex browser test needs |</p>
<h2>Best for most teams: OnlineOrNot</h2>
<p><a href="/">OnlineOrNot</a> is built for developers who want reliable monitoring without enterprise complexity.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>
<p><strong>30-second checks on all paid plans</strong> - Most competitors reserve fast checks for expensive tiers. OnlineOrNot checks every 30 seconds starting at $12/month.</p>
</li>
<li>
<p><strong>Multi-location verification</strong> - Every check runs from multiple regions before alerting. This eliminates false positives from network blips or regional CDN issues.</p>
</li>
<li>
<p><strong>Simple pricing</strong> - $2.40 per monitor per month (monthly billing). No confusing tiers, no "contact sales" for basic features.</p>
</li>
<li>
<p><strong>Status pages included</strong> - Branded status pages on your domain are included, not a paid add-on.</p>
</li>
<li>
<p><strong>Cron job monitoring</strong> - Track scheduled tasks and background jobs alongside your uptime checks.</p>
</li>
<li>
<p><strong>Browser checks</strong> - Run checks in real Chrome browsers to verify JavaScript apps load and work correctly.</p>
</li>
</ul>
<p><strong>What it doesn't do:</strong></p>
<ul>
<li>No APM or distributed tracing</li>
<li>No log aggregation</li>
</ul>
<p><strong>Pricing:</strong> Free tier with 3 monitors. Paid plans from $12/month for 5 monitors.</p>
<p><strong>Best for:</strong> Development teams, SaaS companies, and anyone who wants fast, reliable monitoring without paying for features they won't use.</p>
<h2>Best for enterprise: Datadog Synthetic Monitoring</h2>
<p><a href="https://www.datadoghq.com/product/synthetic-monitoring/">Datadog</a> is the default choice for enterprises already using Datadog for APM and logging.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li>
<p><strong>Deep integration with APM</strong> - Correlate synthetic test failures with traces, logs, and infrastructure metrics. When a test fails, you can see exactly what broke in the backend.</p>
</li>
<li>
<p><strong>Browser testing</strong> - Record user journeys and replay them as synthetic tests. Good for testing complex flows like checkout processes.</p>
</li>
<li>
<p><strong>Mobile app testing</strong> - Test mobile apps alongside web properties.</p>
</li>
<li>
<p><strong>Global coverage</strong> - Run tests from dozens of locations worldwide.</p>
</li>
</ul>
<p><strong>What it doesn't do well:</strong></p>
<ul>
<li>
<p><strong>Pricing is complex</strong> - You pay per test execution, which can add up quickly. Running a check every minute from 5 locations is 5 × 60 × 24 × 30 = 216,000 test executions per month.</p>
</li>
<li>
<p><strong>Overkill for simple monitoring</strong> - If you just want to know if your API is up, you don't need Datadog's full observability platform.</p>
</li>
<li>
<p><strong>Slow minimum interval</strong> - Fastest check interval is 1 minute, compared to 30 seconds elsewhere.</p>
</li>
</ul>
<p><strong>Pricing:</strong> Approximately $5 per 10,000 test executions for API tests, $12 per 1,000 for browser tests. No free tier for synthetics.</p>
<p><strong>Best for:</strong> Large organizations already invested in the Datadog ecosystem who need synthetic monitoring integrated with their existing observability stack.</p>
<h2>Best budget option: UptimeRobot</h2>
<p><a href="https://uptimerobot.com/">UptimeRobot</a> is the tool everyone starts with. It's free for basic monitoring and has been around forever.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Generous free tier</strong> - 50 monitors with 5-minute checks, completely free</li>
<li><strong>Simple interface</strong> - No learning curve, just add URLs and go</li>
<li><strong>Cheap paid plans</strong> - Pro plans start at $7/month</li>
</ul>
<p><strong>What it doesn't do well:</strong></p>
<ul>
<li><strong>Slow check intervals</strong> - Free tier is 5 minutes, which means up to 5 minutes before you know about an outage</li>
<li><strong>Single-location checks</strong> - No multi-location verification means more false positives from network issues</li>
<li><strong>Basic alerting</strong> - Limited options for routing alerts to different team members</li>
<li><strong>Recent ToS changes</strong> - Free tier can no longer be used for commercial projects</li>
</ul>
<p><strong>Pricing:</strong> Free for 50 monitors (personal use). Pro from $7/month.</p>
<p><strong>Best for:</strong> Personal projects, hobby sites, and teams with minimal monitoring budgets who can tolerate slower detection.</p>
<h2>Best self-hosted: Uptime Kuma</h2>
<p><a href="https://github.com/louislam/uptime-kuma">Uptime Kuma</a> is an open-source monitoring tool you run on your own infrastructure.</p>
<p><strong>What it does well:</strong></p>
<ul>
<li><strong>Completely free</strong> - No per-monitor costs, no monthly fees</li>
<li><strong>Full control</strong> - Your data stays on your servers</li>
<li><strong>Feature-rich</strong> - Supports HTTP, TCP, DNS, Docker, and more</li>
<li><strong>Nice UI</strong> - Clean, modern interface that's easy to use</li>
<li><strong>Active development</strong> - Regular updates and community contributions</li>
</ul>
<p><strong>What it doesn't do well:</strong></p>
<ul>
<li><strong>Single location</strong> - You're monitoring from wherever you host it. No multi-region verification unless you set up multiple instances.</li>
<li><strong>Self-hosting overhead</strong> - You need to maintain the server, handle updates, and ensure it stays online. If your monitoring server goes down, you won't know about outages.</li>
<li><strong>No team features</strong> - Limited collaboration tools compared to managed services</li>
<li><strong>No managed alerting infrastructure</strong> - You're responsible for ensuring SMS/email delivery works</li>
</ul>
<p><strong>Pricing:</strong> Free (you pay for hosting)</p>
<p><strong>Best for:</strong> Home labs, self-hosting enthusiasts, teams with strict data residency requirements, or anyone comfortable managing their own infrastructure.</p>
<h2>Other options worth considering</h2>
<h3>Better Stack</h3>
<p><a href="https://betterstack.com/">Better Stack</a> (formerly Better Uptime) combines uptime monitoring with incident management and on-call scheduling.</p>
<ul>
<li>30-second checks</li>
<li>Built-in incident management</li>
<li>On-call scheduling included</li>
<li>From $24/month</li>
</ul>
<p>Good if you want monitoring + incident response in one tool. More expensive than pure monitoring solutions.</p>
<h3>Pingdom</h3>
<p><a href="https://www.pingdom.com/">Pingdom</a> is owned by SolarWinds and has been around since 2007. Reliable but dated.</p>
<ul>
<li>1-minute minimum check interval</li>
<li>Strong brand recognition</li>
<li>From $15/month for 10 monitors</li>
</ul>
<p>Good for enterprises who already use SolarWinds products. Feels dated compared to modern alternatives.</p>
<h3>StatusCake</h3>
<p><a href="https://www.statuscake.com/">StatusCake</a> offers uptime, page speed, and server monitoring.</p>
<ul>
<li>30-second checks require €70/month Business plan</li>
<li>Page speed monitoring included</li>
<li>From €20/month</li>
</ul>
<p>Good if you need page speed testing alongside uptime monitoring. The tiered pricing with monitor caps can get expensive.</p>
<h3>Checkly</h3>
<p><a href="https://www.checklyhq.com/">Checkly</a> focuses on API and browser testing with a developer-first approach.</p>
<ul>
<li>Playwright-based browser checks</li>
<li>Monitoring as code</li>
<li>From $30/month</li>
</ul>
<p>Good for teams with complex browser testing needs or who want to version control their monitoring configuration.</p>
<h2>How to choose the right tool</h2>
<h3>What check frequency do you need?</h3>
<p>If downtime costs you significant money, you want the fastest detection possible. 30-second checks catch issues twice as fast as 1-minute checks.</p>
<ul>
<li><strong>30 seconds:</strong> OnlineOrNot, Better Stack, StatusCake (expensive tier)</li>
<li><strong>1 minute:</strong> Datadog, Pingdom, Checkly</li>
<li><strong>5 minutes:</strong> UptimeRobot (free)</li>
</ul>
<h3>What's your budget?</h3>
<p>Be honest about what you're willing to pay:</p>
<ul>
<li><strong>Free:</strong> UptimeRobot (limited), Uptime Kuma (self-hosted)</li>
<li><strong>$10-30/month:</strong> OnlineOrNot, StatusCake, Pingdom</li>
<li><strong>$50-100/month:</strong> Better Stack, Checkly</li>
<li><strong>$100+/month:</strong> Datadog, enterprise tiers of others</li>
</ul>
<h3>Do you need multi-location verification?</h3>
<p>Single-location checks trigger false positives when there's a network issue between the monitoring location and your server. Multi-location verification only alerts when multiple regions confirm the outage.</p>
<p>Most paid tools offer this. Free tools and self-hosted options typically don't.</p>
<h3>Do you need browser testing?</h3>
<p>If you need to verify that JavaScript apps load and render correctly, or test complex user flows (login, checkout, multi-step forms), you need a tool with browser testing capabilities:</p>
<ul>
<li>OnlineOrNot (real Chrome browsers)</li>
<li>Datadog Synthetic Monitoring</li>
<li>Checkly</li>
<li>Pingdom</li>
<li>Playwright/Puppeteer + your own infrastructure</li>
</ul>
<p>For simple HTTP/API monitoring, browser testing is overkill.</p>
<h3>What integrations do you need?</h3>
<p>Most tools integrate with Slack, PagerDuty, and webhooks. Check that your specific tools are supported:</p>
<ul>
<li><strong>On-call platforms:</strong> PagerDuty, Opsgenie, Incident.io</li>
<li><strong>Chat:</strong> Slack, Discord, Microsoft Teams</li>
<li><strong>Ticketing:</strong> Jira, Linear, GitHub Issues</li>
</ul>
<h2>FAQ</h2>
<h3>What is an example of synthetic monitoring?</h3>
<p>A synthetic monitor might check <code>https://api.example.com/health</code> every 30 seconds. It sends an HTTP GET request, verifies the response is a 200 status code with <code>{"status": "ok"}</code> in the body, and confirms the response time is under 500ms. If any of those conditions fail, you get an alert.</p>
<h3>What is a synthetic monitoring agent?</h3>
<p>A synthetic monitoring agent is software that runs the actual checks. Managed services run agents in data centers around the world. Self-hosted solutions like Uptime Kuma require you to run the agent on your own server.</p>
<h3>What are the different types of synthetic monitors?</h3>
<p>The main types are:</p>
<ul>
<li><strong>HTTP/API monitors</strong> - Check that URLs respond correctly</li>
<li><strong>Browser monitors</strong> - Simulate user interactions like clicking and form submission</li>
<li><strong>TCP/UDP monitors</strong> - Check that ports are open</li>
<li><strong>DNS monitors</strong> - Verify DNS resolution works</li>
<li><strong>SSL monitors</strong> - Track certificate expiration</li>
</ul>
<h3>Synthetic monitoring vs RUM - which do I need?</h3>
<p><strong>Synthetic monitoring</strong> proactively tests your services on a schedule. It catches outages and degradation before users are affected.</p>
<p><strong>Real User Monitoring (RUM)</strong> passively collects data from actual user sessions. It shows you what real users experience across different devices and network conditions.</p>
<p>Most teams should start with synthetic monitoring because it's simpler and catches the most critical issues. Add RUM when you need to understand real user experience in detail.</p>
<h3>How much does synthetic monitoring cost?</h3>
<p>Costs range from free (UptimeRobot free tier, self-hosted Uptime Kuma) to hundreds of dollars per month for enterprise tools.</p>
<p>For a typical SaaS with 20-50 monitors:</p>
<ul>
<li>Budget option: $7-15/month (UptimeRobot Pro)</li>
<li>Mid-range: $40-100/month (OnlineOrNot, Better Stack)</li>
<li>Enterprise: $200+/month (Datadog, based on usage)</li>
</ul>
<h3>Can I use multiple synthetic monitoring tools?</h3>
<p>Yes, and some teams do. Running two independent monitoring services provides redundancy - if one has an outage, the other still alerts you.</p>
<p>The downside is managing multiple dashboards and paying for both services. For most teams, one reliable tool is enough.</p>
<hr>
<h2>Related comparisons</h2>
<ul>
<li>If you only need scheduled URL checks, compare <a href="/uptime-monitoring-tools">uptime monitoring tools</a>.</li>
<li>If you are replacing UptimeRobot, read the <a href="/uptimerobot-alternative">UptimeRobot alternative</a> comparison or the broader <a href="/best-uptimerobot-alternatives">best UptimeRobot alternatives</a> guide.</li>
<li>If you are replacing Pingdom, read the <a href="/pingdom-alternative">Pingdom alternative</a> comparison or the broader <a href="/best-pingdom-alternatives">best Pingdom alternatives</a> guide.</li>
</ul>
<p>The best synthetic monitoring tool is the one you'll actually use. Start with something simple, get it running on your critical endpoints, and expand from there.</p>
<p>If you're not sure where to start, <a href="/">OnlineOrNot</a> offers a 14-day free trial with 30-second checks - no credit card required.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[API monitoring: A practical guide]]></title>
            <link>https://onlineornot.com/api-monitoring-guide</link>
            <guid>https://onlineornot.com/api-monitoring-guide</guid>
            <pubDate>Tue, 03 Mar 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Your API is the backbone of your product. Mobile apps, integrations, webhooks - they all depend on it working.</p>
<p>When it goes down, everything goes down with it.</p>
<p>The frustrating part is that API failures are often invisible. Your website might load fine, but the API returning data to your mobile app? Down for hours without anyone noticing.</p>
<p>Until customers start complaining.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#what-is-api-monitoring">What is API monitoring?</a></li>
<li><a href="#why-you-need-it">Why you need it</a></li>
<li><a href="#what-to-monitor">What to monitor</a></li>
<li><a href="#setting-up-api-monitoring">Setting up API monitoring</a></li>
<li><a href="#going-beyond-uptime">Going beyond uptime</a></li>
<li><a href="#common-mistakes">Common mistakes</a></li>
</ul>
<h2>What is API monitoring?</h2>
<p>API monitoring is checking that your API endpoints are working correctly - not just that they respond, but that they respond with the right data, fast enough.</p>
<p>At its simplest, you're making requests to your API at regular intervals and checking:</p>
<ol>
<li>Does it respond at all?</li>
<li>Does it return the correct status code?</li>
<li>Is the response what we expect?</li>
<li>How long did it take?</li>
</ol>
<p>If any of those checks fail, you get an alert.</p>
<p>It's basically the same idea as <a href="/uptime-monitoring">uptime monitoring</a> for websites, but with a few extra considerations for APIs.</p>
<h2>Why you need it</h2>
<p>There are a few scenarios where API monitoring saves you:</p>
<p><strong>Silent failures</strong> - Your API starts returning 500 errors, but your website still loads because it degrades gracefully. Users see broken features, but you don't see errors in your logs because the frontend swallows them.</p>
<p><strong>Slow degradation</strong> - Response times creep up from 200ms to 2 seconds over a few days. Not enough to trigger an error, but enough to make your mobile app feel sluggish. Users churn, and you don't know why.</p>
<p><strong>Third-party dependencies</strong> - You call another service's API, and they have an outage. Your API technically works, but returns empty data or errors for certain requests.</p>
<p><strong>Regional issues</strong> - Your API works fine from your office, but users in Europe are timing out because of a CDN misconfiguration.</p>
<p>In all these cases, monitoring catches the problem before your users do.</p>
<h2>What to monitor</h2>
<p>Not every endpoint needs the same level of monitoring. Here's how I think about it:</p>
<h3>Critical endpoints</h3>
<p>These are the endpoints that, if they break, your business stops working:</p>
<ul>
<li>Authentication (login, token refresh)</li>
<li>Core business logic (checkout, payments, data submission)</li>
<li>Public APIs your customers integrate with</li>
</ul>
<p>Monitor these frequently (every 30 seconds to 1 minute) from multiple regions. Set up aggressive alerting - phone calls, not just Slack messages.</p>
<h3>Important endpoints</h3>
<p>Endpoints that matter, but won't immediately break everything:</p>
<ul>
<li>User profile and settings</li>
<li>Search and filtering</li>
<li>Non-critical data fetches</li>
</ul>
<p>Monitor these every few minutes. Slack alerts are usually fine.</p>
<h3>Everything else</h3>
<p>Dashboard data, analytics, internal tools. If these break, it's annoying but not catastrophic.</p>
<p>Monitor less frequently, or don't monitor at all. Focus your attention on what matters.</p>
<h2>Setting up API monitoring</h2>
<p>Here's how to set up basic API monitoring with <a href="/api-monitoring">OnlineOrNot</a>:</p>
<p><strong>1. Start with your most critical endpoint</strong></p>
<p>Pick one endpoint - probably your main API health check or authentication endpoint. Don't try to monitor everything at once.</p>
<p><strong>2. Create the check</strong></p>
<p>You'll need:</p>
<ul>
<li>The URL to monitor</li>
<li>The HTTP method (GET, POST, etc.)</li>
<li>Any required headers (like <code>Authorization</code> or <code>Content-Type</code>)</li>
<li>The expected response (status code, maybe specific JSON fields)</li>
</ul>
<p><strong>3. Set the check interval</strong></p>
<p>For critical APIs, every 30 seconds is a good starting point. You can always adjust later.</p>
<p><strong>4. Add assertions</strong></p>
<p>This is where API monitoring gets more useful than basic uptime checks. You can verify:</p>
<ul>
<li>Status code is 200</li>
<li>Response contains specific JSON fields</li>
<li>Response time is under a threshold</li>
<li>Response body matches a pattern</li>
</ul>
<p>For example, if you're monitoring a <code>/health</code> endpoint that returns <code>{"status": "ok"}</code>, you'd assert that the response contains that exact JSON.</p>
<p><strong>5. Configure alerts</strong></p>
<p>Choose where alerts go. For critical endpoints, I recommend at least two channels:</p>
<ul>
<li>Immediate: Phone/SMS/PagerDuty for wake-you-up urgency</li>
<li>Awareness: Slack/Email for the rest of the team</li>
</ul>
<p><strong>6. Add more endpoints gradually</strong></p>
<p>Once your first check is running smoothly, add the next most critical endpoint. Build up your coverage over time.</p>
<h2>Going beyond uptime</h2>
<p>Basic "is it up?" monitoring is table stakes. Here are some things worth tracking as you mature:</p>
<h3>Response time trends</h3>
<p>A slow API is almost as bad as a down API. Track your p50, p95, and p99 response times over time. If your p95 suddenly jumps from 500ms to 2 seconds, something changed.</p>
<h3>Error rate</h3>
<p>What percentage of requests are returning errors? Even if the API is "up", a 5% error rate means 1 in 20 users is having a bad time.</p>
<h3>Multi-region checks</h3>
<p>Your API might work perfectly from AWS us-east-1 but be slow or broken from Europe. Running checks from multiple locations catches regional issues.</p>
<h3>Authenticated checks</h3>
<p>Your public health endpoint might return 200, but what about authenticated requests? Sometimes auth middleware breaks while the rest of the API is fine.</p>
<h3>Dependency health</h3>
<p>If your API depends on a database or third-party service, monitor those too. When something breaks, you want to know if it's your code or a dependency.</p>
<h2>Common mistakes</h2>
<p>A few things I've seen teams get wrong:</p>
<h3>Only monitoring the health endpoint</h3>
<p><code>/health</code> returning 200 doesn't mean your API works. I've seen health checks succeed while the actual API was completely broken because the database was down.</p>
<p>Monitor real endpoints that exercise real functionality.</p>
<h3>Ignoring response content</h3>
<p>Checking for a 200 status isn't enough. An API can return 200 with an error message in the body, or with empty data when there should be results.</p>
<p>Use assertions to verify the response makes sense.</p>
<h3>Too many alerts</h3>
<p>If every alert goes to the same place with the same priority, your team will start ignoring them. Differentiate between "wake someone up" and "look at this tomorrow".</p>
<p>I wrote more about this in <a href="/saving-your-team-from-alert-fatigue">saving your team from alert fatigue</a>.</p>
<h3>Not testing from where users are</h3>
<p>If all your monitors run from the same region as your servers, you'll miss regional issues. Users in Australia don't care that your API is fast from Virginia.</p>
<h3>Forgetting about dependencies</h3>
<p>Your API might be working perfectly, but if Stripe is down, your checkout is broken. Consider monitoring critical third-party APIs too, or at least having alerts for when they have issues.</p>
<hr>
<p>API monitoring isn't complicated, but it does require some thought about what matters most to your business.</p>
<p>Start with your most critical endpoint. Get that working well. Then expand from there.</p>
<p>The goal isn't to monitor everything - it's to make sure you find out about problems before your users do.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Cron job monitoring: How to know when your scheduled tasks fail]]></title>
            <link>https://onlineornot.com/cron-job-monitoring-guide</link>
            <guid>https://onlineornot.com/cron-job-monitoring-guide</guid>
            <pubDate>Tue, 03 Mar 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Cron jobs have a nasty habit of failing silently.</p>
<p>Your backup script stops running. Your daily report never gets sent. Your database cleanup job crashes halfway through. And you don't find out until someone asks "hey, why haven't we had a backup in three weeks?"</p>
<p>The problem is that cron doesn't care if your job succeeds or fails. It just runs the command and moves on. No alerts, no notifications, nothing.</p>
<p>Let's fix that.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#why-cron-jobs-fail-silently">Why cron jobs fail silently</a></li>
<li><a href="#how-cron-job-monitoring-works">How cron job monitoring works</a></li>
<li><a href="#setting-up-cron-job-monitoring">Setting up cron job monitoring</a></li>
<li><a href="#common-issues-with-cron-jobs">Common issues with cron jobs</a></li>
<li><a href="#what-to-monitor">What to monitor</a></li>
</ul>
<h2>Why cron jobs fail silently</h2>
<p>Cron was designed in the 1970s. Back then, the assumption was that a sysadmin would check the server regularly and notice if something was wrong.</p>
<p>That's not how most of us work today.</p>
<p>When a cron job fails, a few things might happen:</p>
<ol>
<li><strong>Nothing</strong> - The job crashes, cron shrugs, and nobody knows</li>
<li><strong>An email gets sent</strong> - Cron can email output to root, but who checks that?</li>
<li><strong>A log entry appears</strong> - Somewhere, buried in <code>/var/log/syslog</code></li>
</ol>
<p>None of these are great for catching problems quickly.</p>
<p>The real killer is when a job <em>doesn't run at all</em>. Maybe the server rebooted and cron didn't start. Maybe someone accidentally deleted the crontab. Maybe the disk filled up and cron couldn't write its lockfile.</p>
<p>In these cases, there's nothing to log. The job just... doesn't happen.</p>
<h2>How cron job monitoring works</h2>
<p>The solution is a "dead man's switch" (also called heartbeat monitoring).</p>
<p>The idea is simple:</p>
<ol>
<li>At the end of your cron job, you ping a URL</li>
<li>A monitoring service tracks these pings</li>
<li>If a ping doesn't arrive when expected, you get an alert</li>
</ol>
<p>It's called a dead man's switch because the alert triggers on <em>absence</em> of activity, not presence. If your job stops running for any reason - crash, server down, crontab deleted - you'll know.</p>
<p>Here's what it looks like in practice:</p>
<pre><code class="language-bash"># Before: Your cron job
0 2 * * * /home/user/backup.sh

# After: With monitoring
0 2 * * * /home/user/backup.sh &#x26;&#x26; curl -fsS --retry 3 https://oonchk.com/abc123
</code></pre>
<p>The <code>&#x26;&#x26;</code> is important - the curl only runs if <code>backup.sh</code> exits successfully. If your script fails, the ping doesn't get sent, and you get an alert.</p>
<h2>Setting up cron job monitoring</h2>
<p>Here's how to set it up with <a href="/cron-job-monitoring">OnlineOrNot</a>:</p>
<p><strong>1. Create a heartbeat monitor</strong></p>
<p>Give it a name (like "nightly-backup") and set the expected schedule. If your job runs daily at 2am, tell the monitor to expect a ping every 24 hours.</p>
<p><strong>2. Add a grace period</strong></p>
<p>Jobs don't always run at exactly the same time. A backup might take 5 minutes one day and 20 minutes the next. Set a grace period that accounts for normal variation.</p>
<p>For a daily job, 30-60 minutes of grace is usually fine. For an hourly job, maybe 10 minutes.</p>
<p><strong>3. Add the ping to your cron job</strong></p>
<p>You'll get a unique URL. Add it to the end of your cron command:</p>
<pre><code class="language-bash">0 2 * * * /home/user/backup.sh &#x26;&#x26; curl -fsS --retry 3 https://oonchk.com/your-unique-id
</code></pre>
<p>The flags:</p>
<ul>
<li><code>-f</code> - Fail silently on HTTP errors</li>
<li><code>-s</code> - Silent mode (no progress output)</li>
<li><code>-S</code> - Show errors if they occur</li>
<li><code>--retry 3</code> - Retry a few times if the network hiccups</li>
</ul>
<p><strong>4. Configure your alerts</strong></p>
<p>Choose where you want to be notified: Slack, email, PagerDuty, SMS, etc.</p>
<p>For critical jobs (like backups), I'd recommend at least two channels. Slack for awareness, phone/SMS for wake-you-up urgency.</p>
<p>That's it. If your job stops running, you'll know within minutes instead of weeks.</p>
<h2>Common issues with cron jobs</h2>
<p>Over the years, I've seen the same problems come up again and again:</p>
<h3>The job runs but fails</h3>
<p>Your script starts, hits an error, and exits with a non-zero status. If you're using <code>&#x26;&#x26;</code> before your ping (as above), this is caught automatically.</p>
<p>If you're not using <code>&#x26;&#x26;</code>, your monitoring will think everything's fine when it isn't.</p>
<h3>The job takes too long</h3>
<p>A job that usually takes 10 minutes suddenly takes 2 hours. This might be fine, or it might mean something's wrong.</p>
<p>Good monitoring tools let you track job duration, not just completion. If a job starts taking significantly longer than usual, that's worth investigating.</p>
<h3>The job overlaps with itself</h3>
<p>Your hourly job takes 90 minutes to run. Now you've got two instances running at once, possibly fighting over the same resources.</p>
<p>This isn't strictly a monitoring problem, but it's something to watch for. Use a lockfile or flock to prevent overlapping runs:</p>
<pre><code class="language-bash">0 * * * * flock -n /tmp/myjob.lock /home/user/myjob.sh &#x26;&#x26; curl ...
</code></pre>
<h3>Environment differences</h3>
<p>This one's classic: the job works when you run it manually, but fails under cron.</p>
<p>Cron runs with a minimal environment. Your <code>$PATH</code> is different, your shell config isn't loaded, and environment variables you take for granted aren't set.</p>
<p>Always use full paths in cron jobs, and set any required environment variables explicitly.</p>
<h3>The server rebooted</h3>
<p>Cron usually starts automatically on boot, but not always. And even if cron is running, your job might depend on other services that aren't ready yet.</p>
<p>This is where heartbeat monitoring really shines - if the job doesn't run after a reboot, you'll know.</p>
<h2>What to monitor</h2>
<p>Not every cron job needs monitoring. Here's how I think about it:</p>
<p><strong>Always monitor:</strong></p>
<ul>
<li>Backups (you really don't want to find out these stopped working when you need them)</li>
<li>Jobs that affect money (billing, invoices, payments)</li>
<li>Data pipelines and syncs</li>
<li>Security-related jobs (certificate renewal, log rotation)</li>
</ul>
<p><strong>Probably monitor:</strong></p>
<ul>
<li>Reports and notifications</li>
<li>Cleanup jobs</li>
<li>Health checks</li>
</ul>
<p><strong>Maybe skip:</strong></p>
<ul>
<li>Low-importance jobs you'd barely notice if they stopped</li>
<li>Jobs that have other visible effects when they fail</li>
</ul>
<p>The rule of thumb: if you'd want to know within an hour that this job stopped running, monitor it.</p>
<hr>
<p>Silent failures are the worst kind of failures. By the time you notice something's wrong, the damage is already done.</p>
<p>Adding a simple ping to your cron jobs takes about 30 seconds per job. It's one of those small investments that pays off enormously when things go wrong.</p>
<p>And things always go wrong eventually.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[MTTR: What it means and how to improve it]]></title>
            <link>https://onlineornot.com/mttr</link>
            <guid>https://onlineornot.com/mttr</guid>
            <pubDate>Tue, 03 Mar 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>It's 2am. Your phone buzzes with an alert. Your site is down.</p>
<p>You fumble for your laptop, try to remember where the logs are, SSH into the server, and eventually figure out the problem. By the time you've fixed it, it's 4am.</p>
<p>The next morning, someone asks: "How long did it take to resolve?"</p>
<p>That's your MTTR.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#what-is-mttr">What is MTTR?</a></li>
<li><a href="#the-four-types-of-mttr">The four types of MTTR</a></li>
<li><a href="#how-to-calculate-mttr">How to calculate MTTR</a></li>
<li><a href="#whats-a-good-mttr">What's a good MTTR?</a></li>
<li><a href="#mttr-vs-mtbf">MTTR vs MTBF</a></li>
<li><a href="#how-to-improve-your-mttr">How to improve your MTTR</a></li>
</ul>
<h2>What is MTTR?</h2>
<p>MTTR stands for Mean Time to Resolution (or Recovery, or Repair - more on that in a second). It measures the average time it takes your team to resolve incidents.</p>
<p>The idea is simple: the faster you can get your service back online, the less pain for your customers and your business.</p>
<p>If you had three incidents last month that took 30 minutes, 45 minutes, and 15 minutes to resolve, your MTTR would be 30 minutes.</p>
<h2>The four types of MTTR</h2>
<p>Here's where things get confusing. MTTR can actually mean four different things:</p>
<p><strong>Mean Time to Resolution</strong> - The total time from incident start to "everything's back to normal". This includes detection, diagnosis, fixing, and verification. This is the most common definition for software teams.</p>
<p><strong>Mean Time to Recovery</strong> - How long until service is restored. You might not know <em>why</em> it broke, but users can use the product again.</p>
<p><strong>Mean Time to Repair</strong> - Just the time spent actively fixing. Doesn't include waiting around or diagnosis. More common in hardware/manufacturing contexts.</p>
<p><strong>Mean Time to Respond</strong> - How quickly someone acknowledges the incident after an alert fires. Sometimes called MTTA (Mean Time to Acknowledge).</p>
<p>For most of us running web services, <strong>Mean Time to Resolution</strong> is what we care about - it captures the full customer impact.</p>
<p>The key thing is to pick one definition and stick with it. Otherwise your metrics become meaningless.</p>
<h2>How to calculate MTTR</h2>
<p>The formula is straightforward:</p>
<pre><code>MTTR = Total downtime / Number of incidents
</code></pre>
<p>So if you had 4 incidents last month with a combined downtime of 2 hours:</p>
<pre><code>MTTR = 120 minutes / 4 incidents = 30 minutes
</code></pre>
<p>That's it. Nothing fancy.</p>
<p>One thing to watch out for: make sure you're measuring from when the incident <em>actually started</em>, not when you first noticed it. If your site was down for an hour before anyone noticed, that hour counts.</p>
<p>This is why <a href="/uptime-monitoring">monitoring your uptime</a> matters - it reduces the gap between "incident started" and "someone's working on it".</p>
<h2>What's a good MTTR?</h2>
<p>It depends on what you're building, but here are some rough benchmarks from DORA's research:</p>
<p>| Performance | MTTR |
|-------------|------|
| Elite | Less than 1 hour |
| High | Less than 1 day |
| Medium | Less than 1 week |
| Low | More than 1 week |</p>
<p>If you're running a SaaS product, you probably want to aim for under an hour. For critical infrastructure like payments, under 15 minutes.</p>
<p>But here's the thing: <strong>don't obsess over the number</strong>.</p>
<p>I've seen teams game their MTTR by closing incidents early, or not counting certain outages as "real" incidents. That's worse than having a high MTTR, because now you're lying to yourself.</p>
<p>Focus on actually getting better at resolving incidents, and the number will follow.</p>
<h2>MTTR vs MTBF</h2>
<p>You'll often see MTTR mentioned alongside MTBF (Mean Time Between Failures).</p>
<p><strong>MTTR</strong> tells you how fast you fix things.</p>
<p><strong>MTBF</strong> tells you how often things break.</p>
<p>Both matter. A 5-minute MTTR is great, but if you're having incidents every day, something's fundamentally broken.</p>
<p>Ideally you want:</p>
<ul>
<li>High MTBF (things rarely break)</li>
<li>Low MTTR (when they do break, you fix them fast)</li>
</ul>
<p>If you're having lots of incidents, work on reliability first. If incidents are rare but take forever to resolve, work on your incident response.</p>
<h2>How to improve your MTTR</h2>
<p>Here are practical things you can do today:</p>
<h3>Detect incidents faster</h3>
<p>You can't fix what you don't know about. The time between "something broke" and "someone's looking at it" is often the biggest chunk of your MTTR.</p>
<p>Set up <a href="/uptime-monitoring">uptime monitoring</a> that checks your site frequently (every 30 seconds is ideal) from multiple locations. When something goes wrong, you want to know within minutes, not hours.</p>
<h3>Write runbooks</h3>
<p>At 3am, you won't remember exactly how to restart that one service, or where the logs are, or who to escalate to.</p>
<p>Write it down. Keep a document for each service with:</p>
<ul>
<li>What does this service do?</li>
<li>How do I check if it's working?</li>
<li>Common problems and how to fix them</li>
<li>Who to escalate to if you're stuck</li>
</ul>
<p>I've written more about this in <a href="/incident-management/incident-response/writing-your-first-runbooks">writing your first runbooks</a>.</p>
<h3>Reduce noise</h3>
<p>If your team is drowning in alerts, the important ones get lost. When everything's urgent, nothing is.</p>
<p>Review your alerts regularly. If an alert fires and there's no action to take, delete it or change its threshold. I wrote a whole article on <a href="/saving-your-team-from-alert-fatigue">saving your team from alert fatigue</a> if you want to dig deeper.</p>
<h3>Use status pages</h3>
<p>During an incident, your support team gets flooded with "is it down?" messages. This takes time away from actually fixing the problem.</p>
<p>A <a href="/status-pages">status page</a> lets customers check the status themselves. Less noise for your team, faster resolution.</p>
<h3>Run postmortems</h3>
<p>After every significant incident, figure out what happened and why. Not to blame anyone, but to learn.</p>
<p>Ask:</p>
<ul>
<li>What broke?</li>
<li>How did we find out?</li>
<li>What slowed us down?</li>
<li>How do we prevent this from happening again?</li>
</ul>
<p>The goal is to get a bit better each time. Over months, these small improvements compound.</p>
<h3>Keep it simple</h3>
<p>This one's harder to act on, but worth mentioning: complex systems have more ways to fail, and take longer to debug.</p>
<p>If you're a small team and your architecture looks like a distributed systems textbook, maybe it's time to simplify. I wrote about this more in <a href="/how-to-handle-monitoring-small-team">monitoring your web application as a small team</a>.</p>
<hr>
<p>MTTR isn't the only metric that matters, but it's a good one to track. It keeps you honest about how your team handles incidents.</p>
<p>Start measuring it, pick one thing to improve, and iterate from there.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from February 2026]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2026-february</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2026-february</guid>
            <pubDate>Tue, 03 Mar 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>February was all about making OnlineOrNot better for dev teams. I shipped audit logs for teams, added support for two-factor authentication and passkeys, unified the webhooks system across uptime checks, heartbeats, and status pages, and made it possible to connect multiple Discord channels.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-the-platform">Features for the platform</a>
<ul>
<li><a href="#audit-logs">Audit logs</a></li>
<li><a href="#two-factor-authentication-and-passkeys">Two-factor authentication and passkeys</a></li>
<li><a href="#webhooks-api-and-custom-api-token-expiry">Webhooks API and custom API token expiry</a></li>
</ul>
</li>
<li><a href="#features-for-checks">Features for Checks</a>
<ul>
<li><a href="#unified-webhooks">Unified webhooks</a></li>
<li><a href="#multiple-discord-channels">Multiple Discord channels</a></li>
</ul>
</li>
</ul>
</li>
</ul>
<h2>What's new</h2>
<h3>Features for the platform</h3>
<h4>Audit logs</h4>
<p>If you're on a team, at some point someone's going to change a check or update a status page and nobody will remember who did it, or when. You end up asking around in Slack, or worse, guessing.</p>
<p>OnlineOrNot now has an audit log. Every action taken in your organisation is logged: who did it, what they did, when, and from where. Creating checks, updating status pages, deleting maintenance windows, inviting team members - it's all there.</p>
<p>Each entry shows the actor (user or API token), the action performed, the IP address and location, and a timestamp. If the action had extra context (like which fields changed), you can expand the row to see the full metadata.</p>
<p>You'll find the audit log in the sidebar. It's paginated, so even if your team is busy, it stays fast.</p>
<p>The audit log also works through the REST API, so you can pull events programmatically if you need to feed them into your own logging pipeline or compliance tooling.</p>
<h4>Two-factor authentication and passkeys</h4>
<p>OnlineOrNot accounts have always been secured with email-based one-time passwords. Not great when your email service is down in the middle of an incident.</p>
<p>As of February 16th, OnlineOrNot supports two-factor authentication (2FA) using authenticator apps like 1Password, Authy, or Google Authenticator. You can enable it from your account settings: if your account has a password, scan the QR code with your authenticator app, and you're done. You'll also get backup codes in case you lose access to your authenticator.</p>
<p>On top of that, you can now register passkeys for your account. Passkeys let you log in with your fingerprint, face, or a hardware security key instead of typing a password. If you've used passkeys with GitHub or Google, it's the same idea. There's a "Login with passkey" button on the login page, and you can manage your passkeys in account settings.</p>
<p>Both features are optional. If you're happy with email-based login, nothing changes on your end.</p>
<h4>Webhooks API and custom API token expiry</h4>
<p>You've been able to create webhooks in the dashboard for a while now, but a few folks have asked about managing them through the API. Maybe you're spinning up new environments and don't want to click through the UI each time, or you'd rather have your webhooks defined alongside the rest of your infrastructure.</p>
<p>So I've added webhook management to the REST API. You can now list, create, update, and delete webhooks programmatically. Each webhook can subscribe to status page incident events (started, updated, resolved) and be scoped to specific status pages. You'll find the details in the <a href="https://developers.onlineornot.com/api/webhooks">API docs</a>.</p>
<p><img src="/assets/changelog/2026-02-08/api-token-expiry.png" alt="API token expiry"></p>
<p>While I was in there, I also added customisable API token expiry. You can now pick an expiration date when creating API tokens - handy for CI tokens or anything you'd rather not have floating around forever.</p>
<p><img src="/assets/changelog/2026-02-08/onlineornot-sidebar.png" alt="Sidebar showing Webhooks and API Tokens"></p>
<p>Both API tokens and Webhooks also have their own section in the sidebar now, which should make them easier to find.</p>
<h3>Features for Checks</h3>
<h4>Unified webhooks</h4>
<p>When I first added webhooks to OnlineOrNot, they only worked with status page incidents. Uptime checks and heartbeats had their own separate webhook alert system tucked away in the integrations settings, which meant you had to configure things in two different places.</p>
<p>I've now unified this. The webhooks you create in the Webhooks page can be attached to uptime checks, heartbeats, and status pages all in one place. You pick which events you care about (uptime.down, uptime.up, heartbeat.down, heartbeat.up, plus the existing status page incident events), and which resources should trigger them.</p>
<p>If you had generic webhooks configured before, they've been automatically migrated over. Your existing webhook URLs, check associations, and heartbeat associations are all there. The payload format hasn't changed either, so anything receiving your webhooks will keep working.</p>
<p>On-call integrations like PagerDuty, Opsgenie, Grafana OnCall, and Spike now live under <strong>Settings > Integrations > On-call</strong>.</p>
<h4>Multiple Discord channels</h4>
<p>Up until now, you could only connect one Discord channel to OnlineOrNot. If you wanted alerts in different channels for different checks, you were out of luck.</p>
<p>This worked fine for smaller teams, but once you start shipping a bit, having a single channel for all of your services' alerts starts to get noisy.</p>
<p>You can now connect multiple Discord channels, and pick which ones should receive alerts for each uptime check and heartbeat. The flow is the same as before: click "Add to Discord", authorize in Discord, and give your integration a name so you can tell them apart later.</p>
<p><img src="/assets/changelog/2026-02-10/discord-integration.png" alt="Discord channel naming"></p>
<p>If you already had a Discord integration set up, nothing changes on your end. Your existing checks will keep alerting to the same channel. You can add more channels whenever you need them.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Website monitoring: What it is and why you need it]]></title>
            <link>https://onlineornot.com/website-monitoring-guide</link>
            <guid>https://onlineornot.com/website-monitoring-guide</guid>
            <pubDate>Tue, 03 Mar 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Your website is down. You don't know it yet.</p>
<p>A customer tries to visit, gets an error, and leaves. Then another. And another. An hour passes before someone finally emails you: "Hey, is your site working?"</p>
<p>By then, the damage is done.</p>
<p>This is why website monitoring exists - to make sure you find out about problems in minutes, not hours.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#what-is-website-monitoring">What is website monitoring?</a></li>
<li><a href="#how-it-works">How it works</a></li>
<li><a href="#what-can-go-wrong">What can go wrong</a></li>
<li><a href="#setting-up-monitoring">Setting up monitoring</a></li>
<li><a href="#what-to-monitor">What to monitor</a></li>
<li><a href="#alerting-that-works">Alerting that works</a></li>
</ul>
<h2>What is website monitoring?</h2>
<p>Website monitoring is exactly what it sounds like: regularly checking that your website is working.</p>
<p>At its simplest, a monitoring service visits your site every few minutes and checks if it loads. If it doesn't, you get an alert.</p>
<p>But "working" can mean different things:</p>
<ul>
<li><strong>Is it reachable?</strong> Can users get to it at all?</li>
<li><strong>Is it fast?</strong> Does it load in a reasonable time?</li>
<li><strong>Is it correct?</strong> Does it show the right content?</li>
</ul>
<p>Good monitoring checks all three.</p>
<h2>How it works</h2>
<p>A monitoring service runs checks from servers around the world. Every 30 seconds (or whatever interval you set), it makes a request to your site and checks the response.</p>
<p>If your site returns a 200 status code and loads within your threshold, the check passes. If it returns an error, times out, or is missing expected content, the check fails.</p>
<p>When checks fail, you get notified. Most services let you configure multiple notification channels - email, Slack, SMS, phone calls - depending on how urgently you need to know.</p>
<p>The "from around the world" part matters. Your site might work perfectly from your office in London but be completely unreachable from Australia because of a CDN issue. Multi-region monitoring catches these problems.</p>
<h2>What can go wrong</h2>
<p>Websites go down for all sorts of reasons. Some common ones:</p>
<p><strong>Server crashes</strong> - Your web server runs out of memory, hits 100% CPU, or just dies. The process stops, and requests go unanswered.</p>
<p><strong>Deployment failures</strong> - You push a bad deploy. The new code crashes on startup, or has a bug that breaks the homepage.</p>
<p><strong>SSL certificate expiry</strong> - Your certificate expires, and browsers refuse to load your site. Users see scary security warnings.</p>
<p><strong>DNS issues</strong> - Your domain stops resolving. Maybe your registrar had an issue, maybe you forgot to renew, maybe someone misconfigured something.</p>
<p><strong>Database problems</strong> - Your database runs out of connections, fills up its disk, or just becomes unreachable. Your app can't load data, so pages fail.</p>
<p><strong>Third-party failures</strong> - A CDN, DNS provider, or cloud region goes down. Your code is fine, but users can't reach it.</p>
<p><strong>Traffic spikes</strong> - You get featured on Hacker News or Reddit. Your server can't handle the load and starts dropping requests.</p>
<p>The frustrating thing is that many of these failures are silent. Your server might be happily serving errors while you have no idea anything's wrong.</p>
<h2>Setting up monitoring</h2>
<p>Here's how to get started with <a href="/website-monitoring">website monitoring</a>:</p>
<p><strong>1. Start with your homepage</strong></p>
<p>Add your main URL. This is the most important page - if it's down, your whole site is effectively down.</p>
<p><strong>2. Set a reasonable check interval</strong></p>
<p>Every 30 seconds to 1 minute is good for most sites. Faster catches issues quicker but generates more traffic.</p>
<p><strong>3. Choose your regions</strong></p>
<p>Pick regions where your users actually are. If you're serving a US audience, check from US locations. If you're global, check from multiple continents.</p>
<p><strong>4. Add content verification</strong></p>
<p>Don't just check for a 200 status code. Add a text check to verify the page contains expected content.</p>
<p>For example, if your homepage always has "Welcome to Acme Corp", check for that text. This catches cases where your server returns 200 but serves an error page or blank content.</p>
<p><strong>5. Set up alerts</strong></p>
<p>Configure where notifications go. At minimum, email. For critical sites, also Slack and SMS.</p>
<p><strong>6. Test it</strong></p>
<p>Deliberately break something and verify you get an alert. You don't want to find out your alerting is broken during a real outage.</p>
<h2>What to monitor</h2>
<p>Start with the basics and expand from there:</p>
<h3>Essential</h3>
<ul>
<li><strong>Homepage</strong> - Your main entry point</li>
<li><strong>Login/signup pages</strong> - Critical for user access</li>
<li><strong>Core product pages</strong> - Whatever your users actually use</li>
</ul>
<h3>Important</h3>
<ul>
<li><strong>API endpoints</strong> - If you have an <a href="/api-monitoring">API</a>, monitor it separately</li>
<li><strong>Payment flows</strong> - Checkout pages, subscription management</li>
<li><strong>SSL certificate</strong> - Monitor expiry dates</li>
</ul>
<h3>Nice to have</h3>
<ul>
<li><strong>Marketing pages</strong> - Blog, docs, landing pages</li>
<li><strong>Third-party dependencies</strong> - Status pages of services you rely on</li>
<li><strong>Performance baselines</strong> - Alert if response time degrades significantly</li>
</ul>
<p>You don't need to monitor everything. Focus on pages where downtime actually hurts your business.</p>
<h2>Alerting that works</h2>
<p>The point of monitoring is to get alerts when things break. But alerting is easy to get wrong.</p>
<p><strong>Too many alerts</strong> and your team ignores them. Every alert becomes noise.</p>
<p><strong>Too few alerts</strong> and you miss real problems.</p>
<p><strong>Wrong channels</strong> and the right people don't see the alert in time.</p>
<p>Here's what I recommend:</p>
<h3>For critical pages</h3>
<ul>
<li>Check frequently (every 30 seconds)</li>
<li>Alert immediately on failure</li>
<li>Use aggressive channels: SMS, phone call, PagerDuty</li>
<li>Multiple team members should receive alerts</li>
</ul>
<h3>For important pages</h3>
<ul>
<li>Check every few minutes</li>
<li>Allow 2-3 failures before alerting (to filter out blips)</li>
<li>Use Slack or email</li>
<li>Route to the relevant team</li>
</ul>
<h3>For everything else</h3>
<ul>
<li>Check less frequently</li>
<li>Higher failure threshold before alerting</li>
<li>Email or a low-priority Slack channel</li>
</ul>
<p>The key is matching the alert urgency to the business impact. Your marketing blog being slow for 5 minutes is not the same as your checkout being down.</p>
<p>I wrote more about this in <a href="/saving-your-team-from-alert-fatigue">saving your team from alert fatigue</a>.</p>
<hr>
<p>Website monitoring is one of those things that feels unnecessary until you need it.</p>
<p>The first time you catch an outage in 2 minutes instead of 2 hours, you'll wonder how you ever operated without it.</p>
<p>Start simple: one check on your homepage, alerts going to your email. You can always add more later.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from January 2026]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2026-january</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2026-january</guid>
            <pubDate>Wed, 04 Feb 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Hopefully this will be one of the last major "behind-the-scenes" updates for a while, because OnlineOrNot's frontend now runs on a React framework that's easy to deploy across multiple providers, and is fully off GraphQL, being powered by its own <a href="https://developers.onlineornot.com/">REST API</a>.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-the-platform">Features for the platform</a>
<ul>
<li><a href="#expanded-rest-api">Expanded REST API</a></li>
</ul>
</li>
<li><a href="#features-for-checks">Features for Checks</a>
<ul>
<li><a href="#per-region-response-time-breakdown">Per-region response time breakdown</a></li>
</ul>
</li>
<li><a href="#features-for-status-pages">Features for Status Pages</a>
<ul>
<li><a href="#theyre-even-faster-now">They're even faster now</a></li>
<li><a href="#incident-improvements">Incident improvements</a></li>
</ul>
</li>
</ul>
</li>
</ul>
<h2>What's new</h2>
<h3>Features for the platform</h3>
<h4>Expanded REST API</h4>
<p>As of today, you can manage almost everything in OnlineOrNot programmatically:</p>
<ul>
<li>Heartbeat checks - create, list, update, pause, mute, and delete heartbeat checks</li>
<li>Maintenance windows - create, update, list, and delete maintenance windows</li>
<li>Status page incidents - create and manage incidents and incident updates, including per-component status changes</li>
<li>Scheduled maintenance - schedule, update, and cancel maintenance events on your status pages</li>
<li>Status page subscribers - manage email subscribers for your status pages</li>
<li>Status page components - add, reorder, and remove components</li>
<li>Team management - invite, list, and remove team members</li>
</ul>
<p>As I mentioned earlier, the OnlineOrNot dashboard itself has been rewritten to use the same REST API instead of GraphQL, so the API is a first-class citizen that can't go out of date or become unmaintained. As a side-effect, we get a faster and more reliable dashboard.</p>
<p>Every new endpoint follows the same conventions as the existing API, with consistent error responses and <a href="https://developers.onlineornot.com/">OpenAPI documentation</a>.</p>
<h3>Features for Checks</h3>
<h4>Per-region response time breakdown</h4>
<p>You know your API is slow, but is it slow everywhere, or just from one region? Maybe it's a database replica lagging in Europe, or a CDN cache miss in Asia-Pacific.</p>
<p>OnlineOrNot used to average response times across all monitoring regions into a single number, which made it hard to tell at a glance.</p>
<p><img src="assets/changelog/2025-12-03/regional-response-times.png" alt="OnlineOrNot regional response times"></p>
<p>Your uptime check detail pages now show a per-region breakdown of response times, so you can see exactly where things are slow and narrow down the cause faster.</p>
<h3>Features for Status Pages</h3>
<h4>They're even faster now</h4>
<p>When your service goes down, the last thing you want is for your status page to be slow to load too. Your customers are already stressed, and a sluggish status page doesn't help.</p>
<p>I've spent some time optimizing how OnlineOrNot's public status pages are served. Behind the scenes, status pages now run with targeted placement to minimise round-trip time, lazy loading has been replaced with optimized server-side queries, and the overall page weight has been reduced.</p>
<p>The result: your status page loads faster, and incidents display more quickly during the moments that matter most.</p>
<h4>Incident improvements</h4>
<p>Speaking of incidents: after an incident gets resolved, one of the first things your customers want to know is "how long was it down?"</p>
<p>Previously, they'd have to do the mental math themselves by comparing timestamps.</p>
<p>Resolved incidents on your status pages now display how long they took to resolve (e.g., "Resolved after 7 minutes"), so your customers can see the impact at a glance.</p>
<p><img src="assets/changelog/2025-12-01/status-page-timeline.png" alt="OnlineOrNot Incidents view"></p>
<p>I've also added incident timelines to the recent incidents section of your status page. Your customers can now see the full progression of an incident without having to click into each one individually.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Looking back on 2025, and what's next]]></title>
            <link>https://onlineornot.com/2025</link>
            <guid>https://onlineornot.com/2025</guid>
            <pubDate>Sun, 11 Jan 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Continuing on with the tradition I started to wrap up <a href="/2024">2024</a>, in this article I'll go over what's new in OnlineOrnot from 2025.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-went-live-in-2025">What went live in 2025</a>
<ul>
<li><a href="#general">General</a></li>
<li><a href="#checks">Checks</a></li>
<li><a href="#status-pages">Status Pages</a></li>
<li><a href="#under-the-hood">Under the hood</a></li>
</ul>
</li>
</ul>
<h2>What went live in 2025</h2>
<p>In 2025, I committed changes to OnlineOrNot 1347 times across 229 days (35% more than in 2024). The way I work remains unchanged - folks contact me about things they wish OnlineOrNot could do, I look at how that fits in with the existing roadmap, and I build it in the two hours a day I have before my workday starts.</p>
<p>It's always an odd feeling looking at these stats at the end of a year as a founder - you feel as though you didn't release nearly enough features that your customers ask for while simultaneously being impressed at the amount of work shipped.</p>
<h3>General</h3>
<ul>
<li>finally added <a href="/changelog#2025-05-27-dark-mode-and-more">dark mode</a> to the OnlineOrNot dashboard</li>
<li>added a <a href="/changelog#2025-08-04-microsoft-teams-integration">Microsoft Teams integration</a> for alerting</li>
<li>built a <a href="/changelog#2025-05-09-webhooks-for-status-pages">webhooks system</a> for receiving notifications whenever your status page updates</li>
<li>added a <a href="/changelog#2025-03-07-timezones-for-alerts">user-level timezone setting</a> so alert emails use your local timezone preference</li>
</ul>
<h3>Checks</h3>
<ul>
<li>added <a href="/changelog#2025-04-07-recurring-maintenance">maintenance windows</a> so you can avoid getting alerts during planned downtime</li>
<li>added the ability to <a href="/changelog#2025-07-07-alert-mute-checks">mute checks</a> when you need to temporarily silence alerts without pausing the check</li>
<li>added <a href="/changelog#2025-11-11-response-header-assertions">response header assertions</a> so you can verify Content-Type, caching headers, and more</li>
<li>added <a href="/changelog#2025-11-25-html-assertions">HTML body assertions</a> to check for specific content in responses</li>
<li>added <a href="/changelog#2025-06-24-confirmation-recovery-periods">recovery and confirmation periods</a> to reduce alert noise from flaky endpoints</li>
<li>built a new <a href="/changelog#2025-05-06-patch-endpoint-for-uptime-checks">HTTP API endpoint for updating checks</a></li>
<li>added an <a href="https://developers.onlineornot.com/api/checks#modify-a-check">HTTP API endpoint for pausing and muting checks</a></li>
</ul>
<h3>Status Pages</h3>
<ul>
<li>migrated status page queries to ClickHouse for significantly faster load times</li>
<li>added incident resolution time display so subscribers can see how long incidents took to resolve</li>
<li>improved the subscriber modal experience</li>
<li>added an <a href="/changelog#2025-08-28-rss-feed-for-status-pages">RSS feed</a> for status page updates</li>
<li>aligned status page dark mode with the main app's theme</li>
</ul>
<h3>Under the hood</h3>
<p>This section is a bit more technical than the others, and is mainly for folks that want to know how OnlineOrNot runs under the hood.</p>
<p>2025 was a big year for revisiting infrastructure decisions I made while building OnlineOrNot. While many decisions made sense for the product at the time they were made, the product has grown, and boring technologies that were hard to adopt in the past became much easier.</p>
<ul>
<li>Authentication rewrite
<ul>
<li>Migrated from Auth.js (originally known as NextAuth) to Better Auth, simplifying the codebase significantly and unlocking the ability to add SSO, password auth, and more.</li>
</ul>
</li>
<li>Migrated timeseries data from Postgres to Clickhouse
<ul>
<li>Despite being well-optimised with covering indexes, the queries powering many of OnlineOrNot's dashboards were slowing down when you had too many uptime checks. Some queries were taking 600ms on average, with spikes above 10 seconds. Moving that data into Clickhouse dropped that down to 60ms on average.</li>
</ul>
</li>
<li>Dropped Next.js and Remix, adopted React Router 7
<ul>
<li>Both the frontend for status pages (what you see when you load a status page such as the <a href="https://hackernews.onlineornot.com/">Hacker News status page</a>) and the frontend for OnlineOrNot now run React Router 7 and the latest version of React, significantly simplifying both frontends.</li>
</ul>
</li>
<li>Dropping GraphQL in favor of REST
<ul>
<li>While GraphQL let me iterate rapidly on OnlineOrNot's dashboard, each feature I built in GraphQL then had to be rewritten for REST. As a result, priorities got in the way, and the REST API lagged behind.</li>
<li>The fix is simple: all new code is written in REST, and I got to work migrating off GraphQL. The migration is still in progress, but it'll finish early in 2026.</li>
</ul>
</li>
</ul>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot's lessons from Cloudflare's outage on 2025-11-18]]></title>
            <link>https://onlineornot.com/onlineornot-lessons-from-cloudflare-outage-2025-11-18</link>
            <guid>https://onlineornot.com/onlineornot-lessons-from-cloudflare-outage-2025-11-18</guid>
            <pubDate>Wed, 19 Nov 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>On 2025-11-18 at 11:48 UTC, Cloudflare declared an <a href="https://www.cloudflarestatus.com/incidents/8gmgl950y3h7">incident</a> affecting the global network (that also affected OnlineOrNot). OnlineOrNot monitors websites, APIs, web apps, and cron jobs, while providing status pages as well.</p>
<p>While we partially mitigated the issue by enabling a fallback to AWS-based monitoring, between 13:00 UTC and 14:33 UTC failing checks went unreported, heartbeat checks over-reported, and status pages were unavailable. This postmortem outlines what went wrong, and what will be done to fix the issues.</p>
<p><strong>Table of contents:</strong></p>
<ul>
<li><a href="#onlineornots-dashboard-and-public-api">OnlineOrNot's dashboard and public API</a></li>
<li><a href="#onlineornots-uptime-checking-system">OnlineOrNot's uptime checking system</a></li>
<li><a href="#onlineornots-heartbeat-cron-monitoring-system">OnlineOrNot's heartbeat (cron) monitoring system</a></li>
<li><a href="#onlineornots-status-pages">OnlineOrNot's status pages</a></li>
<li><a href="#remediation-steps">Remediation steps</a></li>
</ul>
<h2>OnlineOrNot's dashboard and public API</h2>
<p>OnlineOrNot's frontend (<code>onlineornot.com</code>) is currently deployed in a single place: AWS Lambda (through Vercel) in us-east-1, with the DNS proxied on Cloudflare. The public API on the other hand, runs purely on Cloudflare Workers.</p>
<p>OnlineOrNot itself paged me to say that something was wrong with the public API at 11:44 UTC. Clicking on the link in the alert showed me Cloudflare's typical error screen, with a key difference - my host was still working, Cloudflare was the one with the error.</p>
<p>As a means of working around it, I disabled DNS proxying in OnlineOrNot's Cloudflare DNS settings, and got access back to the frontend. I then noticed the API was still down.</p>
<p>In disabling DNS proxying, I discovered a bug in OnlineOrNot's frontend that relied on Cloudflare's proxying to function properly.</p>
<p>It would remain down until 15:41 UTC, when I was able to login to the Cloudflare dashboard again, and re-enable DNS proxying. Had I taken no action with the DNS settings, OnlineOrNot's dashboard would have been accessible at 14:33 UTC.</p>
<h2>OnlineOrNot's uptime checking system</h2>
<p>OnlineOrNot's uptime checking system is both multi-cloud and multi-region, running in AWS and Cloudflare Workers. Having learned our lesson from <a href="/onlineornot-aws-outage-retrospective">a large us-east-1 outage in 2021</a>, the failover part of the system is frequently tested. Things still went wrong in new and different ways.</p>
<p>Each part of OnlineOrNot's Cloudflare system continuously reads config from KV, and writes health data to Workers Analytics Engine (WAE). There's also an autonomous Cloudflare Worker process that queries WAE every minute, and if the health of the system degrades, it updates the config in KV, and the system starts running on AWS.</p>
<p>During this incident, no part of this autonomous health system was functioning as expected, and I needed to manually trigger the CI process to redeploy the service to make it default to running in AWS. At 12:12 UTC, 28 minutes after noticing the disruption, OnlineOrNot was back up and running checks in AWS, and sending alerts.</p>
<p>These checks were however, less accurate than usual. The uptime checks in AWS relied on Cloudflare to double-check a URL is down. When the connection to Cloudflare was completely severed between 13:00 UTC and 14:33 UTC, the checks timed out in the queue, resulting in failing checks being skipped:</p>
<p><img src="/assets/onlineornot-lessons-from-cloudflare-outage-2025-11-18/skipped-checks.png" alt="OnlineOrNot missing checks"></p>
<p>(This chart displays time in UTC+1)</p>
<h2>OnlineOrNot's heartbeat (cron) monitoring system</h2>
<p>OnlineOrNot's heartbeat system suffered similar issues to the uptime checking system. Heartbeats provide healthcheck URLs (also hosted on Cloudflare), which were unreachable between 13:00 UTC and 14:33 UTC.</p>
<p>The other side of the system that looks for heartbeats to alert on successfully ran from AWS thanks to the failover.</p>
<p>As a result, the heartbeat system started sending alerts for cron jobs that were still running.</p>
<h2>OnlineOrNot's status pages</h2>
<p>OnlineOrNot's status pages are powered by a Cloudflare Worker serving the frontend, and a Cloudflare Worker proxying the private API.</p>
<p>During the outage, status pages were completely inaccessible between 13:00 UTC and 14:33 UTC.</p>
<p>For customers experiencing their own incidents at the same time as the Cloudflare outage, they had no way to update their status pages or show their users what was happening.</p>
<h2>Remediation steps</h2>
<p>Now that the systems are back online and functioning normally, I have begun work to harden them against failures like this in the future. In particular:</p>
<ul>
<li>Adding multi-cloud fallback services to all APIs and frontends - one vendor is not enough for services like OnlineOrNot</li>
<li>Eliminating the cross-cloud double checking on uptime checks in favor of same-region retries</li>
<li>Implementing automatic DNS failover for status page domains to redirect to the fallback version during outages</li>
<li>Reviewing and improving the autonomous health monitoring system that failed to trigger AWS failover automatically</li>
</ul>
<p>An outage like this is unacceptable for a monitoring system like OnlineOrNot. I've architected our systems to be multi-cloud and resilient, but this incident exposed gaps in our failover automation and dependencies I hadn't fully accounted for.</p>
<p>I apologize for the disruption this caused to your monitoring and incident response yesterday.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from August 2025]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2025-august</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2025-august</guid>
            <pubDate>Wed, 03 Sep 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>A bit more behind-the-scenes work than usual this month, but I still managed to ship some public-facing features you might be interested in. Logging in, and clicking around the dashboard just got 60% faster, and we're just getting started.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#general-features-for-onlineornot">General features for OnlineOrNot</a>
<ul>
<li><a href="#one-time-password-login">One-time password login</a></li>
</ul>
</li>
<li><a href="#features-for-checks">Features for Checks</a>
<ul>
<li><a href="#historical-overview">Historical overview</a></li>
</ul>
</li>
<li><a href="#features-for-status-pages">Features for Status Pages</a>
<ul>
<li><a href="#links-to-external-status-pages">Links to external status pages</a></li>
<li><a href="#rss-feed">RSS feed</a></li>
</ul>
</li>
</ul>
</li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>General features for OnlineOrNot</h3>
<h4>One-time password login</h4>
<p>To login to OnlineOrNot, you used to receive a magic link to let you sign-in on the same device. Various email-scanning providers broke this functionality from time to time, and unfortunately the experience wasn't as smooth as I'd have liked.</p>
<p>As a result, we now use one-time passwords:</p>
<p><img src="/assets/onlineornot-updates-from-2025-august/otp-email-login.png" alt="OnlineOrNot&#x27;s OTP login screen"></p>
<p>The idea remains the same (you'll get an email containing a code, and need it to login), except you can check your email on one device, and login on another, and email-scanning technology won't interfere with your login attempt.</p>
<h3>Features for Checks</h3>
<h4>Historical overview</h4>
<p>Up until recently, you couldn't fetch more than 24 hours of uptime data about a check in the OnlineOrNot dashboard.</p>
<p>In August, I took the time to start migrating OnlineOrNot's timeseries data into a Clickhouse instance, and now for the first time, you can query up to 30 days worth of data in a few hundred milliseconds:</p>
<p><img src="/assets/onlineornot-updates-from-2025-august/uptime-check-historical-overview.png" alt="OnlineOrNot&#x27;s uptime check overview"></p>
<p>(every paginated screen in OnlineOrNot now remembers how many results you wanted to see too)</p>
<p>There's more work to be done to finish the migration (we'll be moving heartbeats over too, and you'll be able to see more detailed logs), but it's looking promising.</p>
<p>Additionally, we're now using OnlineOrNot's own <a href="https://developers.onlineornot.com/api/checks#list-all-checks">public HTTP API</a> to power this screen, and over time we should have 100% of the functionality provided by the public API.</p>
<h3>Features for Status Pages</h3>
<h4>Links to external status pages</h4>
<p>Sometimes your providers have incidents, and you want to give your users more context about the issue. We now link to your external provider's status page when you link a third-party component:</p>
<p><img src="/assets/onlineornot-updates-from-2025-august/external-status-page-linking.png" alt="OnlineOrNot external status page linking"></p>
<h4>RSS feed</h4>
<p>There's now an RSS feed available on each status page:</p>
<p><img src="/assets/onlineornot-updates-from-2025-august/rss-feed.png" alt="OnlineOrNot&#x27;s status page RSS feed"></p>
<p>Letting your users subscribe to your status page's updates without needing to provide their email address.</p>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, or time to do it all myself.</p>
<p>OnlineOrNot sustains itself from folks like you enjoying it enough to tell your friends and colleagues about it, and if you don't enjoy it - let me know where it's falling short of your expectations.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from June/July 2025]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2025-june-july</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2025-june-july</guid>
            <pubDate>Mon, 04 Aug 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>In this latest update, I'll walk you through a few features I added that will make working with uptime checks less noisy, an alerts integration with Teams, and a few behind-the-scenes changes that will finally let me build mobile apps for OnlineOrNot.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-checks">Features for Checks</a>
<ul>
<li><a href="#alert-muting">Alert muting</a></li>
<li><a href="#confirmation-and-recovery-periods">Confirmation and Recovery periods</a></li>
<li><a href="#microsoft-teams-integration">Microsoft Teams integration</a></li>
</ul>
</li>
<li><a href="#behind-the-scenes-updates-no-one-will-notice">Behind-the-scenes updates no one will notice</a></li>
<li><a href="#new-reddit-community">New reddit community</a></li>
</ul>
</li>
</ul>
<h2>What's new</h2>
<h3>Features for Checks</h3>
<h4>Alert muting</h4>
<p>Sometimes you need to keep monitoring running, but temporarily silence the alerts. Maybe you're investigating an issue, or your service is in a known degraded state while you work on a fix.</p>
<p>OnlineOrNot now lets you manually mute individual uptime and browser checks, giving you the flexibility to silence alerts without stopping monitoring entirely.</p>
<p>When you mute a check, OnlineOrNot will:</p>
<ul>
<li>Continue running the uptime check</li>
<li>Keep collecting response time and availability data</li>
<li>Stop sending alerts and status page updates about downtime or issues</li>
</ul>
<p>Perfect for those times when you're aware of an issue and actively working on it, but don't want to be bombarded with alerts or confuse your users with status page notifications.</p>
<p><img src="assets/changelog/2025-07-07/mute-alerts.png" alt="OnlineOrNot - alert muting"></p>
<p>You can also <a href="https://developers.onlineornot.com/api/checks#modify-a-check">use the API</a> to mute or pause checks from your own codebase.</p>
<h4>Confirmation and Recovery periods</h4>
<p>Ever been in the middle of fixing an incident, and your service starts recovering but isn't quite 100% yet? Some requests succeed, others fail, and your monitoring keeps flip-flopping between "down" and "up" while you're still working on the fix?</p>
<p>You end up frustrated, and with noisy alerts and status pages that don't accurately reflect that the incident isn't fully resolved.</p>
<p>To solve this, OnlineOrNot now supports confirmation periods and recovery periods:</p>
<ul>
<li>Confirmation period - wait for consecutive failures over a specified time period before marking a service as down and sending alerts</li>
<li>Recovery period - wait for consecutive successes over a specified time period before marking a service as recovered and closing incidents</li>
</ul>
<p>For example, you can now configure a check to wait for five minutes of consecutive failures before alerting (avoiding false alarms), and wait for fifteen minutes of consecutive successes before marking the service as fully recovered (ensuring it's actually stable before closing the incident and confusing your users).</p>
<p>You can configure these settings in the "Advanced" section when creating or editing any uptime check.</p>
<p><img src="assets/changelog/2025-06-24/uptime-check-recovery-confirmation-periods.png" alt="OnlineOrNot - confirmation and recovery periods"></p>
<h4>Microsoft Teams integration</h4>
<p>OnlineOrNot now integrates with Microsoft Teams, allowing you to receive uptime and heartbeat alerts directly in your Microsoft Teams channels. With this integration, you'll get clear, actionable alerts when services go down, and recovery notifications when they're back up, as well as reminder alerts (if configured).</p>
<p>To get started, head to your OnlineOrNot settings, click the Integrations tab, and connect your Microsoft Teams channel. You can configure different channels for different services, just like with our other integrations.</p>
<p><img src="assets/changelog/2025-08-04/microsoft-teams-integration.png" alt="OnlineOrNot - Microsoft Teams integration"></p>
<h3>Behind-the-scenes updates no one will notice</h3>
<p>I'm not one to publicize these types of updates too much (most folks don't care that the SQL query powering a commonly-visited page got 80% faster), but this one's worth it.</p>
<p>In July I took the time to rewrite OnlineOrNot's auth engine (the thing that lets you sign-up/login). The original system served its purpose, and was starting to block new features from being built.</p>
<p>The new system runs faster, and we finally have the ability to add:</p>
<ul>
<li>Email + password login (for new users only for now, it will be possible to set a password for existing accounts soon)</li>
<li>SSO/SAML login (work in progress)</li>
<li>Mobile apps (work in progress)</li>
</ul>
<h3>New reddit community</h3>
<p>In case you missed it, there's a bunch of places (<a href="https://www.linkedin.com/company/onlineornot/">LinkedIn</a>, <a href="https://www.youtube.com/@OnlineOrNot">YouTube</a>, <a href="https://twitter.com/OnlineOrNot">Twitter</a>, <a href="https://discord.onlineornot.com/">Discord</a>) you can stay up to date on what I'm doing with OnlineOrNot.</p>
<p>As of this month, there's a new one: <a href="https://www.reddit.com/r/onlineornot/">reddit</a>.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from May 2025]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2025-may</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2025-may</guid>
            <pubDate>Mon, 02 Jun 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>As OnlineOrNot has grown, I've been building features quickly to get them into your hands as fast as possible. However, this meant I ended up with multiple versions of similar pages that looked and worked differently from each other. This month, I focused on putting systems in place to create a consistent experience across all parts of the dashboard, making everything look and feel unified.</p>
<p>In terms of what's new to see: there's a new navigation sidebar (so there's room for new features in the menu), and the whole app supports dark mode now.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-the-dashboard">Features for the dashboard</a>
<ul>
<li><a href="#dark-mode">Dark mode</a></li>
<li><a href="#heartbeats-in-greater-detail">Heartbeats in greater detail</a></li>
</ul>
</li>
</ul>
</li>
</ul>
<h2>What's new</h2>
<h3>Features for the dashboard</h3>
<h4>Dark mode</h4>
<p>While OnlineOrNot's status pages have supported dark mode for a long time now, I never found the time to update the dashboard to support it too (particularly when there were useful features to build).</p>
<p>Today, I'm happy to announce that the dashboard finally supports dark mode too:</p>
<p><img src="assets/changelog/2025-05-27/onlineornot-main-dash.png" alt="OnlineOrNot main dashboard with dark mode"></p>
<p>By default, OnlineOrNot will follow your system's dark mode setting, but you can also manually toggle it in the account dropdown in the sidebar:</p>
<p><img src="assets/changelog/2025-05-27/dark-mode-toggle.png" alt="OnlineOrNot main dashboard with dark mode, dropdown"></p>
<p>While adding dark mode, I also took the time to give the oldest parts of OnlineOrNot some polish:</p>
<ul>
<li>The main features are now in a mobile-friendly sidebar</li>
<li>We're now using a design system + component library to ensure a consistent user experience throughout the dashboard</li>
<li>Every date displayed in the dashboard now shows you how long ago it was, the date in your local timezone, as well as the date in UTC time</li>
</ul>
<h4>Heartbeats in greater detail</h4>
<p>Finally, heartbeat checks (what you use to monitor cron jobs and scheduled tasks) now display status changes in a table, rather than just in a graph:</p>
<p><img src="assets/changelog/2025-05-27/heartbeat-checks-glow-up.png" alt="OnlineOrNot Heartbeats UI glow-up"></p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from April 2025]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2025-april</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2025-april</guid>
            <pubDate>Fri, 09 May 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>In April I worked on a behind the scenes refactor, webhooks for status pages, a new endpoint, and more.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#something-entirely-new">Something entirely new</a></li>
<li><a href="#features-for-checks">Features for checks</a>
<ul>
<li><a href="#patch-endpoint-for-uptime-checks">PATCH endpoint for Uptime Checks</a></li>
</ul>
</li>
<li><a href="#features-for-status-pages">Features for status pages</a>
<ul>
<li><a href="#webhooks-for-status-pages">Webhooks for status pages</a></li>
</ul>
</li>
</ul>
</li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Something entirely new</h3>
<p>I spent a solid week or so this month rewriting how OnlineOrNot handles what each account can do. Before, OnlineOrNot would only let folks sign-up for a certain plan, and each plan would have predefined values for how many checks, status pages, and users they could use. Now, each account has its own separate limit, independent of the plan they're on.</p>
<p>In short, instead of forcing you to <a href="https://onlineornot.com/pricing">sign-up for a plan</a> that gives you 100 checks when you only need 5, you can now just use 5 checks.</p>
<p>Alternatively, if you need thousands of checks, you can also just sign-up for thousands of checks without needing OnlineOrNot staff to set-up a new plan for you.</p>
<h3>Features for checks</h3>
<h4>PATCH endpoint for Uptime Checks</h4>
<p>OnlineOrNot has a new early-release HTTP API PATCH endpoint for <a href="https://developers.onlineornot.com/api/checks#modify-a-check">updating uptime checks</a>.</p>
<p>As of today, the endpoint only supports updating your uptime check's HTTP request headers, but will be updated to support more features over time.</p>
<p>PS: In case you weren't aware, OnlineOrNot has an API! You create an API token from the OnlineOrNot dashboard, by going to Settings > Developers and selecting Create Token.</p>
<h3>Features for status pages</h3>
<h4>Webhooks for status pages</h4>
<ul>
<li>
<p>OnlineOrNot can now send you webhooks about your status page incidents.</p>
<p>To start with, the new webhooks system is simple and only supports sending the following events:</p>
<ul>
<li>status_page.incident.started</li>
<li>status_page.incident.updated</li>
<li>status_page.incident.ended</li>
</ul>
<p><img src="/assets/changelog/2025-05-09/statuspage-webhooks.png" alt="OnlineOrNot status page webhooks"></p>
<p>In the longer term, OnlineOrNot will also automatically migrate existing uptime check webhooks to the new system, so that it's easier to setup webhooks for multiple resources.</p>
</li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, or time to do it all myself.</p>
<p>OnlineOrNot sustains itself from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>As always: here's an ask</strong>:</p>
<p>I work on things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[How to get alerted when your EC2 instance shuts down]]></title>
            <link>https://onlineornot.com/how-to-get-alerted-on-ec2-instance-shutdown</link>
            <guid>https://onlineornot.com/how-to-get-alerted-on-ec2-instance-shutdown</guid>
            <pubDate>Tue, 29 Apr 2025 06:26:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Some of your most critical infrastructure runs on AWS EC2, so it's pretty damn important to know when your EC2 instances shut down.</p>
<p>Sure, chances are someone in your organisation will start kicking and screaming within 30 minutes of a particularly important instance shutting down, but we can do better than that.</p>
<p>When it comes to monitoring and customers (whether inside your org or outside), being proactive wins you a lot of points.</p>
<ul>
<li><a href="#aws-native">AWS Native</a>
<ul>
<li><a href="#cloudwatch">CloudWatch</a></li>
<li><a href="#eventbridge">EventBridge</a></li>
<li><a href="#dont-use-cloudtrail">(don't use) CloudTrail</a></li>
</ul>
</li>
<li><a href="#why-you-shouldnt-focus-on-server-status">Why you shouldn't focus on server status</a></li>
<li><a href="#external-monitoring-services">External monitoring services</a>
<ul>
<li><a href="#health-checks-and-heartbeat-monitors">Health checks and heartbeat monitors</a></li>
</ul>
</li>
</ul>
<h2>AWS Native</h2>
<p>There are a few ways to track whenever your EC2 instance shuts down natively within AWS:</p>
<h3>CloudWatch</h3>
<p>Your first idea might be to set up a CloudWatch alarm whenever CPU Utilization is at 0% for 5 minutes, which is close, but <strong>doesn't actually do the job</strong>. An instance shutting down doesn't actually send data about its CPU Utilization.</p>
<p>What you actually want to do is to tell CloudWatch to treat missing data as breaching the threshold.</p>
<p>You can see the <a href="https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/AlarmThatSendsEmail.html#alarms-and-missing-data">official AWS CloudWatch docs</a> for how to do that.</p>
<h3>EventBridge</h3>
<p>Using AWS EventBridge, you can get notified directly when your EC2 instance changes state.</p>
<p>Basically, you create an SNS topic with a subscription for however you want to receive the notification, then add an EventBridge rule to react to when your EC2 instance changes state, and point it at your SNS topic.</p>
<p>You can see the <a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-events-eventbridge-example.html">official AWS EC2/EventBridge docs</a> for how to do that.</p>
<h3>(don't use) CloudTrail</h3>
<p>You might be tempted to use AWS CloudTrail for this, but it's not the right tool for the job. CloudTrail is an audit log. It's more for tracking who in your organization ran which API call against your AWS resources, a while after the fact.</p>
<p>It won't catch your instance deciding to shut itself down, and it'll take a while to get data from it.</p>
<h2>Why you shouldn't focus on server status</h2>
<p>It doesn't directly answer your question (how do you get alerted when an EC2 instance shuts down), but you should probably reconsider your approach. Ideally, a human shouldn't get involved unless you <em>actually need human intervention</em> to restart your application.</p>
<p>What you should do instead is set up EC2 autoscaling with health checks so that the system recovers automatically whenever an instance shuts down (EC2 has <a href="https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-health-checks.html">docs</a> for this).</p>
<p>Basically: stop looking for infrastructure issues, and actually worry about application availability instead.</p>
<p>Once your application can heal itself by adding resources as needed, you can start to worry about getting alerts when your application becomes unreachable.</p>
<h2>External monitoring services</h2>
<p>I'm not going to sugar-coat it, this article was made possible because the author runs an external monitoring service, but even if I didn't, I would still recommend this route by default.</p>
<p>If your application runs an HTTP service, this is the quickest win by far.</p>
<p>Set up an external check to visit your application's HTTP endpoint/URL, and you'll receive an SMS, Email, Slack notification, Pager alert, and more if the application stops being available on the internet.</p>
<p>You can see a <a href="/uptime-monitoring-best-practices">quick guide</a> on how to set that up.</p>
<h3>Health checks and heartbeat monitors</h3>
<p>If your application doesn't expose a HTTP server, you can still use an external monitoring service.</p>
<p>You would just need to set up your application to ping a healthcheck endpoint on a regular basis (either at the instance-level using something like <code>cron</code>, or from within your application)</p>
<p>There's also a <a href="/docs/monitor-cron-job-scheduled-task">quick guide</a> for setting that up.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from 2025 Q1]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2025-first-quarter</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2025-first-quarter</guid>
            <pubDate>Tue, 08 Apr 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>I can't believe it's April already. OnlineOrNot now lets you automatically pause checks on a recurring basis with maintenance windows, there's better support for timezones, loading a <em>lot</em> of data got faster, and more.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-checks">Features for checks</a>
<ul>
<li><a href="#maintenance-windows">Maintenance windows</a></li>
<li><a href="#alerts-in-your-timezone">Alerts in your timezone</a></li>
</ul>
</li>
<li><a href="#features-for-status-pages">Features for status pages</a>
<ul>
<li><a href="#better-timezone-display">Better timezone display</a></li>
<li><a href="#external-status-pages">External status pages</a></li>
</ul>
</li>
<li><a href="#bug-fixes-and-improvements">Bug fixes and improvements</a></li>
</ul>
</li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Features for checks</h3>
<h4>Maintenance windows</h4>
<ul>
<li>
<p>OnlineOrNot defaults to checking your URLs and waiting for your cron jobs twenty-four hours a day, without a break. But what if every weeknight, you had a 30 minute window where you knew your service might go down?</p>
<p>OnlineOrNot now lets you pause your uptime checks and heartbeats for a given period of time, on a recurring basis:</p>
<p><img src="/assets/onlineornot-updates-2025-q1/recurring-maintenance.png" alt="OnlineOrNot maintenance windows"></p>
<p>Once you've set a recurring maintenance window, OnlineOrNot will update your checks and heartbeats to be in "maintenance" mode for the duration of the maintenance period.</p>
</li>
</ul>
<h4>Alerts in your timezone</h4>
<ul>
<li>
<p>OnlineOrNot now lets you select a timezone preference for the email alerts you receive.</p>
<p>This is particularly useful if you have team members in different timezones, and you want to make sure they receive alerts that are easy to understand without having to convert the timezone.</p>
<p>To update your timezone preference, click into your account settings
and select your timezone from the dropdown:</p>
<p><img src="/assets/onlineornot-updates-2025-q1/timezone-preferences.png" alt="OnlineOrNot custom timezone"></p>
</li>
</ul>
<h3>Features for status pages</h3>
<h4>Better timezone display</h4>
<ul>
<li>
<p>Status pages used to default to UTC time for all updates. That's probably fine when you're in Europe, but your users could be from anywhere. Every time now displays:</p>
<ul>
<li>How long ago something happened (relative time)</li>
<li>When something happened in UTC</li>
<li>When something happened in their local timezone</li>
</ul>
<p><img src="/assets/onlineornot-updates-2025-q1/statuspage-timezone-display.png" alt="OnlineOrNot status page timezone display"></p>
</li>
</ul>
<h4>External status pages</h4>
<ul>
<li>OnlineOrNot now supports monitoring <a href="/docs/third-party-status-pages#supported-external-status-pages">over a hundred external status pages</a> on your own status page</li>
</ul>
<h3>Bug fixes and improvements</h3>
<ul>
<li>Fixed a bug to ensure OnlineOrNot can keep checking even during an incident with our main service provider</li>
<li>Rewrote how OnlineOrNot shows multiple pages of data for checks, heartbeats, and status pages. It's now a <em>lot</em> faster to load a <em>lot</em> of data.</li>
<li>To help folks <a href="/cron-job-monitoring">monitoring their cron jobs</a>, I built a <a href="/generate-curl-command">cron command generator</a></li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, or time to do it all myself.</p>
<p>OnlineOrNot sustains itself from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>As always: here's an ask</strong>:</p>
<p>I work on things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Monitoring your web application as a small team]]></title>
            <link>https://onlineornot.com/how-to-handle-monitoring-small-team</link>
            <guid>https://onlineornot.com/how-to-handle-monitoring-small-team</guid>
            <pubDate>Wed, 22 Jan 2025 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>When you're part of a small team running a system with thousands of users or more, it can be pretty daunting to think about going on holiday, or even relaxing for a weekend.</p>
<p>"What if it goes down, and I'm not there to fix it?!" you ask yourself.</p>
<p>While you can never really guarantee that nothing will go wrong, you can take some steps to minimise your risk of things going wrong.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#before-you-start">Before you start</a>
<ul>
<li><a href="#choose-boring-technology">Choose boring technology</a></li>
<li><a href="#choose-technology-you-already-know">Choose technology you already know</a></li>
<li><a href="#opt-for-simplicity-over-complexity-especially-considering-your-team-size">Opt for simplicity over complexity, especially considering your team size</a></li>
</ul>
</li>
<li><a href="#while-running-the-business">While running the business</a>
<ul>
<li><a href="#use-managed-services-where-possible">Use managed services where possible</a></li>
<li><a href="#deploy-at-a-good-time">Deploy at a good time</a></li>
<li><a href="#write-two-way-database-migrations">Write two-way database migrations</a></li>
<li><a href="#write-good-error-pages">Write good error pages</a></li>
<li><a href="#avoid-alert-fatigue">Avoid alert fatigue</a></li>
<li><a href="#write-yourself-a-runbook">Write yourself a runbook</a></li>
<li><a href="#automate-your-recovery">Automate your recovery</a></li>
<li><a href="#keep-track-of-what-goes-wrong">Keep track of what goes wrong</a></li>
<li><a href="#run-post-mortem-meetings">Run post-mortem meetings</a></li>
</ul>
</li>
</ul>
<h2>Before you start</h2>
<h3>Choose boring technology</h3>
<p>This mainly comes to mind for databases, but can also apply to the frameworks you choose to use: pick boring technology that has been around for ages, with a community behind it. That way when things go wrong, you won't be the only person in the world trying to solve your problem.</p>
<p>An early startup I worked on used a NoSQL database for clearly very relational data. Even if you ignore the fact that we weren't using the right tool for the job - we would constantly run into issues that you couldn't find on StackOverflow/Google.</p>
<p>In fact, the <strong>only result on Google</strong> for our problems was often an article written by us.</p>
<p>Since that project, I've pushed for using Postgres (released in 1996) as much as possible, and in the last five years the only outage I've seen was due to under-provisioning (we tried to use a lower RAM server to save on running costs, and it backfired spectacularly).</p>
<h3>Choose technology you already know</h3>
<p>In a similar vein to picking boring technology, using technology you already know leaves you with just the "simple" task of building your product.</p>
<p>For example, if every project your team has worked on in the last year was built in Rails + React, maybe just use Rails + React to build your project. Building your product is time-intensive as it is, without having to worry about whether you're doing it <em>the $INSERT_LANGUAGE_HERE way</em>.</p>
<h3>Opt for simplicity over complexity, especially considering your team size</h3>
<p>By simplicity, I mean perhaps running your service on a sharded multi-region database isn't the best idea for your team of two devs - a single large database instance with redundancy would work just as well (up to a certain point).</p>
<p>While it may technically be "better" or "correct", when you have fewer resources to investigate problems, sometimes it's easier to just bump your database up a few resource tiers (scale vertically), rather than to scale onto multiple servers (scaling horizontally). Of course, if it looks like your service needs to handle hundreds of thousands of users, perhaps <strong>then</strong> you should consider horizontal scaling.</p>
<p>If you find yourself constantly getting alerted and having to fight fires to keep your service running, it might be time to simplify your architecture.</p>
<h2>While running the business</h2>
<h3>Use managed services where possible</h3>
<p>Sure, you can run your database on any VPS hosting provider in the world for cheap, but then it's on you to handle:</p>
<ul>
<li>Keeping the server updated</li>
<li>Monitoring the server</li>
<li>Ensuring the backup script runs</li>
<li>Keeping backups, deleting old ones</li>
<li>Fighting fires when things go wrong</li>
</ul>
<p>Alternatively, you can outsource the actual running of your database to AWS. When things go wrong, you then have access to their immense support resources to resolve the issue.</p>
<p>I take a similar approach with payments (Stripe), email (Postmark/ConvertKit), and tracking errors (Sentry/Bugsnag/Rollbar).</p>
<h3>Deploy at a good time</h3>
<p>Generally speaking, deploying before heading to lunch, dinner, a holiday, or a weekend trip isn't a great idea. You need to ensure you've got time to rollback, or hotfix the change you made.</p>
<p>Some people prefer only releasing at quiet times of the day, when the system doesn't have many users. The upside of this is that subsequent outages would impact the least number of users, but the downside is that it probably won't be the best time for <strong>you</strong>.</p>
<p>Depending on the number of users you have, it might be worth looking into <a href="https://martinfowler.com/articles/feature-toggles.html">feature flags</a>. Feature flags let you separate your deployments, from releases, when it comes to delivering features.</p>
<p>In other words, you "deploy" your change with the feature flag turned off, then once you've checked that it works on a small subset of users, you can roll out the change to your whole userbase.</p>
<h3>Write two-way database migrations</h3>
<p>Sometimes the fastest way to resolve an outage is to rollback the service to the last known "good version". A key part of this is to ensure your database migrations work in both directions - both when deploying the new version, and when tearing down the new version to release a previous version.</p>
<p>The alternative here is to roll forward with new database migrations when there's an issue, but I prefer being able to revert my database to the last known "good version".</p>
<h3>Write good error pages</h3>
<p>Communication is key. <a href="/what-fastly-outage-can-teach-about-writing-error-messages#we-can-write-better-error-messages">Good error pages</a> let the user know that it's not their fault. The last thing you want is for your user to feel stupid after trying to submit a form and your service not responding to the request.</p>
<p>On top of that, a bit of communication goes a long way towards reducing the number of users messaging you when things go wrong.</p>
<h3>Avoid alert fatigue</h3>
<p>I previously wrote about <a href="/saving-your-team-from-alert-fatigue">saving your team from alert fatigue</a>, but the gist of it is that you should tailor your alerting to the business impact of the outage.</p>
<p>If you've got a well-tested application where you know nothing should go wrong, you should monitor your uptime every minute, and send an alert <strong>immediately</strong> via phone call/SMS when the app becomes unresponsive.</p>
<p>On the other hand, if your business's legacy app becomes unreachable at the same time each day due to a database backup job running, you can setup your alerts to send after several minutes of downtime.</p>
<p>For more general alerts, like disk space reaching 75% on your server, or a database backup job failing, send the alerts to a Slack/Discord/Microsoft Teams channel as an FYI.</p>
<h3>Write yourself a runbook</h3>
<p><a href="/incident-management/incident-response/writing-your-first-runbooks">Keep a list of steps</a> to run through when your service has issues. It can act as a checklist so you don't miss anything when you inevitably try to fix things after getting an alert at 3am.</p>
<p>As an added bonus, when you start building your dev team, you'll already have a procedure in place for what to do when the service goes down - they won't have to relearn the same mistakes again.</p>
<p>Mike Julian's <a href="https://www.oreilly.com/library/view/practical-monitoring/9781491957349/">Practical Monitoring</a> recommends the following contents for a runbook:</p>
<ul>
<li>What is this service, what does it do?</li>
<li>Who is responsible for it?</li>
<li>What dependencies does it have?</li>
<li>What does the infrastructure for it look like?</li>
<li>What metrics and logs does it emit, and what do they mean?</li>
<li>What alerts are set up for it, and why?</li>
</ul>
<p>As well as that, make sure your alerts contain links to your runbook, so you're not franctically searching for it.</p>
<h3>Automate your recovery</h3>
<p>If parts of your runbook involve running a series of commands in a terminal, chances are you've got yourself a script you can automatically run without waking up a human.</p>
<p>It's worth keeping in mind - if the system can automatically recover, it's not worth waking up a human to check what happened.</p>
<h3>Keep track of what goes wrong</h3>
<p>Over time, you'll get a chance to observe your system in action. You'll find some parts of the codebase will cause more outages than others, which will nudge you to write better tests for it, or refactor to a better implementation, or better exception handling or validation.</p>
<p>Schedule yourself some time for these fixes, or at least add tasks to the backlog that let you track your effort to fixing these hot zones.</p>
<h3>Run post-mortem meetings</h3>
<p>You need to know what caused an issue to ensure you don't repeat the mistake again.</p>
<p>After you resolve incidents, be sure to give yourself some time (it doesn't <em>have to be</em> immediately after the incident) to review root causes, and come up with actions to ensure the incident doesn't happen again.</p>
<p>As your team grows, you want to foster a blameless culture around post-mortems. If your team fears being in trouble for mistakes, they'll either try to hide them or downplay their impact. Google's <a href="https://sre.google/sre-book/postmortem-culture/">SRE book</a> has some handy tips on building a postmortem culture.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[What is the curl command?]]></title>
            <link>https://onlineornot.com/curl-command</link>
            <guid>https://onlineornot.com/curl-command</guid>
            <pubDate>Wed, 08 Jan 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>curl is one of those programs that feels like its always been there for you in a pinch, like when you're trying to debug what your API is doing, and yet we never take the time to actually learn how to use it (I only ever used it via "copy as curl" from my browser's devtools, prior to writing this article).</p>
<p>It's insanely powerful too, run <code>curl --help all</code> to see what I mean.</p>
<p>In this article, we're going to take the time to learn what we can do with a <em>tiny subset</em> of curl's options, so we don't have to look them up every time.</p>
<p>Not a fan of reading? <a href="/generate-curl-command">Generate a curl command</a> instead.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#what-is-curl">What is curl?</a></li>
<li><a href="#curl-examples">curl examples</a>
<ul>
<li><a href="#how-to-make-a-curl-request">How to make a curl request</a></li>
<li><a href="#how-to-save-curl-output-to-a-file">How to save curl output to a file</a></li>
<li><a href="#how-to-make-a-post-request-with-curl">How to make a POST request with curl</a></li>
<li><a href="#how-to-post-custom-headers-with-curl">How to POST custom headers with curl</a></li>
<li><a href="#how-to-post-json-data-with-curl">How to POST JSON data with curl</a></li>
<li><a href="#how-to-post-a-json-file-with-curl">How to POST a JSON file with curl</a></li>
<li><a href="#how-to-change-the-user-agent-header-curl-uses">How to change the user-agent header curl uses</a></li>
<li><a href="#how-to-follow-redirects-with-curl">How to follow redirects with curl</a></li>
<li><a href="#how-to-see-the-full-http-request-and-response">How to see the full HTTP request and response</a></li>
<li><a href="#how-to-handle-cookies-with-curl">How to handle cookies with curl</a></li>
<li><a href="#how-to-simulate-different-http-methods">How to simulate different HTTP methods</a></li>
<li><a href="#how-to-handle-authentication-methods-with-curl">How to handle authentication methods with curl</a></li>
<li><a href="#how-to-handle-file-uploads">How to handle file uploads</a></li>
<li><a href="#how-to-time-your-requests">How to time your requests</a></li>
<li><a href="#putting-it-all-together">Putting it all together</a></li>
</ul>
</li>
<li><a href="#common-curl-command-patterns">Common curl command patterns</a></li>
</ul>
<h2>What is curl?</h2>
<p><a href="https://curl.se/">curl</a> is an open-source program that lets you transfer data to, or from a server, using URLs. If you've used it before, chances are you were just checking if an HTTP API was responding with the data you expected it to - but it supports <em>a lot more</em> than just HTTP.</p>
<p>Excerpt from <code>man curl</code>:</p>
<blockquote>
<p>curl is a tool for transferring data from or to a server using URLs. It supports these protocols: DICT, FILE, FTP, FTPS, GOPHER, GOPHERS, HTTP, HTTPS, IMAP, IMAPS, LDAP, LDAPS, MQTT, POP3, POP3S, RTMP, RTMPS, RTSP, SCP, SFTP, SMB, SMBS, SMTP, SMTPS, TELNET, TFTP, WS and WSS.</p>
</blockquote>
<p>That being said, in this article we're just going to learn how to use curl with HTTP, because that's already a challenge.</p>
<h2>curl examples</h2>
<p>Here are some examples of how to make requests with curl.</p>
<p>We'll be using https://echo.onlineornot.com - a free API that returns your request data back to you (useful for verifying your curl command does what you expect).</p>
<h3>How to make a curl request</h3>
<p>At its simplest, you can give curl any URL, and it will show you what's at that URL via a GET request:</p>
<pre><code class="language-bash">curl https://echo.onlineornot.com
</code></pre>
<p>(<a href="/generate-curl-command?url=https://echo.onlineornot.com">see this example in our curl generator</a>) and you get a response:</p>
<pre><code class="language-js">{
  "method": "GET",
  "pathname": "/",
  "searchParams": {},
  "headers": {
    "accept": "*/*",
    "accept-encoding": "gzip, br",
    "connection": "Keep-Alive",
    "host": "echo.onlineornot.com",
    "user-agent": "curl/8.7.1"
  },
  "body": "",
  "powered-by": "https://onlineornot.com"
}
</code></pre>
<p>As I mentioned earlier, if you don't specify the HTTP method you want to use, by default curl will send a GET request. If you want to send a POST request, you'll need to provide an additional flag, as shown <a href="/curl-command#make-a-post-request-with-curl">below</a>.</p>
<h3>How to save curl output to a file</h3>
<p>There are two ways to save curl's output to a file, instead of displaying the result in your terminal:</p>
<ol>
<li>Using the <code>-o</code> flag</li>
</ol>
<p>For example:</p>
<pre><code class="language-bash">curl -o output.json https://echo.onlineornot.com
</code></pre>
<p>(<a href="/generate-curl-command?url=https://echo.onlineornot.com&#x26;outputToFile=true&#x26;filename=output.json">see this example in our curl generator</a>) this creates a file in the same directory as where I ran the command, with the contents:</p>
<pre><code class="language-js">{
  "method": "GET",
  "pathname": "/",
  "searchParams": {},
  "headers": {
    "accept": "*/*",
    "accept-encoding": "gzip, br",
    "connection": "Keep-Alive",
    "host": "echo.onlineornot.com",
    "user-agent": "curl/8.7.1"
  },
  "body": "",
  "powered-by": "https://onlineornot.com"
}
</code></pre>
<ol start="2">
<li>Using STDIN redirection (the <code>></code> character)</li>
</ol>
<p>For example:</p>
<pre><code class="language-bash">curl https://echo.onlineornot.com > output.json
</code></pre>
<p>This creates the same output.json as before.</p>
<h3>How to make a POST request with curl</h3>
<p>You can specify the HTTP method using the <code>-X</code> flag, like so:</p>
<pre><code class="language-bash">curl -X POST https://echo.onlineornot.com
</code></pre>
<p>(<a href="/generate-curl-command?url=https://echo.onlineornot.com&#x26;method=POST">see this example in our curl generator</a>) which outputs the following in our terminal:</p>
<pre><code class="language-js">{
  "method": "POST",
  "pathname": "/",
  "searchParams": {},
  "headers": {
    "accept": "*/*",
    "accept-encoding": "gzip, br",
    "connection": "Keep-Alive",
    "host": "echo.onlineornot.com",
    "user-agent": "curl/8.7.1"
  },
  "body": "",
  "powered-by": "https://onlineornot.com"
}
</code></pre>
<p>POST requests are cool and all, but they're not much use to the APIs you're using if you can't authenticate, or tell the API which format the data you're sending it is in.</p>
<h3>How to POST custom headers with curl</h3>
<p>You can provide headers to the curl command using the <code>-H</code> or <code>--header</code> flag, like so:</p>
<pre><code class="language-bash"># we've used the backslash (\) below to split
# the command across multiple lines
curl -H 'Content-Type: application/json' \
      -X POST https://echo.onlineornot.com
</code></pre>
<p>(<a href="/generate-curl-command?url=https://echo.onlineornot.com&#x26;method=POST&#x26;headers=%5B%7B%22name%22:%22Content-type%22,%22value%22:%22application/json%22%7D%5D">see this example in our curl generator</a>) which outputs the following in our terminal:</p>
<pre><code class="language-js">{
  "method": "POST",
  "pathname": "/",
  "searchParams": {},
  "headers": {
    "accept": "*/*",
    "accept-encoding": "gzip, br",
    "connection": "Keep-Alive",
    "content-type": "application/json",
    "host": "echo.onlineornot.com",
    "user-agent": "curl/8.7.1"
  },
  "body": "",
  "powered-by": "https://onlineornot.com"
}
</code></pre>
<p>To add multiple headers, you'll need to add additional <code>-H</code> flags, like so:</p>
<pre><code class="language-bash">curl -H 'Content-Type: application/json' \
      -H 'Authorization: Bearer a_real_token_goes_here' \
      -X POST https://echo.onlineornot.com
</code></pre>
<p>Now that we've provided an <code>Authorization</code> header and a <code>Content-Type</code> header, we can send data to most APIs (that support JSON data, anyway).</p>
<h3>How to POST JSON data with curl</h3>
<p>In case you're wondering:</p>
<blockquote>
<p>but when would I send data TO an API?</p>
</blockquote>
<p>You would typically send data to an API when creating or updating records.</p>
<p>You can send data with your request using the <code>-d</code> or <code>--data</code> flag, like so:</p>
<pre><code class="language-bash">curl -H 'Content-Type: application/json' \
      -H 'Authorization: Bearer a_real_token_goes_here' \
      -d '{"name": "My website", "url": "https://onlineornot.com"}' \
      -X POST https://echo.onlineornot.com
</code></pre>
<p>(<a href="/generate-curl-command?url=https://echo.onlineornot.com&#x26;method=POST&#x26;headers=%5B%7B%22name%22:%22Content-type%22,%22value%22:%22application/json%22%7D,%7B%22name%22:%22Authorization%22,%22value%22:%22Bearer+a_real_token_goes_here%22%7D%5D&#x26;body=%7B%22name%22:+%22My+website%22,+%22url%22:+%22https://onlineornot.com%22%7D">see this example in our curl generator</a>) which outputs the following in our terminal:</p>
<pre><code class="language-js">{
  "method": "POST",
  "pathname": "/",
  "searchParams": {},
  "headers": {
    "accept": "*/*",
    "accept-encoding": "gzip, br",
    "authorization": "Bearer a_real_token_goes_here",
    "connection": "Keep-Alive",
    "content-length": "56",
    "content-type": "application/json",
    "host": "echo.onlineornot.com",
    "user-agent": "curl/8.7.1"
  },
  "body": "{\"name\": \"My website\", \"url\": \"https://onlineornot.com\"}",
  "powered-by": "https://onlineornot.com"
}
</code></pre>
<h3>How to POST a JSON file with curl</h3>
<p>You can also send files with curl using the <code>-d</code> flag as in the previous example, and adding <code>@</code> in front of the file path.</p>
<p>For example, if I was in a directory containing a file called <code>my_data.json</code> , I could use:</p>
<pre><code class="language-bash">curl -H 'Content-Type: application/json' \
      -H 'Authorization: Bearer a_real_token_goes_here' \
      -d @my_data.json \
      -X POST https://echo.onlineornot.com
</code></pre>
<p>(in this case, I've put the data from the previous example into <code>my_data.json</code>), which outputs the following in our terminal:</p>
<pre><code class="language-js">{
  "method": "POST",
  "pathname": "/",
  "searchParams": {},
  "headers": {
    "accept": "*/*",
    "accept-encoding": "gzip, br",
    "authorization": "Bearer a_real_token_goes_here",
    "connection": "Keep-Alive",
    "content-length": "56",
    "content-type": "application/json",
    "host": "echo.onlineornot.com",
    "user-agent": "curl/8.7.1"
  },
  "body": "{\"name\": \"My website\", \"url\": \"https://onlineornot.com\"}",
  "powered-by": "https://onlineornot.com"
}
</code></pre>
<h3>How to change the user-agent header curl uses</h3>
<p>By default, curl will use the following user-agent header for sending requests: <code>curl/{version_number_here}</code></p>
<p>For all of the requests in this article, curl has been sending the default user-agent (in my case, <code>curl/8.7.1</code>).</p>
<p>If you're building your own tool that fetches URLs on top of curl, you might want to <strong>customize</strong> the user-agent header curl sends. You can do that by passing a <code>user-agent</code> header, as in the <a href="#how-to-post-custom-headers-with-curl">how to post custom headers with curl</a> example above:</p>
<pre><code class="language-bash">curl -H 'user-agent: testing-echo' \
      https://echo.onlineornot.com
</code></pre>
<p>(<a href="/generate-curl-command?url=https://echo.onlineornot.com&#x26;headers=%5B%7B%22name%22:%22user-agent%22,%22value%22:%22testing-echo%22%7D%5D">see this example in our curl generator</a>) which outputs the following in our terminal:</p>
<pre><code class="language-js">{
  "method": "GET",
  "pathname": "/",
  "searchParams": {},
  "headers": {
    "accept": "*/*",
    "accept-encoding": "gzip, br",
    "connection": "Keep-Alive",
    "host": "echo.onlineornot.com",
    "user-agent": "testing-echo"
  },
  "body": "",
  "powered-by": "https://onlineornot.com"
}
</code></pre>
<h3>How to follow redirects with curl</h3>
<p>By default, curl won't automatically follow redirects (when a website sends you to a different URL). You can enable this with the <code>-L</code> or <code>--location</code> flag:</p>
<pre><code class="language-bash">curl -L https://echo.onlineornot.com/redirect
</code></pre>
<p>This is particularly useful when working with URLs that might redirect you to a login page or a different domain.</p>
<h3>How to see the full HTTP request and response</h3>
<p>When debugging APIs, it's often useful to see exactly what curl is sending and receiving. The <code>-v</code> (verbose) flag shows you the complete HTTP conversation:</p>
<pre><code class="language-bash">curl -v https://echo.onlineornot.com
</code></pre>
<p>This will show you:</p>
<ul>
<li>DNS lookup details</li>
<li>TLS/SSL handshake information</li>
<li>Request headers sent</li>
<li>Response headers received</li>
<li>The response body</li>
</ul>
<h3>How to handle cookies with curl</h3>
<p>Building on our knowledge of headers, sometimes you need to handle cookies. You can:</p>
<ol>
<li>Send a cookie:</li>
</ol>
<pre><code class="language-bash">curl -H "Cookie: session=123abc" https://echo.onlineornot.com
</code></pre>
<ol start="2">
<li>Save received cookies to a file:</li>
</ol>
<pre><code class="language-bash">curl -c cookies.txt https://echo.onlineornot.com
</code></pre>
<ol start="3">
<li>Use saved cookies in a request:</li>
</ol>
<pre><code class="language-bash">curl -b cookies.txt https://echo.onlineornot.com
</code></pre>
<h3>How to simulate different HTTP methods</h3>
<p>We've covered GET and POST, but REST APIs often use other HTTP methods. Here's how to use them:</p>
<pre><code class="language-bash"># PUT request
curl -X PUT -H "Content-Type: application/json" \
     -d '{"name": "Updated website"}' \
     https://echo.onlineornot.com

# DELETE request
curl -X DELETE https://echo.onlineornot.com/resource/123

# PATCH request
curl -X PATCH -H "Content-Type: application/json" \
     -d '{"name": "Partial update"}' \
     https://echo.onlineornot.com
</code></pre>
<h3>How to handle authentication methods with curl</h3>
<p>Building on our earlier Authorization header example, here's how to handle different auth methods:</p>
<ol>
<li>Basic Auth (username/password):</li>
</ol>
<pre><code class="language-bash">curl -u username:password https://echo.onlineornot.com
</code></pre>
<ol start="2">
<li>Bearer token (as shown <a href="/curl-command#how-to-post-custom-headers-with-curl">above</a>):</li>
</ol>
<pre><code class="language-bash">curl -H "Authorization: Bearer a_real_token_goes_here" https://echo.onlineornot.com
</code></pre>
<ol start="3">
<li>API Key in query string:</li>
</ol>
<pre><code class="language-bash">curl "https://echo.onlineornot.com?api_key=a_real_token_goes_here"
</code></pre>
<h3>How to handle file uploads</h3>
<p>For APIs that accept file uploads, you can use the <code>-F</code> flag for form data:</p>
<pre><code class="language-bash"># Upload a single file
curl -F "file=@local-file.pdf" https://echo.onlineornot.com/upload

# Upload multiple files with additional form fields
curl -F "files[]=@file1.pdf" \
     -F "files[]=@file2.pdf" \
     -F "description=My files" \
     https://echo.onlineornot.com/upload
</code></pre>
<h3>How to time your requests</h3>
<p>When working with APIs, you often want to know how long requests take. The <code>-w</code> flag lets you format timing data:</p>
<pre><code class="language-bash">curl -w "\nTime taken: %{time_total}s\n" https://echo.onlineornot.com
</code></pre>
<p>You can get detailed timing metrics with this format string:</p>
<pre><code class="language-bash">curl -w "\
    Time DNS lookup: %{time_namelookup}s\n\
    Time to connect: %{time_connect}s\n\
    Time to first byte: %{time_starttransfer}s\n\
    Total time: %{time_total}s\n" \
    https://echo.onlineornot.com
</code></pre>
<h3>Putting it all together</h3>
<p>Here's an example that combines multiple concepts we've learned:</p>
<pre><code class="language-bash">curl -v -L \
     -H "Content-Type: application/json" \
     -H "Authorization: Bearer a_real_token_goes_here" \
     -d '{"name": "Complex request"}' \
     -w "\nTotal time: %{time_total}s\n" \
     https://echo.onlineornot.com/api/v1/resource
</code></pre>
<p>This command:</p>
<ul>
<li>Shows verbose output (-v)</li>
<li>Follows redirects (-L)</li>
<li>Sets content type and auth headers</li>
<li>Sends JSON data</li>
<li>Shows timing information</li>
<li>Makes a POST request (implied by -d)</li>
</ul>
<h2>Common curl command patterns</h2>
<p>To wrap up, here are some common patterns you might use in your day-to-day work:</p>
<ol>
<li>API testing:</li>
</ol>
<pre><code class="language-bash">curl -v -H "Authorization: Bearer a_real_token_goes_here" https://echo.onlineornot.com
</code></pre>
<ol start="2">
<li>Download a file:</li>
</ol>
<pre><code class="language-bash">curl -L -o output.zip https://echo.onlineornot.com/file.zip
</code></pre>
<ol start="3">
<li>API creation request:</li>
</ol>
<pre><code class="language-bash">curl -X POST \
     -H "Content-Type: application/json" \
     -H "Authorization: Bearer token" \
     -d @create-payload.json \
     https://echo.onlineornot.com/resource
</code></pre>
<p>Remember, you can always combine these patterns and flags to build the exact curl command you need for your use case.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[2024, and what's next]]></title>
            <link>https://onlineornot.com/2024</link>
            <guid>https://onlineornot.com/2024</guid>
            <pubDate>Mon, 06 Jan 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>The other day I realized that while I write monthly updates on what's new in OnlineOrNot, and keep the <a href="/changelog">changelog</a> updated, I didn't have a yearly review of what was shipped, and what folks could expect from the coming year - so here we are!</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-went-live-in-2024">What went live in 2024</a>
<ul>
<li><a href="#general">General</a></li>
<li><a href="#checks">Checks</a></li>
<li><a href="#status-pages">Status Pages</a></li>
</ul>
</li>
<li><a href="#whats-next">What's next</a></li>
</ul>
<h2>What went live in 2024</h2>
<p>Taking a quick look through OnlineOrNot's code history for 2024 shows that I've made 1002 individual releases to OnlineOrNot's production environment over 235 working days, with quite a few bugfixes and UI tweaks too small to mention here.</p>
<p>The vast majority of these were implemented in the two hours I allow myself before full-time work, and sometimes I work on Saturday mornings.</p>
<p>Here are the highlights:</p>
<h3>General</h3>
<ul>
<li>adopted OpenAPI for <a href="/built-my-http-docs-from-scratch">automated API documentation</a>, and more</li>
<li>integrated with <a href="/onlineornot-updates-from-2024-april#onlineornot-x-incidentio-integration">incident.io</a> for alerting</li>
<li>rewrote OnlineOrNot for Cloudflare Workers, and backported the changes to AWS to allow the same code to be run redundantly across the two cloud providers</li>
</ul>
<h3>Checks</h3>
<ul>
<li><a href="/onlineornot-updates-from-2024-january#onlineornot-what-did-you-see">improved observability</a> for <em>why</em> OnlineOrNot thinks a check failed</li>
<li>added <a href="/onlineornot-updates-from-2024-march#reminder-alerts">reminder alerts</a> for all uptime checks, and scheduled task monitors</li>
<li>added <a href="/onlineornot-updates-from-2024-april#cron-expression-support">cron expression support</a> for monitoring cron jobs</li>
</ul>
<h3>Status Pages</h3>
<ul>
<li>made it possible to <a href="/onlineornot-updates-from-2024-april#status-page-component-groups">group status page components</a></li>
<li>made it possible to <a href="/onlineornot-updates-from-2024-march#external-status-pages">track external services</a> on your status page</li>
<li>made it possible to <a href="https://developers.onlineornot.com/api/status-pages#update-a-status-page">manage status pages</a> through the API</li>
<li>added <a href="/onlineornot-updates-from-2024-june-july#retrospective-incidents">retrospective incidents</a></li>
<li>added <a href="/onlineornot-updates-from-2024-june-july#incident-history">incident history</a> for every status page</li>
<li>added <a href="/onlineornot-updates-from-2024-june-july#ip-allowlisting">ip allowlisting</a> for finer-grained status page privacy</li>
<li>added <a href="/onlineornot-updates-from-2024-august#scheduled-maintenance">scheduled maintenance</a></li>
<li>made it possible to <a href="/onlineornot-updates-from-2024-september#custom-logos-and-favicons-for-status-pages">customize your favicon and logo</a></li>
<li>integrated with <a href="/onlineornot-updates-from-2024-august#pingdom-integration">pingdom</a> as a source of incidents</li>
<li>built an <a href="https://developers.onlineornot.com/api/status-pages#get-a-status-pages-summary-of-its-overall-status-components-active-incidents-and-scheduled-maintenance">API endpoint for retrieving a status page's summary</a></li>
</ul>
<h2>What's next</h2>
<p>As I mention on the <a href="/about">About</a> page, OnlineOrNot is 100% backed by its customers, and I expect to still be running it in 10 years. It goes without saying that I'm going to continue to focus on reliability, and ensuring OnlineOrNot is a tool that software teams can trust.</p>
<p>More specifically, I'm going to wrap up some status page features I didn't get to in 2024 (subscribe webhooks to a status page, display data from third-party sources, show a component's uptime history, incident postmortems, custom templates for creating incidents).</p>
<p>Uptime Checking and Cron Job Monitoring will see some love too, with better data display options in the dashboard and UI improvements.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[How OnlineOrNot more than halved its AWS bill]]></title>
            <link>https://onlineornot.com/how-onlineornot-halved-aws-bill</link>
            <guid>https://onlineornot.com/how-onlineornot-halved-aws-bill</guid>
            <pubDate>Mon, 28 Oct 2024 08:45:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Chances are, you've heard one of the main promises of serverless compute: "you only pay for what you use". While convenient when starting a new project, once you start to get continuous usage on that compute, you start to realize you're paying a hefty premium for that convenience.</p>
<p>If you're using AWS Lambda like I am, chances are that "for what you use" part isn't completely true either. If your function spends the majority of its time waiting for I/O to complete (fetching external URLs, querying DBs), you're paying a serverless premium mostly for your compute to sit around, waiting for something to do.</p>
<p>My solution to this was to migrate to <a href="https://developers.cloudflare.com/workers/platform/pricing/#workers">Cloudflare Workers</a>. With Workers, I went from paying up to 30000ms for a function to wait around while fetching a URL, to paying for the 1.4ms (p99) of CPU time to process the data before and after fetching.</p>
<p><strong>Table of contents:</strong></p>
<ul>
<li><a href="#the-project">The project</a></li>
<li><a href="#but-why-dont-you-just-use-vms">"But why don't you just use VMs?"</a></li>
<li><a href="#why-workers-worked-out-to-be-significantly-cheaper">Why Workers worked out to be significantly cheaper</a>
<ul>
<li><a href="#and-then-workers-started-talking-to-postgres">and then Workers started talking to Postgres</a></li>
</ul>
</li>
<li><a href="#the-approach-to-moving-off-aws-lambda">The approach to moving off AWS Lambda</a></li>
<li><a href="#rolling-out-the-change">Rolling out the change</a></li>
<li><a href="#its-faster-too">It's faster, too?</a></li>
<li><a href="#summary">Summary</a></li>
</ul>
<h2>The project</h2>
<p><a href="https://onlineornot.com/">OnlineOrNot</a> is a bootstrapped business that provides status pages with uptime monitoring attached, for software teams. When I started out, it was the 200th alternative uptime monitor for the Internet. To save money on the off-chance no one other than me used it, I built it on AWS Lambda. It did indeed cost nothing while no one was using it.</p>
<p>Building an uptime checker is a simple matter of code, you need:</p>
<ul>
<li>a process to query a database, find checks that are ready to be run, and queue them up in the right location</li>
<li>a process that picks up checks, runs them, and sends the results somewhere</li>
</ul>
<p>On AWS Lambda, this worked out to be a single function for picking up checks, an SNS topic, another function for running checks, a queue for batching writes to the DB, and a final function for writing results to the DB.</p>
<p>Fast-forward to a few weeks ago, and the vast majority of OnlineOrNot's cloud spend came from paying for AWS Lambda, for functions that spent most of their time waiting for things to do.</p>
<h2>"But why don't you just use VMs?"</h2>
<p>I've tried moving off of AWS Lambda before, too.</p>
<p>More than two years ago, after running OnlineOrNot for a bit and seeing costs steadily grow, I tried moving the service to continuously running VMs. The particular vendors I picked at the time struggled to maintain more than a few nines of reliability, which is awkward when you're trying to be accurate at measuring the reliability of others.</p>
<p>Long story short, I moved back to AWS, happy to continue paying the premium, knowing it would at least be reliable.</p>
<h2>Why Workers worked out to be significantly cheaper</h2>
<p>If you're curious, there's a good explanation of <a href="https://developers.cloudflare.com/workers/reference/how-workers-works/">how Workers works</a>, but the gist of it is: classic serverless platforms run a chunky VM under the hood, which can't be easily paused and restarted while waiting for I/O to complete, while Workers is built on the V8 runtime where the isolates that run your code are so lightweight that they <em>can</em> be easily restarted.</p>
<p>The most expensive part of my AWS Lambda functions was waiting for DB queries to finish and fetching external URLs - Cloudflare doesn't bill for that.</p>
<p>So while Workers does provide significantly cheaper compute, what stopped me (until recently) from using them as the main compute for my own applications was the inability to talk directly to my existing database.</p>
<h3>and then Workers started talking to Postgres</h3>
<p>Cloudflare <a href="https://blog.cloudflare.com/hyperdrive-making-regional-databases-feel-distributed/">announced Hyperdrive</a> back in late 2023, it became <a href="https://blog.cloudflare.com/making-full-stack-easier-d1-ga-hyperdrive-queues/">generally available</a> in April 2024, and it was <em>exactly</em> what I was looking for to unlock building entire applications on Workers.</p>
<p>Hyperdrive lets you query Postgres/MySQL databases from Cloudflare Workers (it's a connection pooler with an optional query cache service built-in).</p>
<h2>The approach to moving off AWS Lambda</h2>
<p>I was initially worried about how much time I'd spend rewriting, but it turned out that all of my AWS Lambda code was compatible with Cloudflare Workers. In the move I decided to change my Postgres client library from pg-promise to Postgres.js (Hyperdrive's recommended library), and remove the queue for batching writes, since I now had a connection pooler.</p>
<p>I started by creating a new Durable Object with a single purpose: run an alarm every 15 seconds, and start the checker.</p>
<p>Next, I moved the check-fetcher function across. On AWS, this would talk to SNS topics to send checks to the right region. I needed a Cloudflare-native means of placing Workers in different regions around the world. To do this, I reached for Durable Objects, with their <a href="https://developers.cloudflare.com/durable-objects/reference/data-location/#provide-a-location-hint">location hints</a>.</p>
<p>So I created another new Durable Object cleverly named "placer" with a single purpose: run a Worker in the region you wake up in.</p>
<p>The Worker itself I copied over from AWS Lambda, and since it used native JavaScript, it didn't need rewriting.</p>
<p>Finally I moved over the function to upload results to the database, it too was fully compatible with Workers, but I chose to change from pg-promise's <code>db.one/db.many</code> syntax to Postgres.js's <code>sql</code> syntax because I liked it better.</p>
<h2>Rolling out the change</h2>
<p>To start with, every check in my database has a "version" field, that I use to control which runtime it should run on. This made for a nice feature-flag for opting customers into the new system.</p>
<p>I wasn't <em>too</em> worried, as I had been using Workers to double-check failing uptime checks from AWS for years now, but there are always edge-cases.</p>
<p>I started by running my personal checks on Cloudflare for a few days, before finding some early adopters willing to try it out, then moving over the free tier, followed by the paid tier of users.</p>
<h2>It's faster, too?</h2>
<p>While it's well documented that when you allocate less RAM to an AWS Lambda function, you also get less CPU, I didn't realize that also meant they were throttling network performance.</p>
<p>In moving OnlineOrNot's checks to run primarily from Cloudflare, I noticed response time dropped by up to 75%:</p>
<p><img src="/assets/how-onlineornot-halved-aws-bill/faster-checks.png" alt="OnlineOrNot gets faster checks"></p>
<p>Attempting to investigate further lead me to <a href="https://repost.aws/questions/QUV7aqLIEqTwuUvr9TY9dvoQ/does-lambda-memory-allocation-impact-network-bandwidth">this answer</a> on the AWS forums:</p>
<blockquote>
<p>Yes, the network performance will increase with memory. There are no specifics as there is no definitive performance since Lambda compute is abstracted from you and it is also protocol and application specific. You should benchmark your function with various memory sized to determine where you meet your requirements.</p>
</blockquote>
<h2>Summary</h2>
<p>To wrap it up, OnlineOrNot is now "multi-cloud": there are multiple replicas of the system ready to run (for free while no one uses it) on AWS, and a version of the system running natively on Cloudflare, where checks complete faster, with running costs of less than half of what I was paying for AWS Lambda.</p>
<p>As we are bootstrapped, this accelerates the timeline for hiring additional help for the business, and will improve our support and product overall in the long term.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from September 2024]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2024-september</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2024-september</guid>
            <pubDate>Mon, 07 Oct 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>A quieter month this time around: you can now customize your status pages, there are additional overall statuses for status pages, and I fixed a few bugs.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-status-pages">Features for status pages</a>
<ul>
<li><a href="#custom-logos-and-favicons-for-status-pages">Custom logos and favicons for Status Pages</a></li>
<li><a href="#changing-how-overall-status-works">Changing how overall status works</a></li>
</ul>
</li>
<li><a href="#bug-fixes-and-improvements">Bug fixes and improvements</a></li>
</ul>
</li>
<li><a href="#whats-next">What's next</a></li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Features for status pages</h3>
<h4>Custom logos and favicons for Status Pages</h4>
<p>For a while now, OnlineOrNot's status pages all looked the same: a brand name, followed by some uptime bars:</p>
<p><img src="/assets/onlineornot-updates-from-2024-september/old-status-pages.png" alt="Scheduled maintenance - OnlineOrNot Status Page"></p>
<p>This worked fine if you wanted a bare-bones status page, but what if you wanted it to display your logo and favicon? Now you can!</p>
<p>In the Settings tab of your status page, you can now customize your favicon, and your light-mode and dark-mode logos:</p>
<p><img src="/assets/onlineornot-updates-from-2024-september/customise-your-status-page.png" alt="Scheduled maintenance - OnlineOrNot Status Page"></p>
<p>This way, your status page can match your website's design, and you'll keep your designers happy:</p>
<p><img src="/assets/onlineornot-updates-from-2024-september/new-status-pages.png" alt="Scheduled maintenance - OnlineOrNot Status Page"></p>
<h4>Changing how overall status works</h4>
<p>Before this update, status pages had two overall statuses: "All Systems Operational" and "There is an ongoing incident". They were linked to incidents, and the only way to update your overall status was by creating or resolving an incident.</p>
<p>With this update, overall status on a status page is now controlled by components. Changing an individual component's status to "Major outage", "Partial outage", "Degraded performance" or "Maintenance" will also change your page's status:</p>
<p><img src="/assets/onlineornot-updates-from-2024-september/status-page-statuses.png" alt="Additional statuses - OnlineOrNot Status Page"></p>
<h3>Bug fixes and improvements</h3>
<ul>
<li>Fixed a bug in the uptime percentage calculation on status pages. It would previously ignore incidents on the first day of a component's existence</li>
<li>Status pages now load approximately 500ms faster thanks to a performance optimization</li>
</ul>
<h2>What's next</h2>
<p>Here's what is still planned for the coming months:</p>
<ul>
<li>Ability to subscribe to status pages via webhook</li>
<li>Ability to push third-party metrics and plot graphs to showcase for customers</li>
<li>Display historical uptime stats for the past year in a calendar</li>
<li>Ability to create or update incident updates through APIs</li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, or time to do it all myself.</p>
<p>OnlineOrNot sustains itself from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>As always: here's an ask</strong>:</p>
<p>I work on things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from August 2024]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2024-august</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2024-august</guid>
            <pubDate>Wed, 04 Sep 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>This month I continued adding features to OnlineOrNot's status page functionality, while also taking the time to fix bugs, and add an integration for Pingdom.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-status-pages">Features for status pages</a>
<ul>
<li><a href="#scheduled-maintenance">Scheduled Maintenance</a></li>
<li><a href="#custom-favicon-and-logo">Custom favicon and logo</a></li>
<li><a href="#pingdom-integration">Pingdom integration</a></li>
</ul>
</li>
<li><a href="#bug-fixes-and-improvements">Bug fixes and improvements</a></li>
</ul>
</li>
<li><a href="#whats-next">What's next</a></li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Features for status pages</h3>
<h4>Scheduled Maintenance</h4>
<p>Sometimes you realise ahead of time that you need to take your service offline for a bit while you run updates. You already have enough on your plate, why worry about keeping the status page manually updated too?</p>
<p>You can now keep your customers updated on how your scheduled maintenance is progressing, and when they can use your service again, automatically:</p>
<p><img src="/assets/onlineornot-updates-from-2024-august/scheduled-maintenance.png" alt="Scheduled maintenance - OnlineOrNot Status Page"></p>
<h4>Custom favicon and logo</h4>
<p>We now support custom favicons and logos in status pages:</p>
<p><img src="/assets/onlineornot-updates-from-2024-august/custom-logo-favicon.png" alt="Scheduled maintenance - OnlineOrNot Status Page"></p>
<p>The process requires emailing OnlineOrNot's support, but I figured getting the feature into people's hands was more important than making it fully self-serve.</p>
<p>You can read more about it in the <a href="/docs/status-pages-custom-logo-favicon">documentation</a>.</p>
<h4>Pingdom integration</h4>
<p>You might already have uptime checks that you've been running for years in Pingdom. It does everything you need, and it's a total pain to move uptime monitoring when you've had it running for years, I get it.</p>
<p>What if you still wanted an OnlineOrNot status page for your Pingdom uptime checks?</p>
<p>You're in luck, because OnlineOrNot now integrates with Pingdom to update your status page:</p>
<p><img src="/assets/onlineornot-updates-from-2024-august/onlineornot-pingdom-settings.png" alt="Pingdom integration - OnlineOrNot Status Page"></p>
<h3>Bug fixes and improvements</h3>
<ul>
<li>External status page integrations now update every minute</li>
<li>Fixed bugs in the GitLab and Paddle external status page integration</li>
<li>Fixed a bug that impacted email deliverability when subscribing to a status page</li>
</ul>
<h2>What's next</h2>
<p>Here's what I have planned for these coming months:</p>
<ul>
<li>Upload your own favicon and logo to status pages at any time</li>
<li>Ability to push third-party metrics and plot graphs to showcase for customers</li>
<li>Display historical uptime stats for the past year in a calendar</li>
<li>Ability to subscribe to status pages via webhook</li>
<li>Ability to create or update incident updates through APIs</li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, or time to do it all myself.</p>
<p>OnlineOrNot sustains itself from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>As always: here's an ask</strong>:</p>
<p>I work on things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from June/July 2024]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2024-june-july</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2024-june-july</guid>
            <pubDate>Tue, 06 Aug 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>If you remember the <a href="/onlineornot-updates-from-2024-may">last update</a>, I originally planned to ship quite a few features in June. Life and family had to take priority for me, so they didn't get done.</p>
<p>Thankfully, I'm building OnlineOrNot for the extremely long term, so there's no rush to ship as fast as possible.</p>
<p>A fair few things <em>did</em> get shipped though.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#features-for-status-pages">Features for status pages</a>
<ul>
<li><a href="#retrospective-incidents">Retrospective incidents</a></li>
<li><a href="#descriptions-and-hiding-from-search-engines">Descriptions and hiding from search engines</a></li>
<li><a href="#incident-history">Incident history</a></li>
<li><a href="#ip-allowlisting">IP allowlisting</a></li>
<li><a href="#additional-external-status-pages">Additional external status pages</a></li>
</ul>
</li>
</ul>
</li>
<li><a href="#whats-next">What's next</a></li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Features for status pages</h3>
<h4>Retrospective incidents</h4>
<p>Ever internally declare an incident, have your team fix the issue, and in the heat of the moment completely forget to update your status page?</p>
<p>It happens more often than you'd think, which is why OnlineOrNot now supports adding incidents that happened in the past:</p>
<p><img src="/assets/onlineornot-updates-from-2024-june-july/status-pages-retroactive.png" alt="Retrospective incidents - OnlineOrNot Status Page"></p>
<h4>Descriptions and hiding from search engines</h4>
<p>You can now add descriptions (in markdown) to status pages, and hide them from being indexed by search engines.</p>
<p><img src="/assets/onlineornot-updates-from-2024-june-july/status-page-description.png" alt="Incident history - OnlineOrNot Status Page"></p>
<p>To update your status page's settings, click into your status page's dashboard and visit the "Settings" tab.</p>
<p>In the "Settings" tab, you'll find a new Description field, and a checkbox for disabling search engines from indexing your status page:</p>
<p><img src="/assets/onlineornot-updates-from-2024-june-july/status-page-search-engine-blocking.png" alt="Incident history - OnlineOrNot Status Page"></p>
<h4>Incident history</h4>
<p>Each status page now displays more than just the last fourteen days of incidents: you can now access your status page's entire incident history.</p>
<p><img src="/assets/onlineornot-updates-from-2024-june-july/incident-history.png" alt="Incident history - OnlineOrNot Status Page"></p>
<p>To access a status page's incident history, scroll down to "Recent Incidents" and click the "view incident history" link to the right.</p>
<h4>IP allowlisting</h4>
<p>You might want to restrict access to your status page, but not want to share a password around with your team.</p>
<p>For these cases, OnlineOrNot now supports IP allowlisting:</p>
<p><img src="/assets/onlineornot-updates-from-2024-june-july/ip-allowlist.png" alt="IP allowlisting - OnlineOrNot Status Page"></p>
<h4>Additional external status pages</h4>
<p>OnlineOrNot can now display the uptime status of the following providers on your status page:</p>
<ul>
<li>Ably</li>
<li>Radar</li>
</ul>
<p>Bringing it to a total of thirty-two external services supported, with more to come!</p>
<h2>What's next</h2>
<p>Here's what I have planned for this coming month:</p>
<ul>
<li>Ability to subscribe to status pages via webhook</li>
<li>Ability to create or update incident updates through APIs</li>
<li>Ability to push metrics and plot graphs to showcase for customers</li>
<li>Display historical uptime stats for the past year in a calendar</li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>So here's an ask</strong>:</p>
<p>I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from May 2024]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2024-may</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2024-may</guid>
            <pubDate>Tue, 04 Jun 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>In May I mostly took the time to plan out the rest of the year, as well as one big feature to make OnlineOrNot's incident management more realistic.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#fine-grained-management-for-your-incidents">Fine-grained management for your incidents</a></li>
<li><a href="#additional-external-status-pages">Additional external status pages</a></li>
</ul>
</li>
<li><a href="#whats-next">What's next</a></li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Fine-grained management for your incidents</h3>
<p>There's no sugar-coating it: incidents can get complicated. To help, OnlineOrNot now supports updating components as your incident occurs, for components that recover before others.</p>
<p>Here's a concrete example of how this works. Say your AI provider goes down, impacting your API:</p>
<p><img src="/assets/onlineornot-updates-from-2024-may/incident-start.png" alt="Incident Start - OnlineOrNot Status Page"></p>
<p>Half an hour passes, your AI provider comes back online, but your API has degraded performance as your users rush to use your service again:</p>
<p><img src="/assets/onlineornot-updates-from-2024-may/incident-update.png" alt="Incident Update - OnlineOrNot Status Page"></p>
<p>Finally, your API scales up to meet demand, and your service is operating as normal, so you resolve the incident:</p>
<p><img src="/assets/onlineornot-updates-from-2024-may/incident-end.png" alt="Incident End - OnlineOrNot Status Page"></p>
<p>In short, OnlineOrNot now supports letting your users know on a per-update level, how your incident is progressing, and what the status of each component is.</p>
<h3>Additional external status pages</h3>
<p>OnlineOrNot can now display the uptime status of the following providers on your status page:</p>
<ul>
<li>Paddle</li>
<li>GitLab</li>
<li>Bunny.net</li>
</ul>
<p>Bringing it to a total of thirty external services supported, with more to come!</p>
<h2>What's next</h2>
<p>This past month I took the time to evaluate the backlog, and see what customers and prospects have been asking for.</p>
<p>Here's the minimum of what I have planned for this coming month:</p>
<ul>
<li>Ability to restrict access to status pages via IP allowlisting</li>
<li>Ability to subscribe to status pages via webhook</li>
<li>Ability to create or update incident updates through APIs</li>
<li>Ability to push metrics and plot graphs to showcase for customers</li>
<li>Have historical details of the events + uptime percentage for the past year</li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>So here's an ask</strong>:</p>
<p>I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from April 2024]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2024-april</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2024-april</guid>
            <pubDate>Mon, 06 May 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>In April I kept focusing on making Status Pages better than ever, and added a new on-call integration: <a href="https://onlineornot.com/docs/incident-io-integration">incident.io</a>.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#the-free-checker-returns">The free checker returns</a></li>
<li><a href="#onlineornot-x-incidentio-integration">OnlineOrNot x Incident.io integration</a></li>
<li><a href="#status-page-component-groups">Status Page Component Groups</a></li>
<li><a href="#cron-expression-support">Cron expression support</a></li>
</ul>
</li>
<li><a href="#whats-next">What's next</a></li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>The free checker returns</h3>
<p>OnlineOrNot originally started back in 2021 as a tiny form at <code>https://onlineornot.com</code>, letting you enter a URL, and the system would check the URL from around the world.</p>
<p>As I added more features, I needed space on the home page to describe all the functionality OnlineOrNot supported, and eventually I removed the free checker.</p>
<p>It's back now, and better than ever at <a href="https://onlineornot.com/website-down-checker">Is this website down?</a></p>
<p><img src="/assets/onlineornot-updates-from-2024-april/free-checker.png" alt="Free Website Checker"></p>
<h3>OnlineOrNot x Incident.io integration</h3>
<p><a href="https://incident.io">Incident.io</a> lets you manage incidents entirely within your Slack, handle on-call, manage runbooks, and learn from your incidents with post-incident reviews.</p>
<p>If you've dreamed of being able to raise an incident because your daily database backup didn't run (frankly, I get it), you'll want to check this out:
<img src="/assets/onlineornot-updates-from-2024-april/incident-io-cron-job-alert.png" alt="Cron Job alerts in Incident.io"></p>
<p>With the integration, OnlineOrNot supports sending alerts from all uptime checks, as well as cron job/scheduled task checks to incident.io. Once in incident.io, the member of your team currently on-call can be notified, an incident raised, and a post-mortem written, all from the comfort of your own Slack.</p>
<h3>Status Page Component Groups</h3>
<p>Your status page can get a little busy once you've added your own system components next to components from your vendors and other third
parties.</p>
<p>Grouping your components lets you separate your system into logical parts, to make it easier to understand your status at a glance.</p>
<p>As a common example, let's say you've integrated deeply with OpenAI (and a few other vendors), and want to let your customers know when issues with their services may impact your own:</p>
<p><img src="/assets/onlineornot-updates-from-2024-april/ungrouped-status-page.png" alt="Ungrouped OnlineOrNot Status Page"></p>
<p>Even with a single third-party vendor, it starts to become difficult to see where your service ends, and theirs begins.</p>
<p>Grouping simplifies things:</p>
<p><img src="/assets/onlineornot-updates-from-2024-april/grouped-status-page.png" alt="Grouped OnlineOrNot Status Page"></p>
<h3>Cron expression support</h3>
<p>OnlineOrNot now supports monitoring even the most complicated cron schedules with <a href="/cron-job-monitoring">heartbeat checks</a>. Particular thanks to Martin for raising this feature request by replying to a monthly update email!</p>
<p>While getting an alert when your cron job fails to run "every day" is great, sometimes you have jobs you only want to run at midnight every weekday (<code>0 0 * * 1-5</code>):
<img src="/assets/onlineornot-updates-from-2024-april/cron-expression.png" alt="Cron Job alerting for OnlineOrNot Checks"></p>
<h2>What's next</h2>
<p>This section is always difficult to write, if I'm totally honest: software estimation is always a challenge.</p>
<p>Regardless, here's the minimum that I aim to release this month:</p>
<ul>
<li>Scheduled maintenance for status pages</li>
<li>Retrospective incidents for status pages</li>
<li>Additional external status pages</li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>So here's an ask</strong>:</p>
<p>I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from March 2024]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2024-march</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2024-march</guid>
            <pubDate>Tue, 02 Apr 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>March was a fun month for OnlineOrNot. I spoke to more customers than ever before, developed a long roadmap for OnlineOrNot's status page functionality (it's going to be an even better stand-alone product), heartbeat checks are going to support cron expressions, and more.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#up-to-date-changelog">Up-to-date changelog</a></li>
<li><a href="#external-status-pages">External Status Pages</a></li>
<li><a href="#manage-status-pages-through-an-api">Manage status pages through an API</a></li>
<li><a href="#reminder-alerts">Reminder alerts</a></li>
</ul>
</li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Up-to-date changelog</h3>
<p>For about a year, the only way to know what was new in OnlineOrNot was to follow me on <a href="https://twitter.com/OnlineOrNot">Twitter</a>. I've realized not everyone is on Twitter these days, so I'm now also publishing the near-daily updates on <a href="https://discord.onlineornot.com/">Discord</a>, and <a href="https://www.linkedin.com/company/onlineornot">LinkedIn</a>.</p>
<p>More substantial feature updates will be added to the <a href="https://onlineornot.com/changelog">changelog</a> (and there's now an in-app notification when that happens). Once a month updates will continue coming on <a href="https://onlineornot.com/articles">this blog</a> and its newsletter.</p>
<h3>External Status Pages</h3>
<p>Web applications are more complicated than ever. Your service depends on other services, and a simple status page displaying "Website" and "API" as components isn't enough any more.</p>
<p>OnlineOrNot now lets you display the live status of 25 other services, including Cloudflare, OpenAI, DigitalOcean, Vercel, and more.</p>
<p>To get started, jump into your status page's list of components, click the "add a component" dropdown, and click "add a third-party component":</p>
<p><img src="/assets/onlineornot-updates-from-2024-march/external-status-pages.png" alt="OnlineOrNot - External Status Pages"></p>
<p>This will be a growing list, and I expect to have hundreds of services supported by the end of the year.</p>
<h3>Manage status pages through an API</h3>
<p>Manually setting up a single status page can be fine, but what if you have dozens, or even hundreds of pages to manage?</p>
<p>To make this easier, OnlineOrNot now provides API endpoints for managing <a href="https://developers.onlineornot.com/api/status-pages">status pages</a>, <a href="https://developers.onlineornot.com/api/status-page-components">status page components</a>, and <a href="https://developers.onlineornot.com/api/status-page-incidents">status page incidents</a>.</p>
<p>I've also begun work on an official terraform provider for OnlineOrNot, soon you'll be able to manage status pages (and more) through an infrastructure-as-code setup.</p>
<h3>Reminder alerts</h3>
<p>You might get a notification from OnlineOrNot during a meeting, and not realize your website never came back online, or cron job never reported successful completion.</p>
<p>To remedy this, OnlineOrNot now sends reminder notifications until your website comes back online. You can tweak the frequency under a Check's alert settings:</p>
<p><img src="/assets/onlineornot-updates-from-2024-march/reminder-alerts.png" alt="OnlineOrNot - Reminder alerts"></p>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p><strong>So here's an ask</strong>:</p>
<p>I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, I'd like to know more!</p>
<p>You can reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a>, if you'd like to chat.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Improving your on-call schedule with runbooks]]></title>
            <link>https://onlineornot.com/improving-oncall-schedule-with-runbooks</link>
            <guid>https://onlineornot.com/improving-oncall-schedule-with-runbooks</guid>
            <pubDate>Tue, 12 Mar 2024 07:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Incidents are a stressful time for your team: your service isn't working the way you expect and your customers/stakeholders want to know what's going on. The last thing you want to do is let your team improvise everything when it comes to responding to incidents.</p>
<p>Google's own <a href="https://sre.google/sre-book/managing-incidents/">SRE book</a> has great overall tips for incident management, part of which involves "develop(ing) and document(ing) your incident management procedures in advance", which this article dives into.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-is-a-runbook">What is a runbook</a></li>
<li><a href="#your-first-runbook">Your first runbook</a>
<ul>
<li><a href="#when-to-write-your-first-runbook">When to write your first runbook</a></li>
<li><a href="#how-to-write-a-runbook">How to write a runbook</a></li>
<li><a href="#do-not-allow-perfectionism-to-block-you-from-releasing-your-runbooks">Do not allow perfectionism to block you from releasing your runbooks</a></li>
<li><a href="#use-new-team-members-as-a-starting-point">Use new team members as a starting point</a></li>
<li><a href="#do-not-worry-about-automation-yet">Do not worry about automation (yet)</a></li>
</ul>
</li>
<li><a href="#improving-your-existing-runbooks">Improving your existing runbooks</a>
<ul>
<li><a href="#optimize-for-editing">Optimize for editing</a></li>
<li><a href="#encourage-newcomers-to-use-your-runbooks-and-make-edits">Encourage newcomers to use your runbooks and make edits</a></li>
<li><a href="#keep-track-of-how-often-you-need-to-use-your-runbooks">Keep track of how often you need to use your runbooks</a></li>
<li><a href="#explicitly-tell-the-reader-what-to-do">Explicitly tell the reader what to do</a></li>
<li><a href="#make-your-runbook-easier-to-search">Make your runbook easier to search</a></li>
</ul>
</li>
</ul>
<h2>What is a runbook</h2>
<p>Depending on where you work, you'll hear different words describing relatively similar concepts around runbooks:</p>
<ul>
<li><strong>Standard Operating Procedure</strong> (or <strong>SOP</strong>): a standard set of steps required by the business to perform a process or task</li>
<li><strong>Runbook</strong>: a set of steps required to <strong>respond to an incident or alert</strong></li>
<li><strong>Playbook</strong>: what we call the overall set of steps when responding to an incident. Can involve several runbooks</li>
</ul>
<p>You shouldn't worry <em>too</em> much about terminology, just know we're going to be writing instructions on how to perform the manual process of resolving an incident with a system your team runs.</p>
<h2>Your first runbook</h2>
<h3>When to write your first runbook</h3>
<p>Ideally, you would write a runbook as you're deploying a service for the first time. You would go through the manual process you take to resolve any issues that come up, and write down each step, in detail, so anyone on your team can follow it while half-asleep.</p>
<p>Realistically though, chances are you'll be writing your first runbook after something goes wrong. For a lot of organisations, it takes an incident to really examine the way you work and learn where the potential improvements are.</p>
<p>Most commonly, something will go wrong, it'll take your on-call engineer a few minutes to figure out what happened, page the team members that know what to do, and it'll take the team a few hours to resolve the issue.</p>
<p>During the post-incident review, someone will ask "why did it take so long to fix?", and one of the inevitable answers will be "we didn't have a runbook in place". One of your team members will go off on a search for how to create a runbook, and they'll probably end up reading this article.</p>
<h3>How to write a runbook</h3>
<p>Honestly, <strong>anything</strong> is better than nothing. As time goes on, and your team gains operational experience, your runbooks will improve as gaps are found.</p>
<p>To start, you want to be documenting the manual steps necessary to get your service working again. You want to tell folks exactly what to do:</p>
<ol>
<li>Login to the Splunk dashboard at URL
<ul>
<li>Run this query: "QUERY GOES HERE"</li>
<li>In case of weird results, run this command: "COMMAND HERE"
<ul>
<li>Weird results look like this: "SCREENSHOT GOES HERE"</li>
</ul>
</li>
</ul>
</li>
</ol>
<p>In other words, favour explicit steps over implicit steps.</p>
<p>If any of your runbook steps say "Do the usual troubleshooting procedure", you need to rewrite that step, following the advice above.</p>
<p>If any step of your runbook contains "If you need to do X, follow the procedure here...", you need to explain the decision making process that leads the reader towards figuring out if they need to do X.</p>
<h3>Do not allow perfectionism to block you from releasing your runbooks</h3>
<p>Depending on your organisation, you may have stakeholders that'll want every aspect of the system documented in the runbook before it's "ready".</p>
<p>Without feedback on using the runbook (such as during real or simulated incidents), you'll quickly start getting diminishing returns on the time you spend trying to perfect your runbooks.</p>
<p>You're better off spending that time resolving issues that cause common alerts to fire for your on-call engineers.</p>
<h3>Use new team members as a starting point</h3>
<p>New team members are a gift for getting started, and improving your runbooks. They aren't yet affected by the <a href="https://en.wikipedia.org/wiki/Curse_of_knowledge">curse of knowledge</a>, and tend to be curious about how things work.</p>
<p>Have them sit in on incidents, join incident response chat rooms, and have them write down any questions they have as the incident is worked on. The questions they come up with are perfect starting points for a runbook.</p>
<h3>Do not worry about automation (yet)</h3>
<p>Automating your runbooks comes later (chances are, it won't cover every single step of your runbook anyway).</p>
<p>To start with though, you're going to want to document the <strong>entire</strong> manual procedure before you begin to automate the easiest steps.</p>
<p>In short: write today, automate tomorrow.</p>
<p>If you're early in your incident response journey, your team isn't ready for automation. Your team likely has too much <a href="https://en.wikipedia.org/wiki/Tribal_knowledge">tribal knowledge</a>, and needs to work on documenting their manual processes first.</p>
<h2>Improving your existing runbooks</h2>
<p>Now imagine it's 2am.</p>
<p>You're getting paged for the latest service your team built and deployed. You have no idea how to debug it, and it's the first time you've been paged for the service.</p>
<p>"No worries, I'll just go through the list of steps in the runbook, I'm sure it'll be fine... wait, runbook just says 'Ask Dave'?!"</p>
<p>You start paging the developers who built the service (especially Dave), and over a few anxious hours, you and the team manage to resolve the issue, and go back to bed.</p>
<p>In the rest of this article, we're going to improve our runbooks <strong>today</strong>, so no one in your team needs to experience the above scenario.</p>
<h3>Optimize for editing</h3>
<p>The <a href="/incident-management/incident-response/writing-your-first-runbooks">first runbook your team writes</a> is probably going to suck, but that's okay!</p>
<p>Getting started is the first step towards being good at something.</p>
<p>You're going to want to put your runbooks somewhere <strong>easy to edit</strong>. While a git repo is fine, something like Notion, Confluence or Google Docs is going to be easier to update.</p>
<p>Note that a self-hosted wiki/repo probably isn't the best idea - particularly if your team needs VPN/office network access to read the runbooks, and your VPN goes down.</p>
<h3>Encourage newcomers to use your runbooks and make edits</h3>
<p>As mentioned earlier, authors of your runbooks are affected by the curse of knowledge, and you will assume people will have the background knowledge to understand what to do.</p>
<p>Newcomers to your team aren't yet affected by this, and can help you uncover assumed knowledge in your runbooks. Have them pair with your on-call team members as they resolve incidents - they'll get more comfortable with on-call before their first roster, and will notice undocumented gaps that your team <em>just knows</em>.</p>
<p>This will also help the team feel comfortable with editing your runbooks - they're not perfect, and will need changing over time.</p>
<h3>Keep track of how often you need to use your runbooks</h3>
<p>As I mentioned earlier, while the initial goal is to document the manual process required to resolve an incident, a later goal is to automate those steps.</p>
<p>The more often you use a runbook, the more evidence you gather that perhaps parts of the runbook should be automated.</p>
<p>A simple table at the bottom of the runbook would help with this:</p>
<p>| Date last used | Used by |
| -------------- | ------- |
| 2022/01/01     | rozenmd |
| ...            | ...     |</p>
<p>Then, at a regular frequency (say, once a month or so), review your runbooks for automation opportunities.</p>
<h3>Explicitly tell the reader what to do</h3>
<p>People shouldn't need to interpret what the steps in your runbook <em>could</em> mean. Write as though the person reading it just woke up at 2am, and just wants to go back to bed.</p>
<p>The steps should <strong><a href="/incident-management/incident-response/guidelines-for-writing-better-runbooks#explicitly-tell-the-reader-what-to-do">explicitly tell them what to do</a></strong>. If you're starting with existing runbooks, it's worth auditing them to ensure the content is explicit, rather than implicit.</p>
<h3>Make your runbook easier to search</h3>
<p>It helps to add key phrases from your alerts, or even the exact error message thrown by your system.</p>
<p>For example:</p>
<pre><code class="language-md"># How to fix: "Error: Query defined in resolvers, but not in schema"

1. If you see this error in environment X, you need to run this command:
   ...
</code></pre>
<p>It'll make it easier to find a solution to the exact problem the system is having (by including the searched phrase in your heading, you'll also game certain <a href="https://docsearch.algolia.com/docs/tips/#structure-the-hierarchy-of-information">documentation search engines</a> to rank the terms higher).</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[How to check if a website is online]]></title>
            <link>https://onlineornot.com/how-check-website-online</link>
            <guid>https://onlineornot.com/how-check-website-online</guid>
            <pubDate>Thu, 07 Mar 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Ever wonder how to check if a website is still working, without having to load up the website and manually check every few minutes? In this article, we'll go over the various ways to check if a website is online.</p>
<p>If you just want a tool to check if a website is down, check out OnlineOrNot's website down checker.</p>
<p>I also have a video guide on <a href="https://www.youtube.com/watch?v=8eP451y6GNE">how to check if a website is online</a>.</p>
<p><strong>Table of contents:</strong></p>
<ul>
<li><a href="#why-do-websites-go-offline">Why do websites go offline?</a>
<ul>
<li><a href="#1-the-server-is-down">1. The server is down</a></li>
<li><a href="#2-the-server-takes-too-long-to-respond">2. The server takes too long to respond</a></li>
<li><a href="#3-the-domain-name-is-not-pointing-to-the-server">3. The domain name is not pointing to the server</a></li>
<li><a href="#4-the-domain-name-has-expired">4. The domain name has expired</a></li>
<li><a href="#5-the-ssl-certificate-has-expired">5. The SSL certificate has expired</a></li>
</ul>
</li>
<li><a href="#how-to-check-if-a-website-is-online">How to check if a website is online</a>
<ul>
<li><a href="#1-visit-the-website">1. Visit the website</a></li>
<li><a href="#2-use-a-command-line-tool">2. Use a command line tool</a></li>
<li><a href="#3-use-an-online-tool">3. Use an online tool</a></li>
<li><a href="#4-monitor-your-website">4. Monitor your website</a></li>
</ul>
</li>
<li><a href="#how-to-monitor-your-website">How to monitor your website</a></li>
<li><a href="#how-to-monitor-keywords-on-a-website">How to monitor keywords on a website</a></li>
</ul>
<h2>Why do websites go offline?</h2>
<p>In its simplest form, a website is a collection of files that are hosted on a server. When you visit a website, your browser makes a request to the server, and the server responds with the files for the website. More recently, websites have started to use APIs to fetch data to generate webpages from other servers, but the core concept is the same.</p>
<p>When we say a website is online, we mean that a server is responding (with the expected files) to requests from the Internet.</p>
<p>Things aren't that simple when a website is offline, though. There are a few different ways a website can be offline:</p>
<h3>1. The server is down</h3>
<p>If the server is down (that is, not running), it can't respond to requests from the Internet, so the website will be unavailable.</p>
<h3>2. The server takes too long to respond</h3>
<p>If the server takes too long to respond, the browser will give up and show an error. This often happens when a server is overloaded with requests, or if the server makes a request to another server, which itself is overloaded.</p>
<p>Smaller companies tend to run into this issue by hosting their websites on small servers without a CDN in front, as the servers often don't have enough RAM/CPU to handle the number of requests they receive.</p>
<p>In my experience, this is the most common way a website can be offline.</p>
<h3>3. The domain name is not pointing to the server</h3>
<p>If the domain name is not pointing to the server, the server won't receive any requests from the Internet.</p>
<p>You might run into this issue when launching your website, as in the worst case, it can take a few hours for the domain name to start pointing to the server. This can also happen when changing domain name registrars.</p>
<h3>4. The domain name has expired</h3>
<p>If the domain name has expired, the domain name registrar will stop pointing the domain name to the server. This is an avoidable problem however, and I wrote about <a href="/guidelines-to-help-avoid-losing-your-domain">how to avoid losing your domain name</a> a while back.</p>
<h3>5. The SSL certificate has expired</h3>
<p>This tends to happen a lot less these days with the advent of automatically renewing SSL certificates from <a href="https://letsencrypt.org/">LetsEncrypt</a> and free SSL certificates from the likes of Cloudflare, but SSL certificates do have an expiry date.</p>
<p>If your server isn't configured to automatically renew its SSL certificate, modern browsers often won't even try to display your website, and users are trained to avoid sites that don't use HTTPS.</p>
<h2>How to check if a website is online</h2>
<p>There are a few ways to check if a website is online, some requiring more manual work than others:</p>
<h3>1. Visit the website</h3>
<p>The most obvious way to check if a website is online is to visit it using your browser. If you can see the website, it's online. If you can't, it's probably offline (assuming your internet connection is fine).</p>
<p>This method isn't bulletproof, as a website appearing online for you doesn't guarantee that it's online for everyone else.</p>
<h3>2. Use a command line tool</h3>
<p>If you're technically-minded, you can use a command line tool like <code>curl</code> in your Terminal to check if a website is online.</p>
<pre><code class="language-bash">curl https://www.google.com
</code></pre>
<p>If you see a bunch of HTML, the website is online. If you see an error, the website is offline.</p>
<p>If you want to take it further, you can turn the above command into a script that runs every few minutes, and sends you an email if the website is offline:</p>
<pre><code class="language-bash">#!/bin/bash

if curl -s https://www.google.com | grep -q "html"; then
  echo "Website is online"
else
  echo "Website is offline" | &#x3C;command to send email goes here>
fi
</code></pre>
<h3>3. Use an online tool</h3>
<p>OnlineOrNot's website down checker is a free tool to check if a website is online.</p>
<p>All you need to do is enter a URL, such as <code>https://www.google.com</code>, and click "Check".</p>
<p>OnlineOrNot will then check the website from the US East Coast, US West Coast, Australia, Europe, and Asia Pacific regions to see if the website is online, and how long it takes to respond:</p>
<p><img src="assets/how-check-website-online/onlineornot-website-down-checker.png" alt="OnlineOrNot website down checker"></p>
<h3>4. Monitor your website</h3>
<p>If you're running a business, you probably don't have time to check if your website is online every few minutes. You also don't want to have customers email/call/text you frantically that your website is down, and they're unable to use your product.</p>
<p>This is where website uptime monitoring services like <a href="https://onlineornot.com/">OnlineOrNot</a> come in. It checks if your website is online up to every 30 seconds, and if it's not, you'll immediately get a notification via email, SMS, Slack, or Discord.</p>
<h2>How to monitor your website</h2>
<p>To monitor a website with OnlineOrNot, all you need is the URL.</p>
<p>I tend to give my checks memorable names, so I can easily tell what they're for, and I set the check to run from around the world, so I can be sure the website is online for everyone:</p>
<p><img src="assets/how-check-website-online/landing-page-settings-onlineornot.png" alt="OnlineOrNot landing page monitoring settings"></p>
<p>After adding the check, you will start to see the results in the dashboard:</p>
<p><img src="assets/how-check-website-online/onlineornot-dashboard.png" alt="OnlineOrNot dashboard"></p>
<h2>How to monitor keywords on a website</h2>
<p>Monitoring for keywords on OnlineOrNot requires one extra step.</p>
<p>In a check's advanced settings I set the 'Text to search for' to look for the keyword I'm interested in (it also supports HTML, if you want to check for a specific HTML element):</p>
<p><img src="assets/how-check-website-online/landing-page-advanced-settings-onlineornot.png" alt="OnlineOrNot landing page monitoring settings"></p>
<p>Once setup, OnlineOrNot will send you an alert whenever that keyword is <strong>not</strong> found on the website.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from January 2024]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-2024-january</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-2024-january</guid>
            <pubDate>Sat, 03 Feb 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Over December and January I focused on OnlineOrNot's public API, making it possible to see what OnlineOrNot saw when it detected a failing uptime check, as well as fixing bugs and cleaning up tech debt across all of OnlineOrNot (it's getting faster and faster to release updates).</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#adopting-openapi-and-publishing-new-api-docs">Adopting OpenAPI and publishing new API docs</a></li>
<li><a href="#onlineornot-what-did-you-see">OnlineOrNot, what did you <em>see</em>?</a></li>
</ul>
</li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Adopting OpenAPI and publishing new API docs</h3>
<p>OnlineOrNot's public API was originally written to serve a single purpose: provide data for the <a href="https://github.com/OnlineOrNot/onlineornot">CLI</a>. The documentation for the API was handwritten, and lovingly placed alongside OnlineOrNot's <a href="https://onlineornot.com/docs/welcome">other hand-written docs</a>.</p>
<p>This made it time-consuming to update the API: first I would write code, then remember to context-switch into "content-mode" and decide where the new API endpoint docs should live, then actually write docs from scratch, and finally deploy the content.</p>
<p>In December, I automated this process (I wrote more about the technical details <a href="/built-my-http-docs-from-scratch">here</a>), and now we now have a <a href="https://github.com/OnlineOrNot/api-schemas">published OpenAPI schema</a> and a <a href="https://developers.onlineornot.com/">new API documentation hub</a>:</p>
<p><img src="/assets/onlineornot-updates-from-2024-january/api-docs.png" alt="New API docs"></p>
<p>The API itself (like OnlineOrNot) is still under active development, and new endpoints will be automatically published to both the documentation site, and the OpenAPI schema in the coming weeks.</p>
<h3>OnlineOrNot, what did you <em>see</em>?</h3>
<p>Up until now, OnlineOrNot would only ever tell you if your website or API was up or down, and roughly how long your website or API took to serve a request. Actually figuring out what was wrong was "<em>left as an exercise for the reader</em>", like an academic textbook.</p>
<p>As a user of OnlineOrNot, you would have one of two reactions to this:</p>
<ol>
<li>
<blockquote>
<p>I better investigate my logs around that time, see if the database backup job isn't accidentally killing our website at the same time.</p>
</blockquote>
</li>
<li>
<blockquote>
<p>Nah, that's a false alarm, <strong>no way</strong> the website goes down at exactly 2am every single day.</p>
</blockquote>
</li>
</ol>
<p>To remedy this, in January I focused on making every request result (HTTP status code, region it checked from, response time, date and time checked) available in the dashboard, as well as the request/response metadata (request headers, response headers and response body).</p>
<p>Every uptime check in OnlineOrNot now displays a list of recent checks:</p>
<p><img src="/assets/onlineornot-updates-from-2024-january/recent-uptime-results.png" alt="Recent uptime checks"></p>
<p>with an option to view all checks, and filter by failing checks:</p>
<p><img src="/assets/onlineornot-updates-from-2024-january/all-uptime-results.png" alt="All uptime results"></p>
<p>and finally, the raw metadata involved in the check:</p>
<p><img src="/assets/onlineornot-updates-from-2024-january/uptime-metadata.png" alt="uptime result metadata"></p>
<p>By the time you see these screens in OnlineOrNot, they will likely be even more featured - it was just so useful that I wanted to release and share an early preview as soon as possible.</p>
<p>In short, performing a root-cause analysis of why an uptime check failed with OnlineOrNot is now significantly easier.</p>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you.</p>
<p>I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you enjoying it enough to tell your friends and colleagues about it.</p>
<p>So here's an ask: I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, reply to this email/email me at <a href="mailto:max@onlineornot.com">max@onlineornot.com</a> - I'd like to know more!</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[I built my HTTP API docs from scratch]]></title>
            <link>https://onlineornot.com/built-my-http-docs-from-scratch</link>
            <guid>https://onlineornot.com/built-my-http-docs-from-scratch</guid>
            <pubDate>Sun, 14 Jan 2024 10:30:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>You might be thinking “building HTTP API docs from scratch? in 2024? wtf?”, and you’re probably right. After all redoc has been around since 2016, and there are hundreds of “generate <em>beautiful documentation</em> from your OpenAPI spec” startups around, some even use AI now.</p>
<p>To be honest, I didn’t even know it was possible to do-it-yourself when I started looking into it. Every search turned up with instructions on how to convert my OpenAPI schema into <em>beautiful documentation</em> via …some third party.</p>
<p>The thing is, I don’t want AI to answer questions about my product. If I can’t guarantee that it will be 100% accurate, it is useless to me. Similarly, I don’t want to put a third party in charge of displaying my API’s documentation to my users.</p>
<p>Frankly, third parties have no skin in the game when generating <em>beautiful documentation</em>. They seem to be more worried about selling to you, the one with the API, than to your users. Who cares if the docs are unusable? They’re beautiful! Have you seen our dark mode? /s</p>
<p>As a developer that started out on the frontend before learning backend and infra/ops, I thought I could do better, so I gave it a shot.</p>
<p>It only took 8 hours in the end, and I'm pleased with <a href="https://developers.onlineornot.com/">the results</a>.</p>
<p>There were two steps:</p>
<ul>
<li><a href="#converting-openapi-schemas-to-markdown">Converting OpenAPI schemas to markdown</a>
<ul>
<li><a href="#introducing-widdershins">Introducing widdershins</a>
<ul>
<li><a href="#customizing-widdershins-output">Customizing widdershins output</a></li>
</ul>
</li>
</ul>
</li>
<li><a href="#converting-markdown-to-html">Converting markdown to HTML</a></li>
</ul>
<h2>Converting OpenAPI schemas to markdown</h2>
<p>I cheated a little bit, I didn’t handwrite an OpenAPI YAML schema from scratch (though I have in the past). My API of choice (<a href="https://hono.dev/">HonoJS</a>) provides a middleware for turning <a href="https://hono.dev/snippets/zod-openapi">Zod validation into OpenAPI</a> schemas, so I used that.</p>
<p>With an OpenAPI spec in hand, I needed a means of converting that data into markdown. Extensive searching came up with a weird word: widdershins.</p>
<h3>Introducing widdershins</h3>
<p><a href="https://github.com/Mermade/widdershins">Widdershins</a> was originally written to convert Swagger/OpenAPI specs to Slate/ReSlate compatible markdown using a templating language called <a href="https://github.com/olado/doT#readme">doT</a>.</p>
<p>I had no interest in using Slate/ReSlate, but since a templating language was involved, I knew I could grab the existing templates, fork them, and have them output <a href="https://mdxjs.com/">MDX</a>-compatible code instead.</p>
<h4>Customizing widdershins output</h4>
<p>Customizing the widdershins templates was a matter of cloning the widdershins repo, copying the files in <a href="https://github.com/Mermade/widdershins/tree/main/templates/openapi3">templates/openapi3</a> into my own directory, and telling widdershins to use my templates instead of the default ones.</p>
<p>widdershins didn't do <em>everything</em> I wanted, but thankfully it was easy to customize the output.</p>
<p><strong>Some hacks involved</strong></p>
<ul>
<li>By default, widdershins outputs a single markdown file, so I had to insert placeholder text in my templates, and later use that to split the markdown by resource
<ul>
<li>Using an OpenAPI resource tag in my template, I then write the desired filename (<code>checks/index.mdx</code> for example) as the first line of the split markdown, and use that as my filename</li>
</ul>
</li>
<li>doT uses curly brackets for variables with no simple way of inserting curly brackets into output, so I had to use <code>&#x26;lcub;&#x26;lcub;</code> and <code>&#x26;rcub;&#x26;rcub;</code> and run a <code>markdownOutput.replace()</code> to fix up my MDX</li>
</ul>
<h2>Converting markdown to HTML</h2>
<p>As a developer, you're spoiled for choice these days when it comes to turning markdown/MDX into HTML. I was already a paying customer of <a href="https://tailwindui.com/">Tailwind UI</a>, so I chose to use their Protocol template (which uses Next.js under the hood), and modify it to my needs (mainly changes to the color-scheme and icon sizing).</p>
<p>I used Cloudflare's <a href="https://github.com/cloudflare/next-on-pages">next-on-pages</a> project so I could host the website on <a href="https://pages.cloudflare.com/">Cloudflare Pages</a>, and not worry about being billed for bandwidth.</p>
<p>** Summary **</p>
<p>In short, it took me 8 hours to go from OpenAPI YAML file to my own custom HTTP API documentation website, and I'm very pleased with <a href="https://developers.onlineornot.com/">the results</a>.</p>
<p>My main aim with this project was being able to have a say over every single aspect of the docs, so that I can react to my own customer feedback. If folks write in asking for extra details in certain places, or that something doesn't make sense, I'm capable of modifying the docs from the OpenAPI spec to the pipeline generating markdown, to the markdown renderer. I don't have to write to a vendor asking them to fix something, and have them tell me it'll be scheduled some time next quarter, maybe.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[How to get your first ten customers]]></title>
            <link>https://onlineornot.com/how-to-get-your-first-ten-customers</link>
            <guid>https://onlineornot.com/how-to-get-your-first-ten-customers</guid>
            <pubDate>Mon, 18 Dec 2023 06:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>It'll soon be the third anniversary of publicly <a href="https://twitter.com/RozenMD/status/1364881512500404224">launching OnlineOrNot on Twitter</a>, and I often get asked what I did to get my first paying customers - so I felt like sharing.</p>
<p>I assume when most folks ask this that they're looking for the <em>one thing</em> they can do to finally start getting paid customers.</p>
<p>Let me be clear: it's never just <em>one thing</em>.</p>
<p>Update: I wrote about my third year of running OnlineOrNot <a href="https://maxrozen.com/lessons-from-my-third-year-running-a-saas">here</a>.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#what-i-actually-did">What I actually did</a></li>
<li><a href="#some-context">Some context</a></li>
<li><a href="#i-read-a-lot">I read, a lot</a></li>
<li><a href="#on-charging-low-and-having-a-generous-free-tier">On charging low and having a generous free tier</a></li>
<li><a href="#rethinking-the-free-tier">Rethinking the free tier</a></li>
<li><a href="#battling-churn">Battling churn</a></li>
<li><a href="#just-keep-shipping">Just keep shipping</a></li>
</ul>
<h2>What I actually did</h2>
<p>Off the top of my head:</p>
<ul>
<li>I tweeted every time I was thinking about a feature, building a feature, and released a feature</li>
<li>I tweeted random things I learned about running a business, as I learned them</li>
<li>I tapped my personal network on LinkedIn to find out what folks were currently using, and how they were dissatisfied</li>
<li>I charged way too low and attracted un-ideal customers</li>
<li>I made a free tier to compete with other providers that also charged way too low</li>
<li>I wrote changelogs for every big feature</li>
<li>I wrote docs for every feature</li>
<li>I kept a mailing list where I would update folks about the product's progress</li>
<li>I kept a mailing list where I would update folks about running the SaaS</li>
<li>I talked about OnlineOrNot on Reddit, Hacker News, in the aim of inspiring people to launch their idea in a highly competitive space</li>
<li>I shared OnlineOrNot on popular lists (Product Hunt, etc)</li>
<li>I made landing pages that resonate with folks looking specifically for my type of product</li>
<li>I built features customers asked for</li>
<li>I would spend <a href="https://codingweekmarketingweek.com/">one week marketing, followed by one week coding</a>, and repeat</li>
<li>I was extremely responsive to customers via email and chat</li>
<li>I kept improving the product, adding features, and revisiting features</li>
</ul>
<p>There's probably a lot more, but that's what comes to mind first.</p>
<h2>Some context</h2>
<p>OnlineOrNot <a href="https://onlineornot.com/">(this website)</a> started because I needed a weekly report for my contracting clients to prove their web host was unreliable, to the point where it was costing them significant money. They were paying for the cheapest possible tier of WordPress hosting at the time, and didn't believe me when I said random 5 minute blocks of downtime throughout the day were adding up.</p>
<p>I built a dirt-simple form that takes a URL and sends an email notification when the site goes down/up, with a weekly summary email.</p>
<p>Then, I kept adding features every day, 2 hours at a time, even after I stopped being a contractor.</p>
<h2>I read, a lot</h2>
<p>When starting off, I did a <em>lot</em> of reading to figure out what others did before me. While many folks tend to hold-off from starting a business until they feel like they've read everything, I started the business first and started reading as I needed help.</p>
<p>The following articles/books were particularly influential:</p>
<ul>
<li>
<p><a href="https://www.momtestbook.com/">The Mom Test by Rob Fitzpatrick</a></p>
</li>
<li>
<p><a href="https://deployempathy.com/">Deploy Empathy by Michele Hansen</a></p>
</li>
<li>
<p><a href="https://www.amazon.com/Traction-Startup-Achieve-Explosive-Customer/dp/1591848369">Traction by Gabriel Weinberg and Justin Mares</a></p>
</li>
<li>
<p><a href="https://stripe.com/guides/atlas/starting-sales">Your first 10 customers by Patio11</a></p>
</li>
<li>
<p><a href="https://www.bannerbear.com/journey-to-10k-mrr/">A Bootstrapped SaaS Journey to $10K MRR by Jon Yongfook</a></p>
</li>
<li>
<p>A random indiehackers forum comment that made me totally re-think marketing:</p>
<ul>
<li>
<blockquote>
<p>Don't believe for one moment that 5 hours on Product Hunt or anywhere else for that matter represents a serious marketing effort.</p>
</blockquote>
<blockquote>
<p>If you want to run a business rather just create stuff, your work has only just begun. In the light of Facebook and other social media revelations, the idea of a truly disposable email address which means your entire life is not analysed and spammed to death has to be worth something.</p>
</blockquote>
<blockquote>
<p>You haven't told anyone about it though. And I mean you shout from the rooftops every day and everywhere you can think of. You market. People are not going to come looking for you. You have to start approaching influencers, be seen and be heard everywhere you think your potential users might lurk.</p>
</blockquote>
<blockquote>
<p>And, by the way, everyone sees a million new ideas a day so you have to be consistent, appear to be permanent and appear to be solid. No-one is going to entrust communications with you if they think you are a small, one-man band with an idea and little else.</p>
</blockquote>
<blockquote>
<p>Time to start reading marketing articles and strategies and applying them.</p>
</blockquote>
<blockquote>
<p>And expect it to take time.</p>
</blockquote>
</li>
</ul>
</li>
</ul>
<h2>On charging low and having a generous free tier</h2>
<p>Charging low <em>did help</em> attract my first customers, but few of them are still subscribed years later, as I did not build OnlineOrNot with them in mind.</p>
<p>Of particular note is the type of customer that would overload their single server with hundreds of websites, and complain that OnlineOrNot did what they hired it to, whenever their server would inevitably crash under load.</p>
<p>Customers that only use your service because it is cheap are also the type of customer to cheap out in other parts of their business, making them unpleasant to deal with. At $9/mo, you can't afford to spend much time on support - and these are the customers who will ask for the most support.</p>
<p>I also initially made a free tier to attempt to compete with other players in the market, offering dozens of uptime checks for free. Again, while this attracted tons of free users, it was expensive with little return to the business, so I eventually had to trim down the free tier to a level that the business <em>could</em> support.</p>
<h2>Rethinking the free tier</h2>
<p>For a very long time, it wasn't possible to even trial OnlineOrNot's paid plans. You would start on the free tier, and if you needed more checks, you would upgrade. This made growth extremely slow and painful as a founder.</p>
<p>I saw this as starting the relationship with a customer like this: "Hey, this is a free tool, if you use it a bit more, you can pay me". I needed to flip the relationship to be more like "Hey, this is a paid tool, if you want slower checks for your personal projects, I can help you out".</p>
<p>It wasn't until I flipped the relationship - starting you off on a free trial of OnlineOrNot's best features - that OnlineOrNot started growing significantly faster.</p>
<h2>Battling churn</h2>
<p>When I think about OnlineOrNot's customers, there are five "journeys" to spend my time building features for:</p>
<ul>
<li>Lead nurture (aka turning random folks browsing the site into free trials)</li>
<li>Trial conversion (turning free trials into paid customers)</li>
<li>Trial abandonment (what happens to the folks that don't subscribe)</li>
<li>Customer success (ensuring the first month of use is without drama)</li>
<li>Customer retention (communicating value over the long term, and keeping the product useful)</li>
</ul>
<p>When launching a product, most folks (myself included) tend to only focus on lead nurture and trial conversion. This makes sense when you don't have customers yet, I'll admit.</p>
<p>It wasn't until I started balancing my feature development between these five points, rather than just the first two that I got churn down to an acceptable level.</p>
<h2>Just keep shipping</h2>
<p>One of the beautiful things about bootstrapping a business while working a full-time job is that it's default-alive.</p>
<p>There is near-zero cash burn.</p>
<p>The only thing stopping folks from reaching their first ten paying customers is persistence. Just keep shipping (features, marketing), and eventually you'll get there.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from November 2023]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-november-2023</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-november-2023</guid>
            <pubDate>Fri, 01 Dec 2023 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Over October and November I focused on adding new features to Status Pages and making it easier to iterate on new features, as well as fixing bugs across all of OnlineOrNot.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#new-graphs-for-status-pages">New graphs for Status Pages</a></li>
<li><a href="#dark-mode-for-status-pages">Dark Mode for Status Pages</a></li>
<li><a href="#documentation-as-a-youtube-channel">Documentation as a YouTube channel</a></li>
</ul>
</li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>New graphs for Status Pages</h3>
<p>While compact, OnlineOrNot's old Status Page graphs were often difficult to read, and tried to pack too much information into a small space.</p>
<p>We now have new graphs:</p>
<p><img src="/assets/onlineornot-updates-from-november-2023/status-pages-graphs.png" alt="New Graphs for Status Pages"></p>
<p>The new graphs are:</p>
<ul>
<li>interactive, with tooltips displaying each data point</li>
<li>visually larger, with y-axis grid lines for easier reading</li>
</ul>
<p>You can add metrics graphs to your OnlineOrNot Status Pages by adding uptime checks to them, as described in <a href="/docs/onlineornot-status-page-integration">this support article</a>.</p>
<h3>Dark Mode for Status Pages</h3>
<p>A common ask from customers (and my customers' customers) is to make OnlineOrNot Status Pages less blinding when viewed on a system that has a dark-mode theme enabled.</p>
<p>So we now support dark mode:</p>
<p><img src="/assets/onlineornot-updates-from-november-2023/status-page-dark-mode.png" alt="Dark Mode for Status Pages"></p>
<p>By default, OnlineOrNot Status Pages will use the user's system theme to decide whether to default to dark mode or not. Dark mode itself can be toggled by clicking the moon/sun icon at the bottom of each status page.</p>
<h3>Documentation as a YouTube channel</h3>
<p>While OnlineOrNot has a fairly comprehensive set of <a href="/docs/welcome">documentation</a>, not everyone learns best by reading. To remedy this, I started a dedicated <a href="https://www.youtube.com/@OnlineOrNot">YouTube Channel for OnlineOrNot</a>, and a <a href="/screencasts">Screencasts page</a>, where folks can visually learn how to configure OnlineOrNot for their use case:</p>
<p><img src="/assets/onlineornot-updates-from-november-2023/screencasts-page.png" alt="Screencasts page"></p>
<p>Want a feature explained in a video? Let me know!</p>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you, enjoying it enough to tell their friends about it.</p>
<p>So here's an ask: I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, reply to this email/email me at <code>max@onlineornot.com</code> - I'd like to know more!</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from September]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-september-2023</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-september-2023</guid>
            <pubDate>Mon, 02 Oct 2023 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Coming back from August holidays, I felt the need to take a hard look at what OnlineOrNot does, and keep improving it.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#whats-new">What's new</a>
<ul>
<li><a href="#cron-job-monitoring">Cron job monitoring</a></li>
<li><a href="#on-call-integrations">On-call integrations</a></li>
<li><a href="#automated-free-tier">Automated free tier</a></li>
</ul>
</li>
<li><a href="#thank-you">Thank you</a></li>
</ul>
<h2>What's new</h2>
<h3>Cron job monitoring</h3>
<p>I mainly spent September refining <a href="/cron-job-monitoring">cron job monitoring</a> functionality - the system handles millions of cron job checks per day, and is now production-ready.</p>
<p>Cron job monitoring flips how OnlineOrNot works in reverse: instead of checking your website to see if it's online, with cron job monitoring you tell OnlineOrNot if your cron job/scheduled task/IoT device/internet connection is still online via HTTP request, and if the time between requests is too long, you'll get an alert.</p>
<p>After a bit of feedback (particular thanks to Jens!) from OnlineOrNot's customers, quite a few things got built and fixed for cron job monitoring:</p>
<ul>
<li>You can now display the status of your cron job on our <a href="/status-pages">status pages</a>, and have OnlineOrNot automatically open and close incidents based on your cron job monitoring</li>
<li>You can now send alerts about your cron jobs to the following places:
<ul>
<li>SMS</li>
<li>Email</li>
<li>Discord/Slack</li>
<li>Webhooks</li>
<li>On-call integrations: PagerDuty/Opsgenie/Grafana OnCall/Spike.sh</li>
</ul>
</li>
<li>Cron job monitoring now lives on its own dedicated infrastructure, redundantly across several AWS regions</li>
<li>Fixed several bugs in the heartbeat monitoring graphs</li>
</ul>
<h3>On-call integrations</h3>
<p>On-call integrations aren't just for cron job monitoring by the way!</p>
<p>Back in late July I added support for uptime checks to alert based on your team's existing on-call roster for:</p>
<ul>
<li><a href="/docs/pagerduty-integration">PagerDuty</a></li>
<li><a href="/docs/opsgenie-integration">Opsgenie</a></li>
<li><a href="/docs/grafana-oncall-integration">Grafana OnCall</a></li>
<li><a href="/docs/spike-integration">Spike.sh</a></li>
</ul>
<h3>Automated free tier</h3>
<p>OnlineOrNot has offered a free trial of its best features at the start for every new account for well over a year now, but what happens at the end of that free trial hasn't been particularly clear.</p>
<p>I used to think I was being clever by sending an email letting folks know they can pick between <a href="/pricing">upgrading to a paid plan</a>, or using the hobby plan for non-commercial projects. If they did nothing, OnlineOrNot would stop monitoring. In reality, many folks would miss that email, and not realize that OnlineOrNot stopped monitoring.</p>
<p>This goes against what I built OnlineOrNot for, so in September I automatically converted thousands of accounts with expired free trials to use the free tier. Going forward, folks that don't decide will automatically have their accounts converted to the free tier.</p>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you, enjoying it enough to tell their friends about it.</p>
<p>So here's an ask: I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, reply to this email/email me at <code>max@onlineornot.com</code> - I'd like to know more!</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[How to monitor your uptime with OnlineOrNot]]></title>
            <link>https://onlineornot.com/uptime-monitoring-best-practices</link>
            <guid>https://onlineornot.com/uptime-monitoring-best-practices</guid>
            <pubDate>Wed, 27 Sep 2023 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Jumping into monitoring software for the first time can be pretty overwhelming. If you're not in an exploring mood it can be easy to get lost, and you're not sure what all the knobs and buttons do.</p>
<p>To help lighten this feeling for OnlineOrNot, I thought it might be useful to let folks know how I use OnlineOrNot, to monitor OnlineOrNot.</p>
<p>You might think it's silly to monitor your own site as an uptime monitoring service, however as I keep the monitoring infrastructure separate from the marketing website and web app, I actually get notified if anything goes wrong.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-to-monitor">What to monitor</a></li>
<li><a href="#settings-for-landing-pages">Settings for Landing Pages</a></li>
<li><a href="#settings-for-apis">Settings for APIs</a></li>
<li><a href="#alert-settings">Alert settings</a></li>
</ul>
<h2>What to monitor</h2>
<p>To start with, I monitor:</p>
<ul>
<li>the main landing page of my marketing site <a href="https://onlineornot.com/">https://onlineornot.com/</a></li>
<li>the URL for our marketing site sitemap: <a href="https://onlineornot.com/sitemap.xml">https://onlineornot.com/sitemap.xml</a></li>
<li>the main public API endpoint <a href="https://api.onlineornot.com/v1/checks/">https://api.onlineornot.com/v1/checks/</a></li>
</ul>
<p>As OnlineOrNot's marketing site is mainly static HTML (not powered by anything server-based like WordPress), monitoring the main landing page covers almost every page that could go down. I also monitor the sitemap as it's generated by a script at build time, and that script has failed in the past.</p>
<h2>Settings for Landing Pages</h2>
<p>I use the following settings for both my main landing page, and the sitemap.xml file.</p>
<p>To start with, I have OnlineOrNot check its own landing page every thirty seconds, from around the world: <img src="assets/uptime-monitoring-best-practices/landing-page-settings-onlineornot.png" alt="OnlineOrNot landing page monitoring settings"></p>
<p>Things can get noisy on the internet, and it's possible for a website to fail uptime checks for a minute or two without it being a particular drama (assuming it's not a regular occurrence). As a result, I only want to be notified if my landing page check fails four times in a row (two minutes of continuous downtime):</p>
<p><img src="assets/uptime-monitoring-best-practices/landing-page-advanced-settings-onlineornot.png" alt="OnlineOrNot landing page monitoring advanced settings"></p>
<p>To be sure it's actually the page I expect that OnlineOrNot is checking, I also set the 'Text to search for' to look for text on the webpage.</p>
<h2>Settings for APIs</h2>
<p>For APIs, things are a little bit different. If the API check fails, I <strong>know</strong> something is wrong, and needs investigating immediately.</p>
<p>I have OnlineOrNot check its own public API every thirty seconds: <img src="assets/uptime-monitoring-best-practices/api-settings-onlineornot.png" alt="OnlineOrNot API monitoring settings"></p>
<p>To make OnlineOrNot actually check the API correctly, I have OnlineOrNot make an API request as a real (test) user, with a valid JWT token.</p>
<p>I set the following HTTP request settings:</p>
<p><img src="assets/uptime-monitoring-best-practices/api-http-request-settings-onlineornot.png" alt="OnlineOrNot API HTTP Request settings"></p>
<p>It's important to test for correctness too, so I set assertions to check the data is coming back correctly:</p>
<p><img src="assets/uptime-monitoring-best-practices/api-assertion-settings-onlineornot.png" alt="OnlineOrNot API assertion settings"></p>
<p>Finally, in advanced settings, I set the check to only alert me if two checks in a row fail.</p>
<p>As I'm already checking the response via Assertions, I don't set 'Text to search for' for my API check.</p>
<p><img src="assets/uptime-monitoring-best-practices/api-advanced-settings-onlineornot.png" alt="OnlineOrNot API monitoring advanced settings"></p>
<h2>Alert settings</h2>
<p>As I check my phone (way too much), I find email notifications to work quite well when things go wrong (no additional settings required).</p>
<p>For added redundancy though, I also have alerts sent to Slack and Discord, which I've added as integrations for my account:</p>
<p><img src="assets/uptime-monitoring-best-practices/alert-settings-onlineornot.png" alt="OnlineOrNot alert settings"></p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Scaling AWS Lambda and Postgres to thousands of simultaneous uptime checks]]></title>
            <link>https://onlineornot.com/scaling-aws-lambda-postgres-to-thousands-of-uptime-checks</link>
            <guid>https://onlineornot.com/scaling-aws-lambda-postgres-to-thousands-of-uptime-checks</guid>
            <pubDate>Sat, 16 Sep 2023 09:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>When you're building a serverless web app, it can be pretty easy to forget about the database. You build a backend, send some data to a frontend, write some tests, and it'll scale to infinity with no effort, right?</p>
<p>Not quite.</p>
<p>Especially not with a tiny Postgres server. As the number of users of your frontend increases, your app will open more and more database connections until the database is unable to accept any more.</p>
<p>That's just the frontend - it gets worse on the backend.</p>
<p>If you're using AWS Lambda functions as asynchronous "job runner" functions that increase in usage as your web app becomes more popular, you're going to run into scaling issues fast.</p>
<p>By revisiting my app's architecture, I removed the need for each worker function to talk to my database, and greatly improved scalability without needing to pay for a bigger database.</p>
<p>Curious? Read on.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#first-some-context">First, some context</a></li>
<li><a href="#the-scalability-issue">The scalability issue</a></li>
<li><a href="#fixing-the-issue">Fixing the issue</a></li>
</ul>
<h2>First, some context</h2>
<p>OnlineOrNot (where this article is hosted) started its life as an extremely simple uptime monitoring service. As a user you could sign in, add a URL you wanted to monitor, get alerts (only via email), and pay for a subscription.</p>
<p><img src="/assets/scaling-aws-lambda-postgres-to-thousands-of-uptime-checks/onlineornot-first-users.jpeg" alt="OnlineOrNot&#x27;s first form for adding URLs to monitor"></p>
<p>Architecture-wise, it was pretty simple too.</p>
<p>It consisted of:</p>
<ul>
<li>a single AWS RDS database instance running Postgres</li>
<li>a Lambda function that fired every minute to check the database for URLs to check</li>
<li>an SQS queue</li>
<li>a Lambda function that got triggered via SQS, which would check the URLs, and write the results to the database</li>
</ul>
<h2>The scalability issue</h2>
<p>If you have a bit of experience with building on serverless, you can probably immediately see where my mistake was. As the web app grew in popularity, each URL that got added for monitoring would need its own database connection when writing results.</p>
<p>That might have been fine if I was developing an internal uptime checker at a small company (assuming you build for the problems you <em>know</em> you have now, rather than problems you <em>might</em> have), but I'm running OnlineOrNot as a growing bootstrapped business.</p>
<p>Shortly after thinking to myself "this is fine", I found myself needing to check up to 100 URLs at a time, every minute. Each time the jobs would finish, all 100 Checker Lambda functions would try to write to the database at (more or less) the same time, my database would refuse to open any more connections, including from my frontend!</p>
<p>To buy myself time, I scaled up the database a few tiers (the more memory your RDS Postgres instance has, the more simultaneous connections it can receive), and went looking for ways to fix the issue.</p>
<h2>Fixing the issue</h2>
<p>While implementing a fix for the issue, I managed to land my largest customer to that point. They wanted to monitor over a thousand URLs at a time, every minute.</p>
<p>I ran the numbers - if I kept upgrading database tiers every time I needed to scale, I would quickly run out of database tiers to upgrade to.</p>
<p>The fix in the end was to:</p>
<ul>
<li>Instead of letting each checker Lambda write to the database, I had them return their result via Simple Notification Service (SNS)</li>
<li>Added a second queue, that batches results from SNS</li>
<li>Made a new AWS Lambda function to receive batched results, and write to the database from there</li>
</ul>
<p>The architecture ended up looking like this:</p>
<p>As a result, the application now scales relative to how many resultfwdr functions are running, rather than the individual checker Lambda functions. This pattern is also known as fan-out/fan-in architecture.</p>
<p>With this fix in place, I was able to downgrade my database instance back to where I started, with only a minor percentage increase in CPU usage.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Guidelines for updating WordPress and its plugins, safely.]]></title>
            <link>https://onlineornot.com/guidelines-for-updating-wordpress-plugins</link>
            <guid>https://onlineornot.com/guidelines-for-updating-wordpress-plugins</guid>
            <pubDate>Sat, 16 Sep 2023 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Chances are, you already know how important it is to keep WordPress and its plugins up to date. If not, let this article be a wake-up call: you <strong>absolutely</strong> need to keep your system up to date.</p>
<p>All software has bugs and vulnerabilities (whether it's WordPress itself, themes/plugins, or Apache, Nginx, Linux, PHP) discovered every day. Updates patch those bugs to keep your website secure and functioning.</p>
<p>That's all well and good, but the problem is: how do you keep your WordPress installation updated, without an update taking your website completely offline without you noticing? It can be pretty challenging - you have a website that works for your business, you don't need to login to update the content more than a couple of times a year - the last thing on your mind is updating your WordPress plugins.</p>
<p>This article can help.</p>
<p><strong>Table of contents:</strong></p>
<ul>
<li><a href="#manually-update-your-wordpress-plugins">Manually update your WordPress plugins</a>
<ul>
<li><a href="#recovering-a-failed-automatic-plugin-update">Recovering a failed automatic plugin update</a></li>
</ul>
</li>
<li><a href="#monitor-your-website-effectively">Monitor your website, effectively</a></li>
</ul>
<h2>Manually update your WordPress plugins</h2>
<p>While automatic updates are best-practice in most other places, such as your phone and laptop, you definitely want to manually update each WordPress plugin one-by-one, on a regular basis (whether that's each week, or each month).</p>
<p>Why manually? So that you can observe the results, and hit "Rollback" in case your website stops showing content after you update the plugin.</p>
<h3>Recovering a failed automatic plugin update</h3>
<p>"What if it's too late?!" you ask?</p>
<p>You can recover from an automatic update taking down your WordPress website in a few steps, assuming you can log in to the server</p>
<ol>
<li>Log in to your WordPress host server</li>
<li>Rename the existing <code>wp-content/plugins</code> folder, I'd call it something like <code>plugins_temp</code></li>
<li>Create a new, empty <code>plugins</code> folder</li>
<li>One plugin at a time: copy your plugins back from the temporary folder, into the <code>plugins</code> folder, and refresh your WordPress website in your browser - repeat until the site breaks</li>
<li>Once the website breaks, you know that's your bad plugin (or one of them)</li>
<li>At that point, you can log into the WordPress admin, and either rollback the plugin, or if it's a premium plugin, download a fresh copy and install it</li>
</ol>
<h2>Monitor your website, effectively</h2>
<p>While the author of this article <strong>does</strong> run a <a href="https://onlineornot.com/">website monitoring service</a>, this tip applies regardless of which uptime monitoring tool you use: it's important to monitor <em>correctly</em>.</p>
<p>The last thing you want is to find out your website has been down for months, while your uptime monitoring tool has been happily reporting your site as "up".</p>
<p>A simple "is my website online?" check will <strong>not</strong> always work for WordPress. Those famous "Critical Error" screens (below) can show up without sending a "down" HTTP status (4xx or 5xx).</p>
<p><img src="/assets/guidelines-for-updating-wordpress-plugins/critical-error-wordpress.png" alt="WordPress Critical Error Screen"></p>
<p>For this reason, I recommend uptime checks with "<a href="https://onlineornot.com/docs/search-for-text-uptime-checks">text to search for</a>" configured. These checks will look for text that only shows up when your content has loaded, while also looking for "down" HTTP status codes.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Domain expiry: How to prevent your domain from expiring]]></title>
            <link>https://onlineornot.com/guidelines-to-help-avoid-losing-your-domain</link>
            <guid>https://onlineornot.com/guidelines-to-help-avoid-losing-your-domain</guid>
            <pubDate>Tue, 25 Jul 2023 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Imagine you're sitting in your office, and you start noticing emails coming in asking if you'd like to buy your domain.</p>
<p>"Huh, that's weird, I already own that domain" you think to yourself.</p>
<p>A few more emails come in, and they're getting past the spam filter, so you decide to double check your domain manager. Doubt starts creeping into your mind, you start <em>panicking</em>, and you frantically scroll down to where the domain should be, and...</p>
<p>It's gone.</p>
<p>The only option you have is to pay the person that grabbed your domain $3000 USD.</p>
<p><em>Hold up, rewind...</em></p>
<p>This sort of scenario can be avoided, yet an entire industry of domain squatters exists due to how commonly it occurs.</p>
<p>In this article, I'll provide advice you can do <strong>today</strong> to keep your domain secure in the long run.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#enable-2fa">Enable 2FA</a></li>
<li><a href="#check-that-your-domain-is-set-to-auto-renew">Check that your domain is set to auto renew</a></li>
<li><a href="#lock-your-domain-from-transfer">Lock your domain from transfer</a></li>
<li><a href="#check-your-payment-details">Check your payment details</a></li>
<li><a href="#use-a-reputable-domain-registrar">Use a reputable domain registrar</a></li>
<li><a href="#be-sure-you-actually-own-your-domain">Be sure you actually own your domain</a></li>
<li><a href="#extend-your-domain-registration">Extend your domain registration</a></li>
<li><a href="#be-aware-of-any-tld-specific-rules-around-renewals">Be aware of any TLD-specific rules around renewals</a></li>
</ul>
<h2>Enable 2FA</h2>
<p>If your domain registrar supports it, enable 2FA (two-factor authentication, also known as MFA/multi-factor authentication). It'll send you an email/SMS/push notification when logging into your domain manager.</p>
<p>While not a foolproof way of stopping hackers (hackers could still phish your employees for their 2FA code - YubiKey devices don't have this issue), it'll slow them down <strong>and</strong> alert you if your account has been compromised.</p>
<h2>Check that your domain is set to auto renew</h2>
<p>Some domain registrars don't enable auto renew by default, particularly when transferring domains.</p>
<p>Check your domain registrar to see that it's enabled. You can usually find it in settings when managing your domain.</p>
<h2>Lock your domain from transfer</h2>
<p>While you're checking your domain manager has auto renewal enabled, also double check that your domain's "transfer lock" is also enabled.</p>
<p>Enabling transfer lock for your domain is effectively like a car alarm for your domain. If someone manages to get into your domain manager account, and tries to transfer the domain, you'll receive quite a few emails about it.</p>
<p>Enabling transfer lock might slow you down in the future when you want to change domain registrars, but if you don't plan on moving any time soon, you may as well turn it on.</p>
<h2>Check your payment details</h2>
<p>I know this one sounds obvious, but if the payment fails, your domain doesn't get renewed, and you typically have about 30 days to notice before your domain disappears.</p>
<p>The most common mistake is that your credit card expires, and you forget to update the payment details your domain registrar has on file.</p>
<p>On the off-chance your domain registrar accepts PayPal (or similar), also double check that the payment details that PayPal have are also up to date.</p>
<h2>Use a reputable domain registrar</h2>
<p>There are a few ways to interpret "reputable" - I mean large companies trust them with their services, and the business itself is trustworthy. Certain domain registrars also own companies that <a href="https://en.wikipedia.org/wiki/Domain_drop_catching">"drop catch"</a> domains that expire from their services. Would <strong>you</strong> want to use a domain registrar that's financially incentivized to let your domain expire?</p>
<p>Here are some reputable domain registrars that immediately come to mind:</p>
<ul>
<li><a href="https://aws.amazon.com/route53/">AWS Route 53</a>
<ul>
<li>Some of the largest internet companies trust AWS to host their services</li>
</ul>
</li>
<li><a href="https://www.cloudflare.com/en-au/products/registrar/">Cloudflare Registrar</a>
<ul>
<li>20% of the internet's traffic runs through them, and the domain registrar service is provided at-cost</li>
</ul>
</li>
<li>NameCheap
<ul>
<li>Decent reputation, has been around a very long time, used to make you pay for privacy, now doesn't</li>
</ul>
</li>
<li>Gandi
<ul>
<li>Decent reputation, has been around a very long time</li>
</ul>
</li>
</ul>
<p>I've personally had negative experiences with GoDaddy and CrazyDomains (in Australia), and would strongly recommend to anyone reading this: transfer your domain ASAP to somewhere else.</p>
<h2>Be sure you actually own your domain</h2>
<p>If you purchased your domain through a third-party, like Wix, WordPress, or maybe the agency or contractor that helped build your site, chances are you're not fully in control of your domain.</p>
<p>Sure, it might be easier for you to have a third-party manage the domain for you, and pass the bill along each year, however this sort of arrangement can become problematic when you want to cancel the service, or move to another provider.</p>
<p>If you're having an agency or contractors build your site for you, and they become unresponsive, you risk losing the domain if you also let them manage it for you. By keeping the login to your domain manager to yourself, you can cut ties with rogue third-parties and move to a different provider.</p>
<h2>Extend your domain registration</h2>
<p>Most domain registrars will let you extend your registration for around $12 USD per year for .com domains. This lets you remove the risk of your automatic renewal not going through by manually renewing.</p>
<p>For example, AWS offers the following:</p>
<p><img src="/assets/guidelines-to-help-avoid-losing-your-domain/extend-domain-registration.png" alt="AWS Extend Domain Registration"></p>
<p>Considering the amount of money domain squatters will try to get from you if you let your domain expire, it's a pretty good deal.</p>
<h2>Be aware of any TLD-specific rules around renewals</h2>
<p>While it's great fun to grab a domain from a country half way across the world from you so you can spell out your brand, different countries have different rules around domain renewals.</p>
<p>As an example, Spain (.es) charges a renewal fee on top of an annual fee. As well as that, if you let the domain expire, there's another renewal fee that ranges from 30 USD to hundreds of dollars (depending on your registrar).</p>
<h2>Know when something goes wrong</h2>
<p>Even with all these precautions, things can still go wrong. Payment failures, registrar issues, DNS misconfigurations - any of these can take your site offline.</p>
<p>Set up <a href="/uptime-monitoring">uptime monitoring</a> so you'll know within minutes if your site goes down. Whether it's an expired domain, a DNS issue, or something else entirely, you want to find out before your customers do.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Our lessons from the latest AWS us-east-1 outage]]></title>
            <link>https://onlineornot.com/our-lessons-from-the-latest-us-east-1-outage</link>
            <guid>https://onlineornot.com/our-lessons-from-the-latest-us-east-1-outage</guid>
            <pubDate>Sun, 18 Jun 2023 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>In case you missed it, AWS experienced an outage or "elevated error rates" on their AWS Lambda APIs in the us-east-1 region between 18:52 UTC and 20:15 UTC on June 13, 2023.</p>
<p>If this sounds familiar, it's because it's almost a replay of what happened on <a href="/onlineornot-aws-outage-retrospective">December 7, 2021</a>, although that outage was significantly more severe and took longer to restore.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#some-background-on-us-east-1">Some background on us-east-1</a></li>
<li><a href="#the-easiest-mitigation-strategy">The easiest mitigation strategy</a></li>
<li><a href="#the-retrospective">The retrospective</a>
<ul>
<li><a href="#some-context">Some context</a></li>
<li><a href="#what-went-wrong">What went wrong</a></li>
<li><a href="#what-went-well">What went well</a></li>
<li><a href="#what-i-learned">What I learned</a></li>
</ul>
</li>
</ul>
<h2>Some background on us-east-1</h2>
<p>us-east-1 is the oldest and most used AWS region, it's where <a href="https://aws.amazon.com/ec2/">AWS EC2</a> was launched from in 2006. It has six availability zones (AZs), instead of the usual two or three.</p>
<p>Bugs that show up under high load tend to show up there first. If you look at the <a href="https://en.wikipedia.org/wiki/Timeline_of_Amazon_Web_Services#Amazon_Web_Services_outages">list of AWS outages</a>, you'll notice almost every major outage since 2017 started in us-east-1:</p>
<p><img src="/assets/our-lessons-from-the-latest-us-east-1-outage/us-east-1-outages.png" alt="a list of aws outages"></p>
<h2>The easiest mitigation strategy</h2>
<p>While folks may say us-east-1 is "just another region", you could avoid being impacted by major outages with one simple strategy:</p>
<p>Migrate to us-east-2.</p>
<p>It's not 100% foolproof, as some AWS services host their dashboard exclusively in us-east-1, so even if your service isn't impacted by an outage, you might not be able to check out the dashboard during an outage.</p>
<p>Unfortunately for OnlineOrNot, as an uptime monitoring service, things aren't that simple.</p>
<h2>The retrospective</h2>
<h3>Some context</h3>
<p>Since OnlineOrNot's job is to measure uptime, every service is designed to run across multiple AWS regions. There's a high availability API that OnlineOrNot's services use to figure out which region to run from, and I use that to dodge major outages (something I learned to build back in 2021).</p>
<h3>What went wrong</h3>
<p>A couple of weeks before the most recent AWS outage as part of an architecture refactor, I spun up a new alerting service in us-east-1, without taking the time to implement the same multi-region failover that OnlineOrNot's other systems use.</p>
<p>As a result, during the us-east-1 outage while uptime checks were successfully being scheduled, and running - the service that collects the results from individual checks was not.</p>
<h3>What went well</h3>
<p>The only redeeming part of this incident, was that it proved the failover system works.</p>
<p>After the outage, I took the time to add failover to the alerting service, and now the entire OnlineOrNot system has the ability to hop AWS regions.</p>
<h3>What I learned</h3>
<ol>
<li>Don't run production workloads on services that don't have failover implemented</li>
<li>Regularly test service failover, end-to-end</li>
<li>It's almost never a good idea to build things that run <em>only</em> in us-east-1</li>
</ol>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[No, the average cost of downtime is not $5600 per minute]]></title>
            <link>https://onlineornot.com/average-cost-of-downtime-not-5600-minute</link>
            <guid>https://onlineornot.com/average-cost-of-downtime-not-5600-minute</guid>
            <pubDate>Wed, 14 Jun 2023 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>A fairly common claim among website uptime monitoring services is that downtime costs $5600 per minute. Chances are, you'll have one of two reactions to this claim:</p>
<ol>
<li>"Argh, that's a lot! We should use uptime monitoring!"</li>
<li>"That's completely made up, no one running a business my size would tell you a number anywhere near this high"</li>
</ol>
<p>The reality of what downtime costs your business lies somewhere in between.</p>
<p>As a company that runs 3.6 million uptime checks per week, we have a bit of insight into the cost of downtime, so if you're curious - read on.</p>
<h2>Why it's not $5600 per minute</h2>
<p>Let's start with the obvious calculation: take your annual revenue, divide it by the number of minutes in a year, and there's your cost per minute of downtime.</p>
<p>The formula for calculating your downtime cost per minute:</p>
<pre><code class="language-js">(annual revenue) / 60 * 24 * 365
</code></pre>
<p>This assumes 100% of your revenue comes from your website and your website being offline drops your revenue to 0, which are fairly strong assumptions to make.</p>
<p>For downtime to cost $5600 per minute by this measure, your business would need to be making $2.9 billion per year. That's almost enough to make it into the Fortune 500.</p>
<h2>How much does downtime <em>really</em> cost your business</h2>
<p>Let's say your business makes $1 million per year, and 100% of revenue comes from your website and web app, and for each minute your website is down, you're unable to generate revenue:</p>
<pre><code class="language-js">1000000 / 60 * 24 * 365 = 1.902
</code></pre>
<p>At this level of revenue, downtime costs you $1.90 per minute. If your website is offline for a couple of hours without you noticing, that's a couple hundred dollars in lost revenue.</p>
<h2>Why this calculation is not enough</h2>
<p>Unfortunately, things aren't this simple in the real world. Downtime also costs your business reputational damage from both prospective customers unable to pay you, and existing customers unable to use your service. For e-commerce businesses, the above calculation is a rough estimate at best.</p>
<p>For Software as a Service (SaaS) businesses, the above calculation doesn't make sense, as SaaS businesses often run their marketing websites separate to the systems that their customers pay them for (and will often have significantly more advanced monitoring to ensure everything is still running). So their marketing websites going offline might be a little embarrassing, but not the business-ending event some folks make it out to be.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[OnlineOrNot updates from May]]></title>
            <link>https://onlineornot.com/onlineornot-updates-from-may</link>
            <guid>https://onlineornot.com/onlineornot-updates-from-may</guid>
            <pubDate>Thu, 08 Jun 2023 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>It's a warm afternoon here in Toulouse, I feel like it's as good a time as any to update folks about what's new in OnlineOrNot.</p>
<h2>What's new</h2>
<p>In May I spent time making OnlineOrNot better for folks monitoring as a team.</p>
<p>The headline feature is that I've updated how Slack integrations work with OnlineOrNot: you can now disable Slack alerts for an uptime check, or even pick which channel alerts should go to. Folks on the Team plan and above can add multiple Slack channels, and send alerts directly to individual people in your Slack.</p>
<p>Not everyone uses Slack of course, there were other updates:</p>
<ul>
<li>Admins of OnlineOrNot teams can now remove people from a team</li>
<li>You can now monitor entire HTML tags with the "Text to search for" setting in Uptime Checks</li>
<li>OnlineOrNot has been accepted as a "known bot" in Google Analytics 4 (GA4). This means once GA4 updates (I'm told this will happen by June 23), uptime checks won't appear in your GA4 analytics anymore</li>
<li><a href="/cron-job-monitoring">Cron Job Monitoring</a> is now available to try out. The feature is still being refined in beta, but now anyone can try them out.</li>
</ul>
<h2>What's next</h2>
<p>Almost all of my product ideas come from a very simple process:</p>
<ul>
<li>When folks sign up, I ask them a single question about what brought them to OnlineOrNot</li>
<li>Some folks will respond with some of the pains they've experienced while monitoring their websites using other tools</li>
<li>I'll propose a solution, and those folks tend to stick around once it's built</li>
</ul>
<p>June will bring a few things:</p>
<ul>
<li>Making it possible to fetch historical uptime data</li>
<li>Making it possible to add multiple Discord integrations, and pick which alerts get sent to which Discord channel</li>
<li>Rethinking how OnlineOrNot handles alerts and incidents: at the moment you get one email when things go wrong - sometimes you want to keep getting reminded until you confirm that you've seen the alert</li>
<li>Rethinking how OnlineOrNot displays who gets alerts: as I add additional integrations, the "Alert Settings" section for Uptime Checks gets harder to understand</li>
</ul>
<h2>Thank you</h2>
<p>OnlineOrNot only exists because of folks like you. I don't have a massive marketing budget, OnlineOrNot gets its users from folks like you, enjoying it enough to tell their friends about it.</p>
<p>So here's an ask: I work on the things that people care enough about to tell me about them. If you think there's something missing, or something that needs improving, reply to this email/email me at <code>max@onlineornot.com</code> - I'd like to know more!</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[On writing better error messages]]></title>
            <link>https://onlineornot.com/write-better-error-messages</link>
            <guid>https://onlineornot.com/write-better-error-messages</guid>
            <pubDate>Mon, 27 Mar 2023 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>You're browsing your favorite website, clicking around, when suddenly, you're rudely interrupted by a white screen, proclaiming:</p>
<blockquote>
<p>Error 503 Service Unavailable</p>
</blockquote>
<p><img src="/assets/write-better-error-messages/awkward-error-message.png" alt="Unhelpful Error Message">
(I don't mean to pick on Varnish cache here, It's just a screenshot I had handy)</p>
<p>As a developer, my eyes scan error messages like these for numbers - in this case, the "503" - indicating that the error isn't my fault, and I can move on with my life.</p>
<p>Unfortunately the majority of internet users aren't trained in the art of reading <a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Status">HTTP status codes</a>, so this error message isn't particularly useful to them.</p>
<h2>How to write better error messages</h2>
<p>The majority of internet users aren't developers, so outputting the HTTP error code and its name (503 Service Unavailable) isn't good enough.</p>
<p>The <a href="https://www.nngroup.com/articles/improving-dreaded-404-error-message/">Nielsen Norman Group</a> back in 1998 (!) provided us with some basic guiding principles for writing better error messages:</p>
<ul>
<li>Write in plain English (or whichever language you're supporting)</li>
<li>Tell the user exactly what went wrong</li>
<li>Tell the user how the problem can be fixed</li>
</ul>
<p>More concretely, we can write better error messages by answering the following four questions:</p>
<ol>
<li><strong>Who</strong> caused the error?</li>
<li><strong>What</strong> happened, and <strong>why</strong>?</li>
<li><strong>When</strong> will it be fixed?</li>
<li><strong>How</strong> can the user respond to the error?</li>
</ol>
<p>If your error message covers those four points, <em>then</em> you can think about adding humour and some brand identity.</p>
<p>Note that humour and brand identity should never interfere with telling the user how to resolve the issue as briefly as possible:</p>
<p><img src="/assets/write-better-error-messages/getting-mocked.jpg" alt="Bad error messages like &#x22;Oopsie woopsie&#x22; are not recommended"></p>
<h3>Who caused the error?</h3>
<p>The last thing you want to do is make your users feel dumb, or as though they're at fault for an issue with the service. Communicating who caused the error helps clear up any confusion.</p>
<p>Explicitly focus on "we" when the error is caused by an issue on your end (typically HTTP status codes in the 5xx range).</p>
<p>An error message such as</p>
<blockquote>
<p>Our service is down for maintenance</p>
</blockquote>
<p>is infinitely better than:</p>
<blockquote>
<p>Uh-oh! Something went wrong!</p>
</blockquote>
<p>For errors caused by the user (typically HTTP status codes in the 4xx range), be explicit about that too. For example, a 403 Forbidden error (where you know the user isn't authorized to view content) could be communicated as:</p>
<blockquote>
<p>Access Denied. You do not have permission to view this page.</p>
</blockquote>
<h3>What happened, and why?</h3>
<p>While users may not be technical, they still need an explanation of why they're seeing your error screen.</p>
<p>Take the classic 404 error message: "404 Not Found". You can make it significantly better for non-technical users with a single word:</p>
<blockquote>
<p>Page not found.</p>
</blockquote>
<p>adding a "why", makes it even better, giving them a way to fix the issue:</p>
<blockquote>
<p>Page not found. You might have mistyped the URL.</p>
</blockquote>
<h3>When will it be fixed?</h3>
<p>It's relatively difficult to keep an error message updated with details of your outage, and when you expect the service to become available again.</p>
<p>A better approach would be to link to either your <a href="/status-pages">status page</a>, or Twitter account, or both, as in GitHub's case:</p>
<p><img src="/assets/what-fastly-outage-can-teach-about-writing-error-messages/github-down-example.png" alt="An example of GitHub&#x27;s service error page"></p>
<h3>How can the user respond to the error?</h3>
<p>In the case of a 404, you might want to list some steps the user can take to fix the issue, such as:</p>
<ul>
<li>Going back to the home page</li>
<li>Using your search bar to find the page if it's been moved</li>
<li>Contacting support</li>
</ul>
<p>Whereas in the case of a 5xx server error, you want to communicate to the user that there isn't much they can do, and that it's not their fault.</p>
<p>My favorite example of a company doing this well is Airbnb:</p>
<p><img src="/assets/what-fastly-outage-can-teach-about-writing-error-messages/airbnb-down-example.png" alt="An example of Airbnb&#x27;s service error page"></p>
<p>They tell users:</p>
<ul>
<li>there's definitely an issue, and they're working on it,</li>
<li>to check out their Twitter account for updates,</li>
<li>a way to get support for urgent issues,</li>
<li>and they set the expectation that they may be slow to respond while the site is experiencing downtime</li>
</ul>
<h2>Summary</h2>
<p>We as developers need to improve our error messages. Try to be helpful, and explain:</p>
<ol>
<li><strong>Who</strong> caused the error?</li>
<li><strong>What</strong> happened, and <strong>why</strong>?</li>
<li><strong>When</strong> will it be fixed?</li>
<li><strong>How</strong> can the user respond to the error?</li>
</ol>
<p>Depending on the type of service you run, you'll likely still need request IDs and other diagnostics in your error message to help your support staff debug the issue.</p>
<p>A human-readable error message doesn't have to come at the expense of removing <em>all</em> technical information.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Monitoring our monitoring]]></title>
            <link>https://onlineornot.com/monitoring-our-monitoring</link>
            <guid>https://onlineornot.com/monitoring-our-monitoring</guid>
            <pubDate>Wed, 08 Mar 2023 08:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Last Saturday, <a href="https://status.onlineornot.com/incidents/wRYnAV2o9WMG">our API went down</a>. Not even a funny error message or slightly slower responses either, it just completely vanished off the internet for 18 minutes.</p>
<p>I'm not normally one to point fingers at my hosting provider when things go wrong (since ultimately, <em>I chose</em> to use them, so it's <em>my problem</em> to fix), but when <a href="https://community.fly.io/t/reliability-its-not-great/11253">fly.io publicly posts on their forums about their reliability issues</a>, I may as well link to them.</p>
<p>That being said, my fix was "simple": our monitoring is multi-region and multi-cloud, our API can be too.</p>
<p>Since I use Cloudflare Workers in front of everything, I started using a Worker as a sort of load-balancer to send traffic to different providers, in the hope that it'll be more resilient in case of an individual host suffering another outage.</p>
<h2>Suddenly, false-positives</h2>
<p>Around the same time, I noticed the number of email alerts OnlineOrNot was sending was a bit high.</p>
<p>On a bad day, we'll send out maybe 300 emails. On Saturday, we sent out 4900.</p>
<p>This Saturday morning, I rolled out a change to OnlineOrNot to migrate from a system that checks from a single region by default (us-east), to one that checks around the world (us-east -> eu-west -> south-east asia -> us-west) significantly more frequently (up to every 30 seconds for paid plans).</p>
<p>As part of this change, a bad deployment to Singapore resulted in elevated error alerts for almost all of our uptime checks from Singapore. At first I tried finding another region that wasn't having reliability issues, before just rolling back the change entirely.</p>
<h2>What I did about it</h2>
<p>This weekend made me realize I didn't have a reliable way of knowing when OnlineOrNot's own monitoring was having issues. While <a href="https://onlineornot.com/uptime-monitoring-best-practices">OnlineOrNot does use OnlineOrNot to monitor OnlineOrNot</a> quite successfully, it only covers its web presence (marketing website, web app, API), and not the services running the system.</p>
<p>Instead, I would rely on a live-tail of the logs to see how things were going (essentially, vibes-based monitoring), and for two years that worked quite well! I also have a sort of "deadman's switch" to automatically change hosting provider running the uptime checks in the case of total outage, but that didn't help here.</p>
<p>To start with, I've started storing the results of each check in Clickhouse and built a Grafana dashboard for a live view of the system:</p>
<p><img src="/assets/monitoring-our-monitoring/grafana.png" alt="OnlineOrNot&#x27;s Grafana dashboard"></p>
<p>While it doesn't sound like much, this sudden visibility into how the system is running (live) gives me the confidence to make improvements I wouldn't otherwise be able to make, such as quadrupling uptime check concurrency on a single VM, and tweaking default timeouts and retries to reduce false positives.</p>
<p>We also have distributed tracing in other parts of the system (status pages), which I'll be adding to the uptime check system to get a bit more of a detailed view into how each check runs (more than just <code>console.log</code>, at least).</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Announcing OnlineOrNot's open-source uptime check CLI]]></title>
            <link>https://onlineornot.com/announcing-onlineornot-uptime-check-cli</link>
            <guid>https://onlineornot.com/announcing-onlineornot-uptime-check-cli</guid>
            <pubDate>Thu, 02 Mar 2023 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>I'm excited to be launching <a href="https://github.com/onlineornot/onlineornot">onlineornot</a> - OnlineOrNot's newly-built, open source (TypeScript) CLI for managing your uptime checks, all from the comfort of your terminal!</p>
<p>Getting started is a matter of:</p>
<ul>
<li><a href="https://onlineornot.com/docs/cli-installation">Installing onlineornot</a> via npm</li>
<li>Running <a href="https://onlineornot.com/docs/cli-login"><code>onlineornot login</code></a></li>
<li>and finally, <code>onlineornot checks list</code> to get a list of your uptime checks</li>
</ul>
<p>You can also check out the <a href="https://onlineornot.com/docs/cli-commands">docs</a> to see what else you can do.</p>
<p>"Why a CLI?" you might ask - OnlineOrNot needed a way of testing the newly built <a href="https://onlineornot.com/docs/api-overview">public API</a>.</p>
<p>To start with, the public API supports fetching data about <a href="https://onlineornot.com/docs/api-uptime-checks">uptime checks</a> and <a href="https://onlineornot.com/docs/api-tokens">API tokens</a>, but I'll be supporting status pages soon too.</p>
<p>If you don't feel like using the API, you can always pass a <code>--json</code> flag to the CLI to get your output as JSON, and pass it to the next command you'd like to run.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[What I learned running a SaaS for a second year]]></title>
            <link>https://onlineornot.com/lessons-from-two-years-of-saas-operation</link>
            <guid>https://onlineornot.com/lessons-from-two-years-of-saas-operation</guid>
            <pubDate>Mon, 20 Feb 2023 06:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Two years ago, OnlineOrNot started as a little toy app I built in an afternoon to see what it's like using the Next.js framework, to see if a URL is down from around the world.</p>
<p>I gave myself a week to turn that toy into a SaaS people could pay for. It looked like this when it went live:</p>
<p><img src="/assets/lessons-from-two-years-of-saas-operation/one-week-of-onlineornot.jpeg" alt="OnlineOrNot after one week"></p>
<p>It wasn't ready for real users, but that didn't matter. I had something out there, that people could sign-up for, tell me what they were expecting, and how OnlineOrNot fell short of their expectations.</p>
<p>Since then, I've been shipping features in two hour blocks, cutting down scope aggressively to ensure something goes out into the world, each time I write code for OnlineOrNot. I then get feedback on what goes out, and the app improves.</p>
<p>These days, OnlineOrNot can do a <em>quite a bit more</em> than just visit a page, and send an email alert when it's down:</p>
<p><img src="/assets/lessons-from-two-years-of-saas-operation/onlineornot-present-day.png" alt="OnlineOrNot, present day"></p>
<p>OnlineOrNot Uptime Checks get used to monitor web apps, APIs, internet connections, residential power, IoT devices, blogs, and of course, regular websites. It's also no longer just an uptime monitor. OnlineOrNot is more of a status page service, with built-in uptime monitoring.</p>
<p>I learned quite a bit in my <a href="/what-learned-running-saas-for-year">first year of running a SaaS</a>, and this year's learning builds on top of that.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#onlineornot-runs-on-opportunity-cost">OnlineOrNot runs on opportunity cost</a></li>
<li><a href="#no-running-two-saas-businesses-isnt-a-good-idea">No, running two SaaS businesses isn't a good idea</a></li>
<li><a href="#ship-small-invalidate-your-assumptions">Ship small, invalidate your assumptions</a></li>
<li><a href="#talk-to-potential-customers-even-if-you-think-you-have-nothing-for-them">Talk to potential customers, even if you think you have nothing for them</a></li>
<li><a href="#what-i-actually-shipped">What I actually shipped</a>
<ul>
<li><a href="#january">January</a></li>
<li><a href="#february">February:</a></li>
<li><a href="#march">March</a></li>
<li><a href="#april">April</a></li>
<li><a href="#may">May</a></li>
<li><a href="#june">June</a></li>
<li><a href="#july">July</a></li>
<li><a href="#august">August</a></li>
<li><a href="#september">September</a></li>
<li><a href="#october">October</a></li>
<li><a href="#november">November</a></li>
<li><a href="#december">December</a></li>
</ul>
</li>
</ul>
<h2>OnlineOrNot runs on opportunity cost</h2>
<p>I don't spend anything on customer acquisition apart from opportunity cost. I build features and write content for OnlineOrNot's customers at the expense of possibly other lucrative activities I could be doing.</p>
<p>Folks often look at my "2 hours per work day, every work day" rule, calculate roughly what they think my hourly rate is, and say "pffft I'd rather be doing nothing".</p>
<p>The thing is, I'd rather be doing this. To me, OnlineOrNot is like painting, and if years down the track I decide to start a new artwork, I don't have to learn to paint again.</p>
<h2>No, running two SaaS businesses isn't a good idea</h2>
<p>Early on in the year, I had the bright idea to spin out OnlineOrNot's internal feature flag service, and make it a SaaS project.</p>
<p>The thought process was essentially:</p>
<p>| <img src="/assets/lessons-from-two-years-of-saas-operation/on-doing-two-projects.png" alt="Working on two projects at once, a tweet"> |
| :---------------------------------------------------------------------------------------------------------------------: |
|                                     <em>it was not different this time, dear reader.</em>                                      |</p>
<p>I spent a weekend building a prototype SaaS app, launched it, and nothing happened. I talked to potential customers throughout my professional network, and folks were just not that keen. It turns out, I wasn't that keen either. I worked a bit on content marketing, let the project sit for a few months, and eventually shut it down.</p>
<p>It cost me time I could have spent building features for OnlineOrNot, and was just a distraction, in the end.</p>
<p>It turns out no one cares if you came up with a really fast way to do feature flags over Cloudflare's shiny new developer platform - your service didn't exist a week ago and has almost no features, and there are feature flag services that have been around for a very long time, with an adequate feature set.</p>
<p>It made me realise OnlineOrNot's moat is that I don't plan on giving up on it, I'm genuinely interested in the problem space, and I keep myself employed full-time so that I can build OnlineOrNot the way I want, with zero risk to my livelihood.</p>
<h2>Ship small, invalidate your assumptions</h2>
<p>The trouble with holding off from releasing a feature because it "isn't ready", is that folks can't tell you if your assumptions are wrong, until you spend months building on top of your incorrect assumptions.</p>
<p>OnlineOrNot's Status Page feature came from a conversation with a customer trying to figure out how to sign up hundreds of email addresses to get notifications when an uptime check fails.</p>
<p>I built something as quickly as possible, showed the customer (and a few other customers, friends, and colleagues), realised where I went wrong, and tried again.</p>
<p>OnlineOrNot itself was released after only <a href="/building-saas-in-one-week-how-built-onlineornot">7 days of part-time development</a>, after all.</p>
<h2>Talk to potential customers, even if you think you have nothing for them</h2>
<p>I was contacted by a CTO at some point during the year asking if OnlineOrNot supported some particular feature.</p>
<p>My normal reaction would've been to just say "sorry, no" and leave it at that, but I got curious and started asking what they'd like to achieve with the feature, and told them how I figured I would build the feature.</p>
<p>They signed up to a paid plan the next day, they've been customers ever since, and the feature gets used by other customers too.</p>
<p>That's all for what I learned this year. Below is a sort of CHANGELOG of what I shipped in OnlineOrNot.</p>
<h2>What I actually shipped</h2>
<p>I recently wrote <a href="https://maxrozen.com/2022-just-keep-shipping">2022: I just kept shipping</a>, and while it's nice to tell folks "just keep shipping", I think it's valuable to point out just how much you can get done if you give yourself 2 hours, and ship ever day.</p>
<p>There were months where my focus was more on marketing/writing docs/helping and talking to customers/drinking wine by the lake/I just didn't feel like writing code, so not all months resulted in the same amount of feature development, but I still released my articles/docs/etc ASAP, and iterated on them while they were live.</p>
<h3>January</h3>
<ul>
<li>The uptime dashboard now actually tracks the last 24h of uptime (used to be response time)</li>
<li>Added auto-refresh to the uptime dashboard so folks don't need to remember to refresh the page</li>
<li>Made the uptime dashboard mobile-friendly</li>
<li>Migrated the database from Intel to ARM</li>
<li>Made the rest of the web app mobile-friendly</li>
<li>Made it clearer <em>why</em> OnlineOrNot thinks a check is failing</li>
<li>Reached 50 million all-time uptime checks</li>
<li>Published <a href="/incident-management/incident-response/communicating-users-incidents">Communicating to Users During Incidents</a></li>
<li>Cleaned up my sign-up form to remove distractions, halving the drop-off rate in the process</li>
<li>Built the first screen of a new onboarding flow and immediately released it</li>
<li>Added a second screen to the onboarding flow to enable people to join the mailing list</li>
<li>Added a third screen to the onboarding flow to figure out how people found out about OnlineOrNot (majority of folks come from me shitposting on Twitter, and commenting on hacker news) and what their name/company name is (so I finally stopped sending emails starting with "Hey there,")</li>
<li>Added a fourth screen to the onboarding flow to let folks add their whole team in one go (it's just a textfield, and I parse it to invite folks)</li>
<li>Added GitHub Auth to the login/signup screen</li>
<li>Added a free trial sign-up option to the onboarding flow (that's right, I only added free trials after 11 months of running the business)</li>
</ul>
<h3>February:</h3>
<ul>
<li>After shipping the free trial sign-up, I had 14 days to implement what happens at the end of a free trial, so I shipped that separately. In the meanwhile, I was manually sending folks their "getting the most out of your free trial" emails</li>
<li>Removed the paywall from a few uptime monitoring features</li>
<li>Wrote the <a href="/what-learned-running-saas-for-year">first year's version of this blog post</a>, around 33k people read it</li>
<li>Wrote <a href="/uptime-monitoring-best-practices">How OnlineOrNot uses OnlineOrNot to run OnlineOrNot</a></li>
<li>Added a "you're on a free trial!" banner to OnlineOrNot</li>
<li>Wrote a pair of guides for the docs to help folks getting started with their free trial</li>
<li>Added human-sounding errors when uptime assertions fail
<ul>
<li>for example: "Looked for a value at <code>$.data.viewer.sendWeeklyReport</code>, found <code>true</code>, which we expected to be <code>false</code>"</li>
</ul>
</li>
<li>Added the ability to report on SSL certificate validity</li>
<li>Worked on the deployment pipeline, getting a 7 minute build down to 1 minute 45 seconds</li>
</ul>
<h3>March</h3>
<ul>
<li>Got tired of forgetting to write docs after releasing features, so I moved the <a href="/docs">docs</a> into my app's monorepo, and gave it the same look and feel as the rest of my web presence (I still get customers telling me they wish their app had this - but I'm not planning on spinning off a docs product)</li>
<li>Made it possible to BYO Twilio account to enable unlimited SMSes for free</li>
<li>Added a "home screen" to OnlineOrNot</li>
<li>Used my new docs as a template for an <a href="/incident-management">Incident Management</a> mini-site. Before this, I had a few blog posts about on-call, run-books, and navigating between them was annoying.</li>
</ul>
<h3>April</h3>
<ul>
<li>Moved house, took the time to explore my new town/region, so didn't ship much</li>
<li>Tested and tweaked copy on my landing pages</li>
<li>Unified browser checks and regular uptime checks into a single dashboard</li>
</ul>
<h3>May</h3>
<ul>
<li>Moved most of my DNS management from AWS to Cloudflare (I got hired by Cloudflare for $DAY_JOB, and it's a lot faster than AWS)</li>
<li>Built and shipped the DNS system behind OnlineOrNot's Status Pages (what allows me to simultaneously display status pages at custom domains and on OnlineOrNot's subdomain)</li>
<li>Started showing the last 14 days of incident history on status pages</li>
<li>Actually made it possible for my paid customers to sign up for a status page at a custom domain</li>
</ul>
<h3>June</h3>
<ul>
<li>Made it possible to manually add incidents to status pages</li>
<li>Made it possible to manually update existing incidents on status pages</li>
<li>Started an <a href="/docs/early-access-program">Early Access Program</a> so I didn't have to feature flag early features one account at a time</li>
<li>Migrated most of my core uptime check workload off AWS and onto fly.io, and <a href="/on-moving-million-uptime-checks-onto-fly-io">wrote about it</a></li>
<li>Made it possible to add components to a status page</li>
<li>Made it possible to pick components affected by an incident</li>
<li>Published <a href="/incident-management/incident-response/writing-your-first-runbooks">Writing your first runbooks</a></li>
<li>Published <a href="/incident-management/incident-response/guidelines-for-writing-better-runbooks">Guidelines for writing better runbooks</a></li>
</ul>
<h3>July</h3>
<ul>
<li>Made it possible to link existing uptime checks to a status page, and automatically start and resolve incidents based on uptime check data</li>
<li>Hacker News went down, 19k people ended up checking out https://hackernews.onlineornot.com, and I <a href="/how-i-accidentally-told-19k-people-hacker-news-was-down">wrote about the experience</a></li>
<li>Made it possible to sign up to status page updates, without needing to sign up to OnlineOrNot</li>
<li>Rewrote the whole status page app from Next.js to Remix, because shipping 100KB of JavaScript for a basic HTML page felt ridiculous</li>
<li>Published <a href="/incident-management/learn/postmortem-templates">Postmortem Templates</a></li>
<li>Edited and republished <a href="/incident-management/on-call/improving-your-teams-on-call-experience">Improving your team's on-call experience</a></li>
</ul>
<h3>August</h3>
<ul>
<li>Summer holidays</li>
<li>Added a way to highlight active incidents on the status page</li>
<li>Made it possible to make a status page private (password-protection)</li>
<li>Added live system metrics to status pages</li>
</ul>
<h3>September</h3>
<ul>
<li>Used OnlineOrNot to acquire the <a href="https://twitter.com/onlineornot">@OnlineOrNot</a> handle on Twitter</li>
<li>Worked on my landing pages</li>
<li>Fought off an attack from spammers looking to abuse OnlineOrNot</li>
<li>Added support for webhook events from Uptime Robot</li>
<li>Migrated the business from Australia to France</li>
</ul>
<h3>October</h3>
<ul>
<li>Holidays in Australia</li>
<li>Added Cloudflare Workers to the list of "backup uptime checkers" that OnlineOrNot uses to verify a website is actually down</li>
<li>Published <a href="/saving-your-team-from-alert-fatigue">Saving your team from alert fatigue</a></li>
</ul>
<h3>November</h3>
<ul>
<li>Made it even clearer <em>why</em> OnlineOrNot thinks a check is failing</li>
</ul>
<h3>December</h3>
<ul>
<li>Added support for HTTP Basic Auth in uptime checks</li>
<li>Hit 120 million uptime checks</li>
<li>Made it possible to duplicate an uptime check</li>
</ul>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Saving your team from alert fatigue]]></title>
            <link>https://onlineornot.com/saving-your-team-from-alert-fatigue</link>
            <guid>https://onlineornot.com/saving-your-team-from-alert-fatigue</guid>
            <pubDate>Tue, 01 Nov 2022 08:45:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>It's a story as old as the web itself: someone on your team gets excited to install a new tool.</p>
<p>The tool promises to finally give you a clear view into the problems your users have with your product.</p>
<p>Your team agrees to give it a go.</p>
<p>The errors start coming...</p>
<p>...and they don't stop coming...</p>
<p>Soon enough, most of your team has either created an email filter to manage all the alerts, or has unsubscribed themselves entirely. Just like all the other tools.</p>
<p>Welcome to alert fatigue.</p>
<p><strong>What this article covers:</strong></p>
<ul>
<li><a href="#what-is-alert-fatigue">What is Alert Fatigue</a></li>
<li><a href="#beating-alert-fatigue">Beating Alert Fatigue</a>
<ul>
<li><a href="#make-alert-management-a-team-effort">Make alert management a team effort</a></li>
<li><a href="#clean-up-your-alerts">Clean up your alerts</a></li>
<li><a href="#categorize-your-alerts-and-notifications">Categorize your alerts and notifications</a>
<ul>
<li><a href="#high-priority-alert---immediate-action-required">High priority alert - immediate action required</a></li>
<li><a href="#medium-priority-alert---awareness-required">Medium priority alert - awareness required</a></li>
<li><a href="#low-priority-alert---things-to-fix-during-spare-timediagnostics">Low priority alert - things to fix during spare time/diagnostics</a></li>
</ul>
</li>
</ul>
</li>
</ul>
<h2>What is Alert Fatigue</h2>
<p>In short, alert fatigue happens when you start to become desensitized to alerts.</p>
<p>It tends to happen after receiving multiple alerts with similar content (often with limited context into the problem, so not enough information to fix the alert).</p>
<p>Alert fatigue can strike relatively quickly when there's a large number of alerts (often irrelevant to the developer), and there's not enough time to resolve all of the issues.</p>
<h2>Beating Alert Fatigue</h2>
<p>This section suggests a few ways you can beat alert fatigue.</p>
<p>It boils down to: making alert management a team effort, continuously cleaning up your alerts, and categorizing your alerts.</p>
<h3>Make alert management a team effort</h3>
<p>There is nothing worse than being paged for a system you didn't create, with an alert you have no autonomy over.</p>
<p>If your team has an on-call roster, it should have autonomy to create, modify, and remove alerts as it requires. Sometimes alerts are flaky, and need to be temporarily disabled or have its thresholds changed, and sometimes an alert no longer fits the team's needs, and should be removed entirely.</p>
<h3>Clean up your alerts</h3>
<p>Cleaning up your alerts doesn't have to be a big task to do in one go, you can continuously clean-up alerts as part of an on-call roster.</p>
<p>The idea is to look at every single alert as they fire, and figure out if there's a clear human response required. If an alert fires, and there's no human action required as a result, it shouldn't be taking your team's attention, and needs to be removed.</p>
<p><strong><em>Getting paged constantly while on-call is a symptom of a broken system</em>.</strong></p>
<p>Either the system is too unstable, and needs time invested to make it resilient, or your alerts are too verbose and their thresholds need tweaking, or they require no action, and need to be deleted.</p>
<h3>Categorize your alerts and notifications</h3>
<p>The beauty of having an <a href="/incident-management/on-call/improving-your-teams-on-call-experience">on-call roster</a> within your team, is that your team agrees that one person gets to focus entirely on either fighting fires, or improving the on-call experience for the week.</p>
<p>Grooming your alerts <em>greatly</em> improves your team's on-call experience. Broadly speaking, sending every notification or alert your system generates to the same location <strong>is a mistake</strong>.</p>
<p>There are three types of alert that we care about, and they need to go to different places, as they help us do different things. Not every alert needs to wake someone up. Not every alert needs to go to a Slack channel.</p>
<p>While it's easy to sit on my armchair here and give you clear categories, in reality your team will often get high priority alerts that aren't really high priority, and low priority alerts that really should be looked at immediately. Part of your team's on-call responsibilities should be to review the alerts that fired each week, and re-categorize alerts as necessary.</p>
<h4>High priority alert - immediate action required</h4>
<p>This covers things like: your website being completely unreachable after several checks in a 5 minute period, your SSL certificate expiring, or customers being unable to pay for your product.</p>
<p>These are the types of alerts that need go to someone's phone (SMS, call or pager app under a high priority), and wake them up if they happen outside of business hours.</p>
<p>Once your team is aware of the alert, they start up an incident room, and start working through your <a href="/incident-management/incident-response/writing-your-first-runbooks">team's runbook</a> for the service affected.</p>
<h4>Medium priority alert - awareness required</h4>
<p>This covers things like: your system's database backup failing, and starting to run out of disk space on your database server.</p>
<p>They're still particularly <em>bad</em> things, but nothing that immediately impacts your customers.</p>
<p>These types of alerts that can go under a lower priority in your pager tool or to a team channel in Slack/Discord/Microsoft Teams. Since they're lower priority, your team should only investigate if there's no active incident, and generally only during business hours.</p>
<h4>Low priority alert - things to fix during spare time/diagnostics</h4>
<p>This covers things like: individual 5xx errors, server timeouts, and JavaScript errors.</p>
<p>Most commonly, these are the alerts your team gets from Sentry. It may be tempting to dump them into a Slack channel, but your team will quickly start to mute/ignore the channel.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[The unreasonable effectiveness of shipping every day]]></title>
            <link>https://onlineornot.com/unreasonable-effectiveness-shipping-daily</link>
            <guid>https://onlineornot.com/unreasonable-effectiveness-shipping-daily</guid>
            <pubDate>Fri, 26 Aug 2022 06:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>It's fairly common for folks in tech to dream of quitting their day job and working on their side projects. I find when you ask them how their projects are going, they tend to have 2-3 projects running at the same time, none of the projects are actually available for potential users to try out.</p>
<p>The question they seem to ask me most is "you seem to complete your projects, how do you stay motivated?"</p>
<p><strong>My secret?</strong> It's a habit. I ship something every day.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#some-context">Some context</a></li>
<li><a href="#forming-the-habit">Forming the habit</a></li>
<li><a href="#how-i-work">How I work</a></li>
<li><a href="#how-this-works-in-practice">How this works in practice</a></li>
<li><a href="#summary">Summary</a></li>
</ul>
<h2>Some context</h2>
<p>For context, since 2017 I have shipped:</p>
<ul>
<li>a job board (shut-down)</li>
<li>an appointment scheduling service (shut-down)</li>
<li>a room booking service (shut-down)</li>
<li>a GraphQL API testing service (shut-down)</li>
<li>a web performance monitoring service (shut-down)</li>
<li>a <a href="https://maxrozen.com/">weekly blog about React</a> (paused while I work on my SaaS)</li>
<li>a <a href="https://rozenmd.gumroad.com/l/useEffect-by-example">book about React's useEffect hook</a> (still gets sales)</li>
<li>an <a href="https://maxrozen.com/beginners-guide-to-react-testing">introductory book for testing React</a></li>
<li>an <a href="/">incident management service</a> (uptime monitoring + status pages)</li>
</ul>
<p>I initially got started because I wanted to learn to use React and GraphQL better for my day job, but these days I enjoy running the business and learning about marketing.</p>
<h2>Forming the habit</h2>
<p>Actually forming the habit is the most difficult part, but I promise it gets easier.</p>
<p>To get started on each project, I prioritize getting an <strong>absolute</strong> minimal "viable" product live and able to accept customers as soon as possible. In the case of my most recent project (that I've now been working on continuously for two hours per workday since February 2021), this took <a href="https://maxrozen.com/2021-strangers-paid-my-macbook#the-approach">one week</a>.</p>
<p><em>I say "viable" because while users could login, subscribe to a paid plan, add pages to monitor, and receive downtime notifications, it wasn't an actually viable solution to their problem. It wasn't until I added significantly more features that users started paying me for the service.</em></p>
<p>Once a product/blog article/book is live, you can start getting users. Once you start getting users, their feedback helps inform what you should work on next.</p>
<h2>How I work</h2>
<p>I tend to work on my personal projects in 2 hour increments (before my work day starts), partly because that's how long I can focus at a time, and partly because I prefer to chip away at a problem, rather than making one big push.</p>
<p>Usually by the end of those two hours, I'll have something running live.</p>
<p>You'll find two hours isn't nearly enough to "complete" a feature. Get used to being embarrassed by your incomplete project.</p>
<p>In other words, limit scope as much as possible.</p>
<p>Practically, this means releasing forms that gather less data than needed, or only building the UI portion of a feature before the backend is ready. For books, it could mean releasing a chapter as an email to your newsletter subscribers. For SaaS, I'd recommend starting an <a href="https://onlineornot.com/docs/early-access-program">early access program</a> and let users that opt-in know when they're looking at incomplete features.</p>
<h2>How this works in practice</h2>
<p>It's all well and good for me to tell you to "ship more!", but it feels pointless without providing a concrete example of how it works in practice for a real SaaS product.</p>
<p>Here's roughly how OnlineOrNot's uptime monitoring was built:</p>
<ol>
<li>Released a simple form that accepted a URL to monitor, and a name for the uptime check</li>
<li>Made it possible to receive email notifications when an uptime check is down</li>
<li>Made it possible to search for text on the page while checking</li>
<li>Made it possible to customize the uptime check frequency</li>
<li>Made it possible to sign up for a Slack notification when your uptime check is down</li>
<li>Made the Slack integration actually send a notification when the uptime check is down</li>
<li>Made it possible to wait for several failed uptime checks before sending a notification</li>
<li>Made it possible to edit existing uptime checks</li>
</ol>
<p>All of these features were delivered while the service was live, accepting users, and those users were helping inform the direction of the product.</p>
<p>You'll find an Early Access Program helps a lot for cases like Steps 5 and 6 above - I'll often release forms to my early access users that only save their preferences, without actually taking the preference into account until a few days later when I fully complete the feature.</p>
<p>Caveat: the types of folks that opt-in to an early access program tend to be accepting of this kind of thing, I wouldn't recommend shipping incomplete product features to your entire user base unless you're extremely early in your product's development.</p>
<h2>Summary</h2>
<p>Consistent effort over a long time helps you accomplish things you think are impossible.</p>
<p>Get used to being embarrassed by your incomplete project.</p>
<p>I wish I realized this in high school/university.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[How I accidentally told 19k people Hacker News was down]]></title>
            <link>https://onlineornot.com/how-i-accidentally-told-19k-people-hacker-news-was-down</link>
            <guid>https://onlineornot.com/how-i-accidentally-told-19k-people-hacker-news-was-down</guid>
            <pubDate>Fri, 12 Aug 2022 08:45:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Hacker News is a forum run by Y Combinator that tends to focus on topics relating to computer science and entrepreneurship. During an average month, Hacker News serves almost 10 million visitors.</p>
<p>Hacker News had a series of extremely rare outages in early July 2022:</p>
<p><img src="/assets/how-i-accidentally-told-19k-people-hacker-news-was-down/onlineornot-downtime-stats-for-hn.png" alt="OnlineOrNot downtime stats for Hacker News"></p>
<p>On July 8, 06:20am UTC, the hard drives in Hacker News's primary server <a href="https://news.ycombinator.com/item?id=32024036">stopped working</a>, and they switched over to the failover server in just over an hour.</p>
<p>The failover server chugged along for almost six and a half hours before the hard drives <a href="https://news.ycombinator.com/item?id=32031330">also failed</a> at 12:47pm UTC.</p>
<p>At this point, it looks like <a href="https://news.ycombinator.com/item?id=32031243">all four hard disks reached 40,000 hours of uptime</a>, and a bug in their firmware caused them to fail.</p>
<h2>About OnlineOrNot</h2>
<p>OnlineOrNot is a new kind of Incident Management software that runs on Cloudflare Pages. OnlineOrNot provides uptime monitoring, with integrated status pages that automatically update.</p>
<p>The architecture for our Status Page service is relatively simple: it's a Remix frontend app (we use Remix as a React framework that lets us write React components, and output HTML/CSS with minimal JS). We do not currently use a client-side framework. As a result of this approach, our average page weight is around 13KB (instead of 200KB).</p>
<p>Our Remix app runs on <a href="https://pages.cloudflare.com/">Cloudflare Pages</a>, which itself uses Cloudflare's global network, without direct access to my database.</p>
<p>To workaround this limitation, we built an Express API server running on <a href="https://fly.io/">fly.io</a>. While I usually opt for a serverless solution for early versions of features, I wanted to avoid cold-starts. Another advantage of using fly.io is that we can deploy replicas of the API server globally. However thanks to Cloudflare Pages' ability to access Cloudflare's cache, this won't be necessary as most requests under high load won't hit my origin server.</p>
<h2>Letting folks know what happened</h2>
<p>While testing the status page feature described above, I decided to monitor Hacker News. To my surprise, the first result came back saying Hacker News was down, and my system automatically created an <a href="https://hackernews.onlineornot.com/incidents/dLKw9DQ-VB2D">incident on the status page</a>.</p>
<p>When Hacker News came back up (running on the failover server), naturally, I <a href="https://news.ycombinator.com/item?id=32023848">created a thread on Hacker News</a>, thinking no one would notice or care that Hacker News went down.</p>
<p>That thread quickly reached the front page, was in the number one spot for about half an hour, and thousands of users were sent to my <a href="https://hackernews.onlineornot.com/">Hacker News Status Page</a>:</p>
<p><img src="/assets/how-i-accidentally-told-19k-people-hacker-news-was-down/onlineornot-cloudflare-stats.png" alt="OnlineOrNot&#x27;s Cloudflare stats"></p>
<p>Eventually Hacker News went down again as the disks in their failover server also failed, and my automated monitoring kicked in again, creating a <a href="https://hackernews.onlineornot.com/incidents/0LB6mQLmkozD">second incident</a> page.</p>
<p>This second outage lasted roughly seven hours, with no updates from the Hacker News team for several hours (understandably so, the failover servers failed in the middle of the night). During this time, Google indexed the status page I built, and temporarily gave me the top search result on Google for "<a href="https://hackernews.onlineornot.com/">hacker news status page</a>".</p>
<p>Now that I know the feature works (and at scale - I'm grateful for the load test!), I'll be releasing an embarrassingly under-featured MVP this week.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Postmortem Templates]]></title>
            <link>https://onlineornot.com/postmortem-templates</link>
            <guid>https://onlineornot.com/postmortem-templates</guid>
            <pubDate>Wed, 27 Jul 2022 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>After things go wrong and the incident is resolved, it's time to learn.</p>
<p>Generally in tech, the act of reviewing the cause of an incident is known as a postmortem, though there may be other names used such as "<a href="https://www.atlassian.com/incident-management/handbook/postmortems#postmortem-process">Post-Incident Review</a> (PIR)" or "<a href="https://wa.aws.amazon.com/wat.concept.coe.en.html">Correction of Error</a> (COE)"</p>
<p>Regardless of what you call this process, you want to use a template so that your team has a standardized way of documenting what went wrong, and what your organization will do differently in the future to avoid repeat issues. It also helps to store these documents in a centralized place so that your teams have the opportunity to learn from each other's mistakes (Google Docs, Confluence/Notion, etc).</p>
<h2>A note on root causes in the cloud</h2>
<p>It's common for postmortem templates to include a <a href="https://en.wikipedia.org/wiki/Five_whys">five whys</a> analysis to determine a root cause.</p>
<p>Be aware that when building software in the cloud, there is rarely a "true" root cause due to the sheer complexity of building distributed software. It's still worth attempting to find improvements to your organization's processes to avoid repeat incidents.</p>
<h2>Amazon's template</h2>
<ul>
<li>What happened?</li>
<li>What was the impact on customers and your business?</li>
<li>What was the root cause?</li>
<li>What data do you have to support this?
<ul>
<li>Especially metrics and graphs</li>
</ul>
</li>
<li>What were the critical pillar implications, especially security?
<ul>
<li>When architecting workloads you make trade-offs between pillars based upon your business context. These business decisions can drive your engineering priorities. You might optimize to reduce cost at the expense of reliability in development environments, or, for mission-critical solutions, you might optimize reliability with increased costs. Security is always job zero, as you have to protect your customers.</li>
</ul>
</li>
<li>What lessons did you learn?</li>
<li>What corrective actions are you taking?
<ul>
<li>Actions items</li>
<li>Related items (trouble tickets etc)</li>
</ul>
</li>
</ul>
<p>Source: https://wa.aws.amazon.com/wat.concept.coe.en.html</p>
<h2>Atlassian's template</h2>
<ul>
<li>Incident summary (what happened, why, incident severity, how long did the incident last?)</li>
<li>Leadup (the series of events that lead to the incident)</li>
<li>Fault (describe how the software misbehaved)</li>
<li>Impact (determine how many users were impacted, over which time period)</li>
<li>Detection (when did the team detect the incident, and how)</li>
<li>Response (who responded to the incident, when, and what did they do?)</li>
<li>Recovery (how was the system restored, what steps were needed to restore the system to a functioning state?)</li>
<li>Timeline (an incident timeline, using the UTC timezone to standardize time)</li>
<li>Five whys (describe the incident, ask why it happened, then ask why <em>that</em> happened, recursively, five times)</li>
<li>Root cause (what needs to be changed to avoid this happening again)</li>
<li>Backlog check (review your backlog to see if any unplanned work could have prevented this issue)</li>
<li>Recurrence (use the root cause to see if other incidents had the same root cause. Ask why the incident happened again)</li>
<li>Corrective actions (describe the work needed to prevent the incident happening again, who is responsible for completing the work, and when)</li>
</ul>
<p>Source: https://www.atlassian.com/incident-management/postmortem/templates</p>
<h2>Summary</h2>
<p>In short, it doesn't matter which template you pick (or if you create one from scratch), as long as it's consistently used by your organization.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Improving your team's on-call experience]]></title>
            <link>https://onlineornot.com/improving-your-teams-on-call-experience</link>
            <guid>https://onlineornot.com/improving-your-teams-on-call-experience</guid>
            <pubDate>Tue, 26 Jul 2022 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Your engineers probably dislike going on-call for your services.</p>
<p>Some might even dread it.</p>
<p>It doesn't have to be this way.</p>
<p>With a few changes to how your team runs on-call, and deals with recurring alerts, you might find your team starting to enjoy it (as unimaginable as that sounds).</p>
<p>I wrote this article as a follow-up to <a href="/getting-over-on-call-anxiety">Getting over on-call anxiety</a>. While that article handles what individual engineers can do to psychologically prepare themselves for on-call, this article describes what the team can do to improve the on-call experience overall.</p>
<h2>You build it, you run it</h2>
<p>This point is important, which is why it's the first tip in this article: you cannot hire an SRE/DevOps team to magically make your operational problems go away.</p>
<p>If your engineers build a service and it's <strong>business-critical</strong>, they need to be on-call for it.</p>
<p>Handling on-call this way incentivizes your team to improve the service's operations over time, and reduces the risk of "just throw it over the fence to the ops folks" mentality that can develop when a team isn't on the hook for their own code.</p>
<h2>Continuously improve your service</h2>
<p>On-call shouldn't just be an additional responsibility to respond to alerts on top of your team's regular scheduled work (no sprint work, no picking up tasks from the backlog, nothing).</p>
<p>Instead when folks are rostered on-call, they should be working on tasks that improve the quality of life for on-call.</p>
<p>This stream of work will come up naturally if you take each alert that fires as an opportunity to improve your service. When things break, take the time to ensure an alert never fires again.</p>
<p><strong>Especially</strong> if the alert is unactionable.</p>
<p><em>A quick note on unactionable alerts</em>: If your service has alerts that do not have a single, clear, human response, <strong>the alert should be deleted</strong>. If it's a flaky alert, your team should prioritize fixing the issue as soon as possible.</p>
<p>Getting paged constantly while on-call is a symptom of a broken system. Either the system is too unstable, and needs time invested to make it resilient, or your alerts are too verbose and their thresholds need tweaking, or they require no action, and need to be deleted.</p>
<p>Over time, applying long-term fixes to the causes of your alerts will make your service more resilient, and make your team's time on-call a lot less stressful.</p>
<p>Google make this possible for their own SRE's by blocking them from spending more than 50% of their time on operational work:</p>
<blockquote>
<p>We cap the amount of time SREs spend on purely operational work at 50%; at minimum, 50% of an SRE’s time should be allocated to engineering projects that further scale the impact of the team through automation, in addition to improving the service.</p>
</blockquote>
<p>— <a href="https://sre.google/sre-book/being-on-call/">Google's SRE Book</a></p>
<h2>Good processes make a good on-call experience</h2>
<p>If your team relies on engineers to regularly improvise while on-call, they (and your customers) are going to have a bad time.</p>
<p><a href="/incident-management/incident-response/writing-your-first-runbooks">Runbooks</a>, up-to-date documentation, and standard operating procedures give your on-call engineers direction on how to respond to, and escalate, incoming alerts. You don't want them to be figuring this stuff out by themselves at 2am.</p>
<p>If you've already got these in place, <a href="/incident-management/incident-response/guidelines-for-writing-better-runbooks#encourage-newcomers-to-use-your-runbooks-and-make-edits">have your newest team members review them</a>. Chances are, there are blindspots that you and your team can't see due to familiarity with the services and their docs.</p>
<p>If <em>you're</em> the new member of the team, take the initiative to improve the documentation yourself. Try going through runbooks and see if you have enough context to complete the steps. A common mistake senior engineers will make is assuming their colleagues know where to run commands, or where certain repos are (particularly when engineers give a service one name, while the codebase lives in a repo with a completely different name).</p>
<p>Your documentation should be as explicit as possible.</p>
<h2>Have a primary on-call, and a secondary on-call</h2>
<p>In order to help drive the continuous improvement aspect of on-call, it helps to have an engineer on secondary on-call.</p>
<p>While the primary on-call engineer answers pages at any time of the day for a week, the secondary on-call engineer's role is to make the primary on-call engineer's life easier. They can assist primary with investigations when a particularly bad incident occurs, but their main focus should be improving service reliability (by improving the codebase/writing tests, fixing flaky alerts, deleting unactionable alerts) and generally ensuring primary doesn't get paged each week for the same reason.</p>
<p>While in the past we've had secondary on-call for only one week at a time, we found that due to ramp up time, a month of secondary on-call worked better.</p>
<p>As secondary on-call doesn't answer pages overnight, you can typically assign newer engineers to the role so they get familiar with your team's documentation and runbooks, and can help improve them further.</p>
<h2>Handover between your on-call engineers</h2>
<p>You should hold a meeting to handover on-call responsibilities between engineers as one shift ends and another begins. We typically hold these in the middle of the week, as most major issues for the week typically surface by then.</p>
<p>In the meeting, the engineer whose on-call shift is ending talks through the alerts that were fired, whether they were resolved, or if they still require attention. Commonly recurring themes (types of alerts that fire often) should be added to the on-call backlog, and prioritized to be fixed by the on-call engineers.</p>
<p>Let anyone on your team attend these meetings (particularly new employees), as it helps them get familiar with the service that they'll eventually be on-call for.</p>
<h2>Pay your staff</h2>
<p>Going on-call overnight is additional work, whether or not an engineer gets paged.</p>
<p>You need to pay a bonus on top of their regular salary for the week your engineers go on-call. In many countries, this is now a legal requirement (you may want to double check this with your legal team).</p>
<p>Particularly if you expect them to do their job well.</p>
<h2>Conclusion</h2>
<p>I wrote this article as a former-employee of a very large tech company (thousands of engineers). In our case, we had our regular software engineers that build the service also handle on-call for it.</p>
<p>Be aware that blindly implementing recommendations will not automatically make your company successful. Large tech companies adopt these practices as a means of managing their success - not because these practices alone make them successful.</p>
<p>That being said, if your on-call staff are constantly fighting fires, or the on-call experience isn't generally improving week to week, hopefully some of these recommendations will help.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[On moving over a million uptime checks per week onto fly.io]]></title>
            <link>https://onlineornot.com/on-moving-million-uptime-checks-onto-fly-io</link>
            <guid>https://onlineornot.com/on-moving-million-uptime-checks-onto-fly-io</guid>
            <pubDate>Wed, 29 Jun 2022 08:45:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>The other day, a friend told me about <a href="https://fly.io">fly.io</a>'s nice developer experience (DX). For my day job, I work on improving <a href="https://github.com/cloudflare/wrangler2">wrangler2</a>'s DX, so naturally it had me curious.</p>
<p>I went from "I'll just play around with it, maybe give it a toy workload" to "holy shit, what if I quickly rewrite my business's AWS Lambda + SQS stack to fit entirely within their free tier" in about 90 minutes.</p>
<p>It wasn't <em>that</em> simple in the end, but I did manage to migrate most of my active workload from AWS Lambda to fly.io.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#the-project">The project</a></li>
<li><a href="#what-that-looks-like-on-aws-lambda">What that looks like on AWS Lambda</a></li>
<li><a href="#the-naive-approach-to-migrating-off-aws-lambda">The naive approach to migrating off AWS Lambda</a></li>
<li><a href="#replacing-sqs-with-redis">Replacing SQS with Redis</a></li>
<li><a href="#messing-around-with-ram-and-finding-out">Messing around with RAM and finding out</a></li>
<li><a href="#the-cost-benefit-calculation">The cost-benefit calculation</a></li>
<li><a href="#what-i-kept-on-aws-lambda">What I kept on AWS Lambda</a></li>
</ul>
<h2>The project</h2>
<p>I built <a href="https://onlineornot.com/">OnlineOrNot</a> as the 200th uptime monitor for the Internet. The idea was: I hate all the other solutions, I'm a frontend developer that can handle infra and backend, let's solve the problem the way I would solve it, while listening to any customers I might get along the way.</p>
<p>On the backend, an uptime monitor is simple(ish):</p>
<ul>
<li>A process to query the database, find the checks that are ready to run, and queue them up as a job</li>
<li>A queue to hold jobs while waiting to be run</li>
<li>A process to pick up jobs from the queue, and run the check (and do something with the result)</li>
</ul>
<p>To save money on the off-chance no one other than me ever used OnlineOrNot, I built it on AWS Lambda. Costs nothing if no one is using it (apart from the database, which I kept within the free-tier while building anyway).</p>
<h2>What that looks like on AWS Lambda</h2>
<p>On AWS Lambda, the system I built looks like this:</p>
<ul>
<li>One AWS Lambda function that runs every minute to query the database</li>
<li>A Simple Queue Service (SQS) queue to hold jobs</li>
<li>One AWS Lambda function to pick up jobs, invoke (yet another) AWS Lambda function that performs the check, get the result, and do something with it</li>
</ul>
<p>The vast majority of my AWS bill was coming from that last function - every time there's an uptime check that hits a timeout, I get billed for the full 10 seconds it spends waiting, twice (once for the check, once for the function waiting for the check), even though the functions aren't doing anything in that time.</p>
<h2>The naive approach to migrating off AWS Lambda</h2>
<p>My initial idea was to take that last AWS Lambda function (the one that has to wait for every timeout) code, use the <a href="https://github.com/bbc/sqs-consumer">SQS Consumer</a> library to pick jobs off the SQS queue, build a Dockerfile to run that code, and boom, easily moved my biggest Lambda cost into a free VM.</p>
<p>It turned out that over a million uptime checks per week is a lot, and the SQS Consumer library couldn't keep up. It also didn't process the jobs in order, which made some jobs never even run their uptime check!</p>
<h2>Replacing SQS with Redis</h2>
<p>With the main solution to processing SQS queues not being fit for my use, I spent a day thinking about how to re-architect my entire solution to run in VMs.</p>
<p>After a bit of research, I found <a href="https://github.com/taskforcesh/bullmq">BullMQ</a>. BullMQ replaces my SQS queue with a Redis instance, and handles the job creating and running for me.</p>
<p>After re-architecting, my fly.io-based solution looks like this:</p>
<ul>
<li>A permanently running VM running Redis with persistent storage (in case it goes down)</li>
<li>A permanently running VM that uses <a href="https://github.com/node-cron/node-cron">node-cron</a> to run every minute, and queues up uptime checks to run</li>
<li>A permanently running VM that picks up jobs from Redis, runs the checks</li>
</ul>
<p>Initially I gave each VM 256MB of RAM, flipped a feature flag to stop processing checks on AWS Lambda, and let fly.io start picking up checks to queue instead.</p>
<p>It worked!</p>
<p>Sort of.</p>
<h2>Messing around with RAM and finding out</h2>
<p>It turns out if you overload a VM with too many jobs, some jobs will "stall" - be neither in the queue ready for work, nor actively worked on by the job runner.</p>
<p>Thankfully I built around this possibility already with my AWS Lambda stack (I requeue if a job hasn't run when we expected it to), and the code held up when permanently running in VMs.</p>
<p>From my customer's perspective, some of their uptime checks would just run every 10 minutes instead of every 5 minutes. While better than failing to run entirely, it wasn't a great outcome.</p>
<p>It took me about 30 minutes each morning over a week to fine tune the number of worker VMs, RAM allocation, and concurrency before working out I just needed a single VM with 512MB RAM.</p>
<h2>The cost-benefit calculation</h2>
<p>I've known this day was coming for about a year, ever since my first customer with over a thousand websites came in, smashed my AWS bill, and made me rethink my unlimited pricing.</p>
<p>Don't make paid resources unlimited, kids. Not even once.</p>
<p>What happened was, my AWS Lambda bill jumped to over $100 USD/mo, and I started thinking "well, surely a continuously running VM would be cheaper", and it is! I roughly worked out how much it'd cost me to rewrite a serverless app to continuously run in a VM (about a week, 2 hours per day before work), and sat on the idea for a year, since building features for my customers was a better use of my time.</p>
<p>What I got wrong in my initial calculation was not factoring in the freedom of not having to worry about a surge of usage (whether from free trial accounts, or broken code). With a continuous running VM, your pricing is capped by:</p>
<p><code>amount of RAM provisioned * number of instances * amount of CPU provisioned</code></p>
<p>whereas with AWS Lambda, your pricing is uncapped:</p>
<p><code>amount of RAM provisioned * number of milliseconds your function takes * number of invocations</code></p>
<p>With AWS Lambda, if you get a sudden rush of curious trial users from a viral blog post, or you've deployed a function that infinitely calls itself, suddenly you're on the hook for a huge bill (entirely bound by how lucky/unlucky you are).</p>
<p>When the same thing happens to a continuously running VM processing jobs in a queue, the wait times increase. If you screwed up the programming, maybe your VM runs out of memory, and reboots. The UX degrades, your monitoring picks it up and lets you know, and you scale up the VM to handle the load (for a couple of bucks a month more).</p>
<h2>What I kept on AWS Lambda</h2>
<p>I'm keeping several full replicas of my stack still sitting on AWS Lambda (in a few AWS regions), ready to turn back on if a certain number of jobs don't run within a given timeframe.</p>
<p>I also still run Google Chrome/Puppeteer for <a href="https://onlineornot.com/browser-checks">Browser Checks</a> on AWS Lambda, but now I'm invoking them from my fly.io VM.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Guidelines for writing better runbooks]]></title>
            <link>https://onlineornot.com/guidelines-for-writing-better-runbooks</link>
            <guid>https://onlineornot.com/guidelines-for-writing-better-runbooks</guid>
            <pubDate>Tue, 14 Jun 2022 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>It's 2am.</p>
<p>You're getting paged for the latest service your team built and deployed. You have no idea how to debug it, and it's the first time you've been paged for the service.</p>
<p>"No worries, I'll just go through the list of steps in the runbook, I'm sure it'll be fine... wait, runbook just says 'Ask Dave'?!"</p>
<p>You start paging the developers who built the service (especially Dave), and over a few anxious hours, you and the team manage to resolve the issue, and go back to bed.</p>
<hr>
<p>In this article, we're going to improve our runbooks <strong>today</strong>, so no one in your team needs to experience the above scenario.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#guidelines">Guidelines</a>
<ul>
<li><a href="#optimize-for-editing">Optimize for editing</a></li>
<li><a href="#encourage-newcomers-to-use-your-runbooks-and-make-edits">Encourage newcomers to use your runbooks and make edits</a></li>
<li><a href="#keep-track-of-how-often-you-need-to-use-your-runbooks">Keep track of how often you need to use your runbooks</a></li>
<li><a href="#explicitly-tell-the-reader-what-to-do">Explicitly tell the reader what to do</a></li>
<li><a href="#make-your-runbook-easier-to-search">Make your runbook easier to search</a></li>
</ul>
</li>
</ul>
<h2>Guidelines</h2>
<p>These aren't in any particular order, and come from a few years experience working at startups, scale-ups, and enterprise.</p>
<h3>Optimize for editing</h3>
<p>The <a href="/incident-management/incident-response/writing-your-first-runbooks">first runbook your team writes</a> is probably going to suck, but that's okay!</p>
<p>Getting started is the first step towards being good at something.</p>
<p>You're going to want to put your runbooks somewhere <strong>easy to edit</strong>. While a git repo is fine, something like Notion, Confluence or Google Docs is going to be easier to update.</p>
<p>Note that an internal wiki/repo probably isn't the best idea - particularly if your team needs VPN/office network access to read the runbooks, and your VPN goes down.</p>
<h3>Encourage newcomers to use your runbooks and make edits</h3>
<p>As the authors of your runbooks, you're affected by what's known as <a href="https://en.wikipedia.org/wiki/Curse_of_knowledge">the curse of knowledge</a> - essentially, you assume people will have the background knowledge to understand what to do.</p>
<p>Newcomers to your team aren't yet affected by this, and can help you uncover assumed knowledge in your runbooks. Have them pair with your on-call team members as they resolve incidents - they'll get more comfortable with on-call before their first roster, and will notice undocumented gaps that your team <em>just knows</em>.</p>
<h3>Keep track of how often you need to use your runbooks</h3>
<p>As I mentioned earlier, while the initial goal is to document the manual process required to resolve an incident, a later goal is to automate those steps.</p>
<p>The more often you use a runbook, the more evidence you gather that perhaps parts of the runbook should be automated.</p>
<p>A simple table at the bottom of the runbook would help with this:</p>
<p>| Date last used | Used by |
| -------------- | ------- |
| 2022/01/01     | rozenmd |
| ...            | ...     |</p>
<h3>Explicitly tell the reader what to do</h3>
<p>People shouldn't need to interpret what the steps in your runbook <em>could</em> mean. Write as though the person reading it just woke up at 2am, and just wants to go back to bed.</p>
<p>The steps should <strong>explicitly tell them what to do</strong>.</p>
<p>For example, instead of writing:</p>
<ol>
<li>Investigate issues in Splunk</li>
</ol>
<p>a better runbook step would be:</p>
<ol>
<li>Login to the Splunk dashboard at URL
<ul>
<li>Run this query: "QUERY GOES HERE"</li>
<li>In case of weird results, run this command: "COMMAND HERE"
<ul>
<li>Weird results look like this: "SCREENSHOT GOES HERE"</li>
</ul>
</li>
</ul>
</li>
</ol>
<h3>Make your runbook easier to search</h3>
<p>It helps to add key phrases from your alerts, or even the exact error message thrown by your system.</p>
<p>For example:</p>
<pre><code class="language-md"># How to fix: "Error: Query defined in resolvers, but not in schema"

1. If you see this error in environment X, you need to run this command:
   ...
</code></pre>
<p>It'll make it easier to find a solution to the exact problem the system is having (by including the searched phrase in your heading, you'll also game certain <a href="https://docsearch.algolia.com/docs/tips/#structure-the-hierarchy-of-information">documentation search engines</a> to rank the terms higher).</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Writing your first runbooks]]></title>
            <link>https://onlineornot.com/writing-your-first-runbooks</link>
            <guid>https://onlineornot.com/writing-your-first-runbooks</guid>
            <pubDate>Tue, 14 Jun 2022 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Incidents are a stressful time for your team: your service isn't working the way you expect and your customers want to know what's going on. The last thing you want to do is let your team improvise everything when it comes to responding to incidents.</p>
<p>Google's own <a href="https://sre.google/sre-book/managing-incidents/">SRE book</a> has great overall tips for incident management, part of which involves "develop(ing) and document(ing) your incident management procedures in advance", which this article dives into.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-is-a-runbook">What is a runbook</a></li>
<li><a href="#when-to-write-your-first-runbook">When to write your first runbook</a></li>
<li><a href="#how-to-write-a-runbook">How to write a runbook</a></li>
<li><a href="#do-not-allow-perfectionism-to-block-you-from-releasing-your-runbooks">Do not allow perfectionism to block you from releasing your runbooks</a></li>
<li><a href="#use-new-team-members-as-a-starting-point">Use new team members as a starting point</a></li>
<li><a href="#do-not-worry-about-automation-yet">Do not worry about automation (yet)</a></li>
</ul>
<h2>What is a runbook</h2>
<p>Depending on where you work, you'll hear different words describing relatively similar concepts around runbooks:</p>
<ul>
<li><strong>Standard Operating Procedure</strong> (or <strong>SOP</strong>): a standard set of steps required by the business to perform a process or task</li>
<li><strong>Runbook</strong>: a set of steps required to <strong>respond to an incident or alert</strong></li>
<li><strong>Playbook</strong>: what we call the overall set of steps when responding to an incident. Can involve several runbooks</li>
</ul>
<p>You shouldn't worry <em>too</em> much about terminology, just know we're going to be writing instructions on how to perform the manual process of resolving an incident with a system your team runs.</p>
<h2>When to write your first runbook</h2>
<p>Ideally, you would write a runbook as you're deploying a service for the first time. You would go through the manual process you take to resolve any issues that come up, and write down each step, in detail, so anyone on your team can follow it while half-asleep.</p>
<p>Realistically though, chances are you'll be writing your first runbook after something goes wrong. For a lot of organisations, it takes an incident to really examine the way you work and learn where the potential improvements are.</p>
<p>Most commonly, something will go wrong, it'll take your on-call engineer a few minutes to figure out what happened, page the team members that know what to do, and it'll take the team a few hours to resolve the issue.</p>
<p>During the post-incident review, someone will ask "why did it take so long to fix?", and one of the inevitable answers will be "we didn't have a runbook in place".</p>
<h2>How to write a runbook</h2>
<p>Honestly, <strong>anything</strong> is better than nothing. As time goes on, and your team gains operational experience, your runbooks will improve as gaps are found.</p>
<p>To start, you want to be documenting the manual steps necessary to get your service working again. You want to tell folks exactly what to do:</p>
<ol>
<li>Login to the Splunk dashboard at URL
<ul>
<li>Run this query: "QUERY GOES HERE"</li>
<li>In case of weird results, run this command: "COMMAND HERE"
<ul>
<li>Weird results look like this: "SCREENSHOT GOES HERE"</li>
</ul>
</li>
</ul>
</li>
</ol>
<p>In other words, favour explicit steps over implicit steps.</p>
<p>If any of your runbook steps say "Do the usual troubleshooting procedure", you need to rewrite that step, following the advice above.</p>
<p>If any step of your runbook contains "If you need to do X, follow the procedure here...", you need to explain the decision making process that leads the reader towards figuring out if they need to do X.</p>
<h2>Do not allow perfectionism to block you from releasing your runbooks</h2>
<p>Depending on your organisation, you may have stakeholders that'll want every aspect of the system documented in the runbook before it's "ready".</p>
<p>Without feedback on using the runbook (such as during real or simulated incidents), you'll quickly start getting diminishing returns on the time you spend trying to perfect your runbooks.</p>
<p>You're better off spending that time resolving issues that cause common alerts to fire for your on-call engineers.</p>
<h2>Use new team members as a starting point</h2>
<p>New team members are a gift for getting started, and improving your runbooks. They aren't yet affected by the <a href="https://en.wikipedia.org/wiki/Curse_of_knowledge">curse of knowledge</a>, and tend to be curious about how things work.</p>
<p>Have them sit in on incidents, join incident response chat rooms, and have them write down any questions they have as the incident is worked on. The questions they come up with are perfect starting points for a runbook.</p>
<h2>Do not worry about automation (yet)</h2>
<p>Automating your runbooks comes later (chances are, it won't cover every single step of your runbook anyway).</p>
<p>To start with though, you're going to want to document the <strong>entire</strong> manual procedure before you begin to automate the easiest steps.</p>
<p>In short: write today, automate tomorrow.</p>
<p>If you're early in your incident response journey, your team isn't ready for automation. Your team likely has too much <a href="https://en.wikipedia.org/wiki/Tribal_knowledge">tribal knowledge</a>, and needs to work on documenting their manual processes first.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[What I learned running a SaaS for a year]]></title>
            <link>https://onlineornot.com/what-learned-running-saas-for-year</link>
            <guid>https://onlineornot.com/what-learned-running-saas-for-year</guid>
            <pubDate>Mon, 21 Feb 2022 06:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>This time last year, I showed the internet a little prototype uptime checker I built using Next.js as the frontend, with services running on AWS Lambda. I gave myself <a href="/building-saas-in-one-week-how-built-onlineornot">one week</a> to put it together.</p>
<p><strong>Update:</strong> I've been running this business for over three years now, if you're interested to see how it's going, check out <a href="https://maxrozen.com/lessons-from-my-third-year-running-a-saas">lessons from my third year running a SaaS
</a>.</p>
<p>The gist of my approach is as follows:</p>
<p>I started with a single Lambda function that checks if static websites were still online, added an email alert if it's offline, wrapped authentication around it, integrated Stripe, and shipped it. Throughout the year, I kept adding features.</p>
<p>My trick for launching into 200 competitors providing the "same" service and still getting customers?</p>
<ul>
<li>
<p>I work two hours a day, every weekday on OnlineOrNot, and no other side projects. I managed to keep this going for about 10 months this year (with a cheeky 2 month burnout in the middle).</p>
</li>
<li>
<p>I focus particularly on features that solve customer pain (and I ask my customers what that pain is)</p>
</li>
<li>
<p>I'm ruthlessly iterative. If I can't get a feature done in two hours, I figure out how to cut scope down to a two hour block, and ship that. Then iterate on it.</p>
</li>
</ul>
<p>Here's what I learned after a year:</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#youre-solving-a-problem-not-selling-a-saas-subscription">You're solving a problem, not selling a SaaS subscription</a></li>
<li><a href="#docs-are-part-of-your-user-experience">Docs are part of your user experience</a></li>
<li><a href="#build-for-mobile">Build for Mobile</a></li>
<li><a href="#ask-people-how-they-found-you">Ask people how they found you</a></li>
<li><a href="#use-analytics-set-up-funnel-tracking">Use analytics, set-up funnel tracking</a></li>
<li><a href="#sometimes-you-need-to-make-your-own-mistakes">Sometimes you need to make your own mistakes</a></li>
<li><a href="#pricing-is-bloody-hard-to-get-right">Pricing is bloody hard to get right</a></li>
<li><a href="#you-probably-focus-too-much-on-mrr">You probably focus too much on MRR</a></li>
<li><a href="#you-still-need-a-free-trial-for-paid-tiers">You still need a free trial for paid tiers</a></li>
<li><a href="#its-hard-to-bring-in-more-traffic-easy-to-change-what-your-current-traffic-does">It's hard to bring in more traffic, easy to change what your current traffic does</a></li>
<li><a href="#content-marketing-buys-you-time">Content marketing buys you time</a></li>
<li><a href="#ship-small-ship-often">Ship small, ship often</a></li>
<li><a href="#ship-first-worry-about-scale-later">Ship first, worry about scale later</a></li>
<li><a href="#you-dont-get-to-spend-as-much-time-working-on-the-problem-as-youd-think">You don't get to spend as much time working on the problem as you'd think</a></li>
</ul>
<h2>You're solving a problem, not selling a SaaS subscription</h2>
<p>When building a product, it helps to think from the perspective of your customers, rather than your own goal of selling subscriptions.</p>
<p>This takes you from a mindset of "I'm just going to keep building features, they'll all come eventually!" to "I should be helping my users solve this annoying problem".</p>
<p>SaaS is just one of many ways to solve a problem. You should be looking at all the ways you could be helping, whether that's screencasts, docs, articles, books, workshops, code samples or software.</p>
<h2>Docs are part of your user experience</h2>
<p>People often say "Developers don't read documentation" - I think that's only partially true.</p>
<p>They don't read, they skim the headings.</p>
<p>In my experience, some people would try to figure out how to do things themselves in OnlineOrNot's UI, get frustrated, check the docs, and either one of two things would happen:</p>
<ol>
<li>They wouldn't find a way to do what they were trying to do, and they'd immediately churn</li>
<li>They'd find the page they were looking for, jump to the heading, and do the thing</li>
</ol>
<p>Successful use of OnlineOrNot's docs drove retention - so I treat it as part of the core product now, not as an afterthought.</p>
<h2>Build for Mobile</h2>
<p>Contrary to popular belief (for B2B SaaS), people actually work from their phones.</p>
<p>Something like 50% of traffic to OnlineOrNot.com comes from folks on mobile. They tend to quickly create an account, add a few pages to monitor, then eventually get on their laptop/desktop to review their checks from time to time.</p>
<p>For about 6 months I didn't support mobile well, and folks that signed up on their phone churned rapidly. I eventually took the time to build responsive views for mobile, and now new mobile users are sticking around.</p>
<h2>Ask people how they found you</h2>
<p>One of the most valuable code changes I made this year was asking people as they signed up: "How did you find out about OnlineOrNot?"</p>
<p>You need to know where your users are finding you.</p>
<p>There are dozens of channels you could be using to attract potential customers, and it's useful to know whether to double-down on paid ads, content marketing, or shitposting on Twitter to bring in users.</p>
<h2>Use analytics, set-up funnel tracking</h2>
<p>Your marketing funnel helps you work out the health of your business. While it's nice to see how many people visit individual pages, it's even better to see how people flow across multiple pages.</p>
<p>When I say marketing funnel, I mean tracking the flow of people that visit your home page, eventually making it to your sign-up form, and finally making it inside the actual product. It helps you diagnose problems with your marketing copy, the sign-up form itself, as well as any onboarding you may have before the user actually sees your app.</p>
<p>You don't need gross privacy-invasive tracking to achieve this by the way - I use a basic privacy-first web analytics service (and have no cookie banner as a result), and it works fine.</p>
<h2>Sometimes you need to make your own mistakes</h2>
<p>I read quite a few business books, mainly from not wanting to repeat mistakes that others have made.</p>
<p>Sometimes though, you need to make mistakes for yourself.</p>
<p>As an example - it took me getting on the front page of Hacker News, having 6000 people visit my landing page, a few hundred people attempt to sign-up, and only single-digits making it through to the app for me to realise something might be wrong.</p>
<p>I had something like a 75% drop-off rate on my sign-up form alone. With a bit of A/B testing, I got it down to 50% just by adding an extra OAuth login provider.</p>
<h2>Pricing is bloody hard to get right</h2>
<p>Price too high, and you'll have churn from folks who expect your app does everything. Too low, and you'll have customers that demand you rewrite your app just because they gave you $9. Refund the difficult customers, raise your prices, and move on.</p>
<p>Be prepared to experiment a lot with pricing.</p>
<h2>You probably focus too much on MRR</h2>
<p>Tracking your MRR is a pretty lousy way to measure how you're doing as a business early on.</p>
<p>Things you did weeks (if not months) ago will affect your MRR today, so you won't really know if pricing changes work until you've already got a decent number of customers going through different stages of their customer journey.</p>
<p>I find measuring your daily active users, or some sort of "success metric" for your customers (like pages checked, images generated, etc) is more helpful than MRR. It lets you figure out if people are actually using your product, and whether it's bringing them value.</p>
<h2>You still need a free trial for paid tiers</h2>
<p>While a free tier is a great way to attract people and get them talking about your product, you still need a way to let them sample "the good stuff", especially if the free tier is significantly less useful than your paid tiers.</p>
<p>It took me 11 months to realise I should probably build an onboarding flow, and start offering free trials. 95% of new users pick a free trial of the pro tier, even though I offer a free tier.</p>
<h2>It's hard to bring in more traffic, easy to change what your current traffic does</h2>
<p>Getting noticed on the internet is a long, slow game.</p>
<p>Eventually over months (if not years), if you're consistent at quality content marketing, the number of readers on your articles will grow from 1-2 a day, to a few hundred per day.</p>
<p>Increasing the number of people landing on your site isn't particularly easy.</p>
<p>On the other hand, what people do once they land on your site is entirely within your influence, and something you can change today (such as adding an additional OAuth login provider to your sign-up form, that I mentioned earlier).</p>
<h2>Content marketing buys you time</h2>
<p>Investing in content marketing gives you the option of letting the business run itself/coast for a while.</p>
<p>Throughout the year, I'd have an occasional article go viral and bring in tens of thousands of visitors over a month, if I did absolutely nothing, about 1500 people would still visit the site organically for the articles I had written.</p>
<p><img src="/assets/what-learned-running-saas-for-year/onlineornot-traffic.png" alt="OnlineOrNot&#x27;s traffic"></p>
<p>This was particularly useful for when it's a pandemic and I was having a bit of a burnout while moving across the world to live in France.</p>
<h2>Ship small, ship often</h2>
<p>People will suggest you should build particular features to improve your product.</p>
<p>They'll probably never use those features.</p>
<p>They're probably just trying to be helpful, and saw a similar feature in another product. Because you're new to running a SaaS, you'll be excited that people are actually talking to you, and rush out to build that feature for them.</p>
<p>I'm not going to tell you not to build the feature (that's the advice I was given, and I built unused features anyway). You should ask how they would use the feature, ask other customers how they deal with the problem, build the smallest possible version of that feature, and see how the rest of your customers use it. You don't want to be building snowflake features only one person uses.</p>
<p>It stings a lot less to remove a feature no one wanted after spending a few hours on it, rather than a few months.</p>
<h2>Ship first, worry about scale later</h2>
<p>In the first iteration of OnlineOrNot, I didn't optimise the architecture at all.</p>
<p>Each uptime check would manage its own database connection, meaning as more users found the service, the harder it would be for additional users to use the app. I also didn't bother making decent error states, so new users would see this while the database was busy:</p>
<p><img src="/assets/what-learned-running-saas-for-year/old-error.jpeg" alt="Old OnlineOrNot Error Screen"></p>
<p>Not a great look.</p>
<p>At the same time, I prefer being embarrassed by incomplete UI than building things people don't need. There was never a guarantee that OnlineOrNot would attract thousands of users, it could have ended up as another SaaS I built only for myself.</p>
<p>I ended up reworking the architecture to handle millions of checks per week on the smallest AWS RDS instance, as well as cleaning up that error screen:</p>
<p><img src="/assets/what-learned-running-saas-for-year/new-error.jpeg" alt="New OnlineOrNot Error Screen"></p>
<h2>You don't get to spend as much time working on the problem as you'd think</h2>
<p>Of the time I spend programming this year, I'd say half of my time went to actually solving the problem I wanted to solve (knowing if a site is down, and alerting folks when that happens). The other half went to building a SaaS platform around that problem.</p>
<p>SaaS platform things you didn't even realise you'd need, like multiple types of authentication and user management, trials, onboarding, team management and invoice management, lifecycle emails, and more.</p>
<p>You can outsource a lot of this (and I do! If Stripe didn't exist, I probably wouldn't be selling a service nor using subscription-based billing), but there's always stuff you don't feel comfortable outsourcing, or that you handle differently, so you need to build it yourself.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Communicating to Users During Incidents]]></title>
            <link>https://onlineornot.com/communicating-users-incidents</link>
            <guid>https://onlineornot.com/communicating-users-incidents</guid>
            <pubDate>Fri, 14 Jan 2022 06:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Imagine you're having a regular day at work, opening up your browser, double checking something for a client in that web app your team built for them, when suddenly, you see this screen:</p>
<p><img src="/assets/communicating-users-incidents/500-error-example.png" alt="500 Error example"></p>
<p>You hit refresh a few times, just to be sure.</p>
<p>Nope. Still down.</p>
<p>What happens next depends on how well your team has planned for incidents like this (some folks call it unplanned downtime).</p>
<h2>Incident Response and Communication</h2>
<p>Incidents can be a frustrating time for everyone.</p>
<p>For you, as someone on the operational side of the app, it's frustrating because you're trying to fix the problem as fast as possible, with as little distraction as possible.</p>
<p>On the user side of things, their work day is interrupted, and without good communication, they're not sure if someone is currently working on the issue, or when it'll be fixed.</p>
<h2>The Incident Communication Scale</h2>
<p>You can think of communicating with your users during an incident as a scale.</p>
<p>On the left, there's "ignore your users", or under-communicate. The vast majority of organisations are on this side of the scale.</p>
<p>On the right, there's "send each of your users an SMS every few minutes to let them know how it's going", or over-communicate.</p>
<p>Both of these extremes are annoying - the sweet spot is somewhere in the middle.</p>
<p>While it's rare for a company to intentionally ignore their users during an incident, even waiting to confirm exactly what's happening before acknowledging an issue can be extremely frustrating for users.</p>
<p>A famous example of under-communication is AWS's <a href="https://status.aws.amazon.com/">Service Health Dashboard</a>. AWS regularly takes over an hour to even acknowledge <em>anything</em> is wrong, let alone describe the issue and whether it's being worked on.</p>
<h2>Reducing Frustration for Everyone</h2>
<p>As I mentioned earlier, chances are, your organisation is on the left hand side of the scale: under-communicating with your users during incidents.</p>
<p>To reduce frustration for everyone, you're going to want to do a few things during your incidents:</p>
<h3>As soon as you're aware of an incident, tell your users!</h3>
<p>You'll want to let your users know something is not right, and someone is investigating.</p>
<p>If you're only working with internal users within your organisation, send an SMS, email or a Slack message letting your users know something is wrong, and that someone is looking into the issue.</p>
<p>If you're working with external users, you're going to want a dashboard to display your product's uptime or status (on separate infrastructure to your app!) that users know to look when things aren't working properly.</p>
<p>You might be wondering whether or not you should notify your users about incidents before you've confirmed there's a user-facing impact. As an idealist, I'd suggest all incidents should be posted on your status dashboard, as the cost of getting it wrong is your user's trust - but certain organisations value other things more than user trust.</p>
<h3>Send regular updates</h3>
<p>When working at a previous employer, we would send an update as often as every 30 minutes, depending on how critical the system was to the business. In this case, "We're still looking into it" was a completely valid update.</p>
<p>For some audiences (particularly if you have a status dashboard), this might be too often, and you'll need to tweak how often you communicate.</p>
<p>General rule of thumb: if your users are reaching out to you (via Phone call, SMS, Tweets, etc), you've probably left it too long without an update.</p>
<h3>Have a dedicated communications person for incidents</h3>
<p>It shouldn't be up to one person to simultaneously communicate with stakeholders AND fix the issue.</p>
<p>At the very least, you'll want a dedicated person in charge of communicating updates from the team fixing the issue to the outside world. Having a separate person here can help improve how long it takes to resolve the issue, as you spend less time context switching between "Oh no, need to fix the problem! aaaahhhh!" and "We are aware of the issue with [SERVICE], and are working to restore it ASAP. We will notify you once we have any updates".</p>
<p>In some organisations this role is called "Communications Engineer" or "Communications Manager" - either way, the person doesn't have to be technical, they just need to be capable of talking to engineers and stakeholders.</p>
<h2>Summary</h2>
<p>In short, once you're aware of an incident, tell your users!</p>
<p>Once you've told your users, send regular updates.</p>
<p>Finally, having a separate person in charge of communication, between the team fixing the issue and the rest of the business, can help reduce your Mean Time to Resolution by reducing context switching.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[What we learned from AWS's us-east-1 outage]]></title>
            <link>https://onlineornot.com/onlineornot-aws-outage-retrospective</link>
            <guid>https://onlineornot.com/onlineornot-aws-outage-retrospective</guid>
            <pubDate>Wed, 08 Dec 2021 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>In case you missed it, for several hours on December 7, 2021, AWS's us-east-1 region had an outage impacting multiple AWS APIs, taking out various websites across the internet.</p>
<p>According to our own monitoring at <a href="https://onlineornot.com/">OnlineOrNot</a>, the outage started at 2021-12-07 15:32 UTC and began to recover well at 2021-12-07 22:48 UTC (with minor signs of life for a few minutes around 2021-12-07 20:08 UTC).</p>
<p>Had we relied solely on AWS to update their status page before reacting, we would have been waiting a while. In fact, AWS took approximately an hour to update their status page to indicate <strong>anything</strong> was happening at all.</p>
<p>After things got serious, they added a banner to the top of their status page, explaining what was happening:</p>
<p><img src="/assets/onlineornot-aws-outage-retrospective/aws-status-page.png" alt="AWS Status Page"></p>
<h2>The retrospective</h2>
<h3>Some context</h3>
<p>OnlineOrNot is an uptime checker service provides uptime and correctness checks for websites, web apps, and APIs. While our infrastructure is hosted across 20 AWS regions, a single point of failure was hosted in us-east-1.</p>
<h3>What went wrong</h3>
<p>OnlineOrNot's uptime check service stopped running regularly when us-east-1 started having issues.</p>
<p>Every 60 seconds, our service is kicked off via an AWS Cloudwatch event rule hosted in us-east-1.</p>
<p>When that event fires, a single AWS Lambda in us-east-1 queries our database to determine which URLs need checking, and sends off messages via SQS queues to trigger our checkers in 20 AWS regions.</p>
<h3>What went well</h3>
<p>Mostly by luck, the database was unaffected and our web app remained online. This let us quickly deploy a banner letting our users know something was wrong:</p>
<p><img src="/assets/onlineornot-aws-outage-retrospective/onlineornot-status-banner.png" alt="OnlineOrNot displaying its warning banner"></p>
<h3>What we learned</h3>
<p>Since the outage we've implemented redundancy and the ability to failover to other AWS regions in our uptime check service.</p>
<p>On top of this, we'll be looking at duplicating more of our stack in other cloud providers to ensure this doesn't happen to OnlineOrNot again on AWS, nor any other cloud service.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Dealing with Noisy Error Monitoring]]></title>
            <link>https://onlineornot.com/dealing-with-noisy-error-monitoring</link>
            <guid>https://onlineornot.com/dealing-with-noisy-error-monitoring</guid>
            <pubDate>Wed, 01 Dec 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Say you've been tasked with monitoring an application, so you set up some alerts to let you know when errors are coming in.</p>
<p>The minutes roll by, the errors start coming...</p>
<p>...and they don't stop coming...</p>
<p>Oh my, there seems to be quite a few errors coming through. Alerting on each error isn't going to help, better report on changes in the error rate instead right?</p>
<p>Not quite.</p>
<p>While there's no shortage of vendors that'll sell you on the benefits of error rate alerting, you need to get back to basics first.</p>
<h2>Instrument your application</h2>
<p>If you haven't already, you'll want to instrument your individual API endpoints.</p>
<p>For each endpoint, you'll want to be monitoring the requests, errors, and how long each call takes. Doing so makes it easier to find regressions in each endpoint when things <em>do</em> go wrong.</p>
<p>There are a few services out there for instrumenting your app, it's worth looking at either <a href="https://sentry.io/welcome/">Sentry</a> or <a href="https://honeycomb.io/">Honeycomb</a>.</p>
<p>For example, tracking down the cause of an increase in 4xx errors becomes a matter of comparing the current week's requests/errors/duration to the previous week's for each endpoint.</p>
<h2>Clean up your alerts</h2>
<p>Go through each type of error you're seeing that triggers an alert, and figure out what the appropriate response would be.</p>
<p>If the error isn't actionable, then it shouldn't create an alert. Have your developers change it to an INFO or a WARN log instead.</p>
<p>If the error IS actionable, and your application has already attempted to self-heal the issue (assuming self-healing is possible), THEN you should alert.</p>
<p>If the error IS actionable, BUT only after a certain threshold, you might be dealing with errors caused by a retry loop. Errors that occur as part of a sequence that eventually succeeds aren't real errors. If after several retry attempts the action still fails, that whole sequence of events should be counted as one error. Have your developers modify the error reporting accordingly.</p>
<p>As part of your team's on-call schedule, you should be documenting every alert that gets made, and actions taken to resolve each issue (and your alerts should provide instructions on what to do via a runbook). Common alerts should be prioritised and fixed by your developers.</p>
<h2>Summary</h2>
<p>If you're struggling with the signal to noise ratio of your company's alerting, you need to reflect on what the point of all of those alerts is.</p>
<p>An alert should mean someone needs to do something NOW.</p>
<p>It should <strong>not</strong> mean "Hey Jordan, check it out, CPU is at 69%!"</p>
<p>Another way to think of this is - if you're getting an alert for an error, it should be a priority to fix that error. If it isn't, what's the point?</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Getting over on-call anxiety]]></title>
            <link>https://onlineornot.com/getting-over-on-call-anxiety</link>
            <guid>https://onlineornot.com/getting-over-on-call-anxiety</guid>
            <pubDate>Thu, 22 Jul 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>You've joined a company, or worked there a little while, and you've just now realised that you'll have to do on-call. You feel like you don't know much about how everything fits together, how are you supposed to fix it at 2am when you get paged?</p>
<p>So you're a little nervous. Understandable. Here are a few tips to help you become less nervous:</p>
<h2>It's not your job to fix everything</h2>
<p>The first thing you need to internalise to do well on-call is: it's not your job to fix every possible thing that could break.</p>
<p>Your job is to make sure that <strong>someone</strong> fixes the issue.</p>
<p>You can <strong>always</strong> page someone else with more expertise to fix the problem (ideally, once you've investigated the alert, and can provide the next person you page with context to get them started).</p>
<h2>Use a different ringtone for your pages</h2>
<p>It's pretty common to have a loud, scary "nuclear alarm" style ringtone for your on-call pages. Or worse, using the same ringtone as the rest of your calls.</p>
<p>You're going to want to use a different ringtone than the rest of your calls - it lets you know immediately whether you're going to need to focus, or if you're about to have a chat with your Mum. On top of that, it helps you avoid associating your phone ringing with anxiety - a fair chunk of the anxiety comes from not knowing whether the call is you being paged, or just another telemarketer.</p>
<p>I'd also recommend using a gradual ringtone that gets louder and louder - this is more of a personal preference, but I find the loud alarm style ringtones create an unproductive sense of panic that gradual ringtones don't.</p>
<h2>Know that it'll get better with time</h2>
<p>Like most things, on-call (should) get easier as time goes on. Eventually, you learn the processes (and help improve the documentation as you discover gaps), who to escalate to when things get <strong>really</strong> bad, and so on.</p>
<p>While the anxiety over getting paged never completely goes away, you gain confidence that you can solve any issue that arises, making the anxiety much easier to manage.</p>
<p>Also know that your team is there to support you - particularly your manager. Chances are, they've guided your colleagues through starting on-call as well, and can help address any concerns you may have about going on-call.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Self-hosting vs Managed Services: Deciding how to host your database]]></title>
            <link>https://onlineornot.com/self-hosting-vs-managed-services-deciding-how-host-your-database</link>
            <guid>https://onlineornot.com/self-hosting-vs-managed-services-deciding-how-host-your-database</guid>
            <pubDate>Thu, 08 Jul 2021 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Like all good things in infrastructure, picking whether or not to self-host your database is full of trade-offs.</p>
<p>On the one hand, you have the absolute freedom to do whatever it is you want with your database - whether it's adding a useful Postgres extension, or experimenting with new technologies. On the other hand, you now have to dedicate resources to keeping your database reliably online.</p>
<p>This article aims to provide an unbiased dive into the benefits of each side, as well as tasks that need doing regardless of what you decide to do, to help you decide whether or not it's worth self-hosting in your circumstances.</p>
<p>Note that while I mention certain services, it's purely out of familiarity - I'm not paid for mentioning them.</p>
<p>You might even decide a traditional database isn't worth the hassle, and opt for an entirely different approach (like MongoDB or CockroachDB) - though that's for another article.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-i-mean-by-self-hosting-and-managed-services">What I mean by self-hosting and managed services</a>
<ul>
<li><a href="#self-hosting">Self-hosting</a></li>
<li><a href="#managed-services">Managed Services</a></li>
</ul>
</li>
<li><a href="#benefits-of-self-hosting">Benefits of self-hosting</a>
<ul>
<li><a href="#price">Price</a></li>
<li><a href="#portability">Portability</a></li>
<li><a href="#control">Control</a></li>
</ul>
</li>
<li><a href="#benefits-of-managed-services">Benefits of managed services</a>
<ul>
<li><a href="#focus">Focus</a></li>
<li><a href="#scalability">Scalability</a></li>
<li><a href="#support">Support</a></li>
<li><a href="#blame-as-a-service">Blame-as-a-service</a></li>
</ul>
</li>
<li><a href="#things-youll-have-to-do-regardless-of-which-option-you-pick">Things you'll have to do regardless of which option you pick</a></li>
<li><a href="#summary">Summary</a></li>
</ul>
<h2>What I mean by self-hosting and managed services</h2>
<p>Not everyone means the same thing when they say "self-hosting" and "managed services", so I just want to be clear on what this article is talking about.</p>
<h3>Self-hosting</h3>
<p>When I say self-hosting, I'm talking about running database software on VMs that you have control over. In other words, running the database on <a href="https://aws.amazon.com/ec2/">AWS EC2</a>, <a href="https://cloud.google.com/compute">Google Cloud Engine</a>, or <a href="https://azure.microsoft.com/en-au/services/virtual-machines/">Azure Virtual Machines</a> or a similar VPS provider.</p>
<p>While you could take the term further, and have "run your own physical hardware" fall under self-hosting, I have no intention of discussing that here.</p>
<h3>Managed Services</h3>
<p>When I say managed services, I'm talking about a "database-as-a-service" type platform, such as <a href="https://aws.amazon.com/rds/">AWS RDS</a>, <a href="https://cloud.google.com/sql">Cloud SQL</a>, or <a href="https://azure.microsoft.com/en-us/products/azure-sql/#product-overview">Azure Database</a>.</p>
<h2>Benefits of self-hosting</h2>
<h3>Price</h3>
<p>One of the biggest factors driving folks to self-host their database is price.</p>
<p>At the low end, the cheapest tier for AWS RDS (the AWS managed database service) is around $15 USD per month. In comparison, installing PostgreSQL or MySQL on your application server is "free". In doing so, you lose the ability to scale parts of your application independently, but for smaller projects this is fine.</p>
<p>Once your application needs a bit more CPU and RAM, and you start looking for say, a 4 vCPU 16 GB RAM server (db.m4.xlarge on AWS RDS), assuming you've got 1TB of data, you're looking at around $266 USD per month on AWS RDS. Over on AWS EC2, that configuration (t4g.xlarge) would cost you only around $98 USD per month.</p>
<h3>Portability</h3>
<p>By self-hosting, you also minimise vendor lock-in. Say one day you notice that your 4 vCPU/16GB RAM/1TB storage configuration costs only $80 USD per month at VPS hosts like Hetzner, moving your self-hosted setup is as simple as running your setup scripts again and restoring from backup.</p>
<p>Of course, using a managed database service doesn't prevent you from using another service, it's just a lot easier when your database setup can be quickly spun up on any commodity Linux VM.</p>
<h3>Control</h3>
<p>Pick your own adventure! By self-hosting, you get to decide:</p>
<ul>
<li>Which OS to run (and ensuring your OS is configured correctly, so your database comes back up after a restart)</li>
<li>How often to update/patch your OS</li>
<li>Which other software runs on the same server as your database</li>
<li>Which database extensions to run (for example, TimescaleDB took years to be <a href="https://github.com/timescale/timescaledb/issues/65">supported by major managed service providers</a>)</li>
<li>How often to upgrade/update/patch your database software</li>
<li>How often to run backups</li>
<li>Which disk configuration to use (to get more performance), as well as how to log (ensuring your logs go to the right place, so they don't fill up your data directory)</li>
<li>and more!</li>
</ul>
<h2>Benefits of managed services</h2>
<h3>Focus</h3>
<p>With managed services, you pick the version of the database you want to run, your desired instance size, click "Create", and that's it.</p>
<p>Backups, minor updates, maintenance tasks and more are all run automatically, leaving you to focus on building your application. You <em>do</em> of course pay for the convenience of having these tasks automated, but it frees up your engineers to work on features of your product, rather than keeping the lights on.</p>
<p>If your core business isn't running a database, why are you wasting your focus on it, when it can be outsourced?</p>
<h3>Scalability</h3>
<p>One of the biggest benefits of managed services is the ease of scaling your database up or down at a moment's notice.</p>
<p>With the press of a button (and a couple of minutes of downtime, depending on how you've set up your database), you can go from the smallest database tier to whichever size you need for your workload, and back down again once you understand the resource requirements of your workload.</p>
<p>You also only pay for the resources you've used - so if you decide you need a significantly beefier machine for 12 hours, you only pay for those 12 hours.</p>
<h3>Support</h3>
<p>When you decide to self-host, the best free support you're likely to get is on mailing lists and forums. Whereas when you use a managed service, part of the cost covers basic access to dedicated support staff that specialise in your database.</p>
<p>While they won't be able to give you free bespoke advice on how to architect your application, they <em>can</em> assist with root cause analysis when incidents happen, and provide general "this is how most people use our databases" advice.</p>
<h3>Blame-as-a-service</h3>
<p>Especially in larger organisations, being able to assign blame to the vendor when things go wrong can save a lot of stress.</p>
<p>In some organisations this doesn't give you a "free pass" for the blame - after all if you decide to outsource to a managed service and it goes down, since you made the decision, you still get the blame for making the decision. At the same time, plausible deniability via "nobody got fired for buying IBM" is still a thing: if you picked a reputable party, how were you to know that they'd screw up?</p>
<p>(While I personally wouldn't use a managed service provider as a scapegoat, there are those that would, so it is worth mentioning.)</p>
<h2>Things you'll have to do regardless of which option you pick</h2>
<p>Whether you decide to self-host your database, or use a managed service, you'll still need to:</p>
<ul>
<li>Secure access to your database</li>
<li>Monitor basic metrics like CPU usage, RAM usage, disk usage, etc as well as slow queries</li>
<li>Ensure your backups are running, and tested for recovery (as well as off-site backups if you want to be particularly strict about your data)</li>
<li>Major updates
<ul>
<li>While most managed services will perform minor updates for you during your scheduled maintenance window, you still need to handle major updates yourself</li>
</ul>
</li>
</ul>
<h2>Summary</h2>
<p>Are you a startup, or a business where you'd prefer all of your engineers to be working on features, rather than keeping a database operational? Comfortable with paying extra to make the problem go away?</p>
<p>Then managed database services such as AWS RDS might be right for you.</p>
<p>Are you an established business with engineers working at a steady pace (i.e. not moving fast and breaking things)? Looking to build something cost efficient in the long-term, that's optimised specifically for the needs of your application?</p>
<p>Then self-hosting with your own employees managing and operating the database may be a better choice.</p>
<p>If you're a large enterprise with a stable product that uses managed services, it might be worth reflecting on whether it's worth bringing your database in-house by self-hosting, as the cost savings can be significant.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[What the Fastly outage can teach us about writing error messages]]></title>
            <link>https://onlineornot.com/what-fastly-outage-can-teach-about-writing-error-messages</link>
            <guid>https://onlineornot.com/what-fastly-outage-can-teach-about-writing-error-messages</guid>
            <pubDate>Wed, 09 Jun 2021 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>In case you missed it, for about 15 minutes on June 8, 2021, Fastly's CDN had an outage, taking some of the internet's largest websites down (including the BBC, UK government, Reddit, and the New York Times - Amazon.com also had its CSS fail to load).</p>
<p>If you happened to visit those websites during Fastly's outage, you saw the relatively unhelpful error message below:</p>
<p><img src="/assets/what-fastly-outage-can-teach-about-writing-error-messages/awkward-error-message.png" alt="Unhelpful Fastly Error Message"></p>
<p>As a frontend developer, my eyes scan error messages like these for numbers - in this case, the "503" - indicating that the error isn't my fault, and I can move on with my life.</p>
<p>Unfortunately the majority of internet users aren't trained in the art of reading <a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Status">HTTP status codes</a>, so this error message wasn't particularly useful to them. Particularly when a solid portion of the error message was an <a href="https://en.wikipedia.org/wiki/Guru_Meditation">in-joke</a>.</p>
<h2>We <em>can</em> write better error messages</h2>
<p>The majority of internet users aren't developers, so just writing the error code and its name (503 Service Unavailable) just isn't good enough.</p>
<p>The <a href="https://www.nngroup.com/articles/improving-dreaded-404-error-message/">Nielsen Norman Group</a> (back in 1998!) provided us with some basic guiding principles for writing better error messages:</p>
<ul>
<li>Write in plain English (or whichever language you're supporting)</li>
<li>Tell the user exactly what went wrong</li>
<li>Tell the user how the problem can be fixed</li>
</ul>
<p>More concretely, we can write better error messages by answering the following four questions:</p>
<ol>
<li><strong>Who</strong> caused the error?</li>
<li><strong>What</strong> happened, and <strong>why</strong>?</li>
<li><strong>When</strong> will it be fixed?</li>
<li><strong>How</strong> can the user respond to the error?</li>
</ol>
<p>If your error message covers those four points, <em>then</em> you can think about adding humour and some brand identity.</p>
<h3>Who caused the error?</h3>
<p>The last thing you want to do is make your users feel dumb, or as though they're at fault for an issue with the service. Communicating who caused the error helps clear up any confusion.</p>
<p>Explicitly focus on "we" when the error is caused by an issue on your end (typically HTTP status codes in the 5xx range).</p>
<p>An error message such as</p>
<blockquote>
<p>Our service is down for maintenance</p>
</blockquote>
<p>is infinitely better than:</p>
<blockquote>
<p>Uh-oh!</p>
</blockquote>
<p>For errors caused by the user (typically HTTP status codes in the 4xx range), be explicit about that too. For example, a 403 Forbidden error (where you know the user isn't authorized to view content) could be communicated as:</p>
<blockquote>
<p>Access Denied. You do not have permission to view this page.</p>
</blockquote>
<h3>What happened, and why?</h3>
<p>While users may not be technical, they still need an explanation of why they're seeing your error screen.</p>
<p>Take the classic 404 error message: "404 Not Found". You can make it significantly better for non-technical users with a single word:</p>
<blockquote>
<p>Page not found.</p>
</blockquote>
<p>adding a "why", makes it even better, giving them a way to fix the issue:</p>
<blockquote>
<p>Page not found. You might have mistyped the URL.</p>
</blockquote>
<h3>When will it be fixed?</h3>
<p>It's relatively difficult to keep an error message updated with details of your outage, and when you expect the service to become available again.</p>
<p>A better approach would be to link to either your status page, or Twitter account, or both, as in GitHub's case:</p>
<p><img src="/assets/what-fastly-outage-can-teach-about-writing-error-messages/github-down-example.png" alt="An example of GitHub&#x27;s service error page"></p>
<h3>How can the user respond to the error?</h3>
<p>In the case of a 404, you might want to list some steps the user can take to fix the issue, such as:</p>
<ul>
<li>Going back to the home page</li>
<li>Using your search bar to find the page if it's been moved</li>
<li>Contacting support</li>
</ul>
<p>Whereas in the case of a 5xx server error, you want to communicate to the user that there isn't much they can do, and that it's not their fault.</p>
<p>My favourite example of a company doing this well is Airbnb:</p>
<p><img src="/assets/what-fastly-outage-can-teach-about-writing-error-messages/airbnb-down-example.png" alt="An example of Airbnb&#x27;s service error page"></p>
<p>They tell users:</p>
<ul>
<li>there's definitely an issue, and they're working on it,</li>
<li>to check out their Twitter account for updates,</li>
<li>a way to get support for urgent issues,</li>
<li>and they set the expectation that they may be slow to respond while the site is experiencing downtime</li>
</ul>
<h2>Summary</h2>
<p>We as developers should take the Fastly outage as an opportunity to improve our error messages. You never know when your witty in-joke might end up being seen by hundreds of millions of internet users, so just try to be helpful, and explain:</p>
<ol>
<li><strong>Who</strong> caused the error?</li>
<li><strong>What</strong> happened, and <strong>why</strong>?</li>
<li><strong>When</strong> will it be fixed?</li>
<li><strong>How</strong> can the user respond to the error?</li>
</ol>
<p>Of course, if you're running a CDN like Fastly is, you'll likely still need request IDs and other diagnostics in your error message to help your support staff debug the issue. A human-readable error message doesn't have to come at the expense of removing <em>all</em> technical information.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Picking a domain name, and which country to host your website from]]></title>
            <link>https://onlineornot.com/picking-domain-name-which-country-to-host-website-from</link>
            <guid>https://onlineornot.com/picking-domain-name-which-country-to-host-website-from</guid>
            <pubDate>Mon, 24 May 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>No matter where you live, if your business targets a global audience, the question of where to host your website comes up pretty often.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#picking-a-domain-name">Picking a domain name</a>
<ul>
<li><a href="#side-note-when-to-get-a-country-domain-name-and-when-to-get-a-com">Side note: When to get a country domain name, and when to get a .com</a></li>
</ul>
</li>
<li><a href="#hosting-a-country-domain-name-anywhere-in-the-world">Hosting a country domain name anywhere in the world</a></li>
<li><a href="#best-practice-use-a-cdn">Best Practice: Use a CDN!</a></li>
</ul>
<h2>Picking a domain name</h2>
<p>A common point of confusion here is your domain name.</p>
<p>I often hear people asking, "If I have a .com.au, do I have to keep my server in Australia?"</p>
<p>In short, the answer is no - It's very common to purchase domains from across the world, whether that's .co.nz (New Zealand), .es (Spain) or .kz (Kazakhstan), and host your content in a different country.</p>
<h3>Side note: When to get a country domain name, and when to get a .com</h3>
<p>If your business is local (think locksmith, plumber, gardener etc) - it might make sense to go for a country-specific domain name.</p>
<p>On the other hand, if you have a global audience go for a .com - for example: <a href="https://onlineornot.com">OnlineOrNot.com</a> (this site!) while based in Australia with me in Sydney, our customers are from across the world, so a country-specific domain name wouldn't make sense.</p>
<h2>Hosting a country domain name anywhere in the world</h2>
<p>Back to the topic at hand - you might be wondering, "How is it possible to have a country domain name anywhere in the world?"</p>
<p>Without going <strong>too technical</strong>, think of a domain name like a record in an address book. Your customers look up your domain name, and then your domain registrar tells your customer's browser where the server is, and the browser starts loading your page.</p>
<p>Of course, if you have a server in Sydney, Australia, and customers from the US are trying to visit your site, it's going to be slow. Unless...</p>
<h2>Best Practice: Use a CDN!</h2>
<p>Essentially, a CDN (Content Delivery Network) makes a copy of your website available in thousands of servers across the world, so when your customers visit your website, the request doesn't have to travel across the world to reach your server. Instead, the requests only have to travel to your customer's nearest CDN location.</p>
<p>Using a CDN is as simple as updating the records in your domain registrar to point at your CDN, instead of your server. It's also one of the quickest and cheapest ways to significantly speed up your site.</p>
<p>I personally use CloudFlare for my own website (<a href="https://maxrozen.com">MaxRozen.com</a>) - their three pricing tiers are free, $20 per month, and $200 per month. Compared to paying a developer $100 per hour to figure out why your website is slow, it's a pretty good deal!</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[How to Improve First Contentful Paint]]></title>
            <link>https://onlineornot.com/how-to-improve-first-contentful-paint</link>
            <guid>https://onlineornot.com/how-to-improve-first-contentful-paint</guid>
            <pubDate>Sun, 02 May 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Chances are, you've run a <a href="https://developers.google.com/speed/pagespeed/insights/">PageSpeed Insights</a> test and noticed "First Contentful Paint" as one of the first numbers in the report.</p>
<p>I've covered most of the metrics before in my article on <a href="/understanding-page-speed-metrics-google-lighthouse">Understanding the Page Speed Metrics in Google Lighthouse</a>, but in this article I wanted to dive deeply into First Contentful Paint - particularly what it is, what a good score is, and how to improve.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-is-first-contentful-paint-fcp">What is First Contentful Paint (FCP)</a></li>
<li><a href="#whats-a-good-first-contentful-paint-fcp-score">What's a good First Contentful Paint (FCP) score?</a></li>
<li><a href="#ways-to-reduce-first-contentful-paint-fcp">Ways to reduce First Contentful Paint (FCP)</a>
<ul>
<li><a href="#reduce-time-to-first-byte-ttfb">Reduce Time To First Byte (TTFB)</a></li>
<li><a href="#resize-your-images">Resize your images</a></li>
<li><a href="#lazy-load-your-images-below-the-fold">Lazy load your images below the fold</a></li>
</ul>
</li>
</ul>
<h2>What is First Contentful Paint (FCP)</h2>
<p>First Contentful Paint is a performance metric measured by Google Lighthouse and PageSpeed Insights. It measures the amount of time it takes the browser to show the first visible (as in, <a href="https://en.wikipedia.org/wiki/Above_the_fold#In_web_design">above the fold</a> - the portion of the webpage visible without scrolling) block of text, or image (including SVGs) on your page. Iframes aren't counted when measuring FCP.</p>
<h2>What's a good First Contentful Paint (FCP) score?</h2>
<p>Google considers FCP values of:</p>
<ul>
<li>less than 2 seconds as <strong>fast</strong></li>
<li>greater than 2 seconds, but less than 4 seconds as <strong>moderate</strong></li>
<li>greater than 4 seconds as <strong>slow</strong></li>
</ul>
<h2>Ways to reduce First Contentful Paint (FCP)</h2>
<h3>Reduce Time To First Byte (TTFB)</h3>
<p>In case you're not familiar, TTFB measures the amount of time your server takes to respond to the browser's request. Since First Contentful Paint includes the amount of time it takes the server to send us our data, improving TTFB will also improve our First Contentful Paint.</p>
<p>There are two ways to improve your TTFB:</p>
<ol>
<li>
<p>Reduce the amount of time spent on the server processing the request (this includes the time your server spends on database queries, API calls, and load balancing).</p>
<ul>
<li>The simplest way to reduce this is to upgrade your server's RAM and CPU specs, or pick a better hosting provider.</li>
<li>Alternatively, look at optimising your backend code, though this is significantly easier said than done, particularly if using a CMS like WordPress or Drupal.</li>
</ul>
</li>
<li>
<p>Reduce the amount of time spent sending the request from the server to the browser</p>
<ul>
<li>The simplest way to reduce this is to <a href="https://onlineornot.com/ways-to-improve-page-speed#1---use-a-cdn">use a CDN</a>, and enable GZIP or Brotli compression.</li>
</ul>
</li>
</ol>
<h3>Resize your images</h3>
<p>This an extremely simple, yet often overlooked step. Some frameworks like WordPress have plugins that do this automatically, but some plugins fail to do this well.</p>
<p>There are two steps:</p>
<ol>
<li>Open your image in macOS Preview or Microsoft Paint, resize the image, making it smaller. For example, if you have a photo that's 4000x3000, I'd make it 1000x750.</li>
<li>Now that the image is physically smaller, run it through an image optimiser. I use <a href="https://tinyjpg.com/">TinyJPG</a> for this, but there are hundreds of options.</li>
</ol>
<p>These two steps alone will save you 67% or more.</p>
<h3>Lazy load your images below the fold</h3>
<p>In other words - ensure your above the fold images aren't being lazy loaded. All other content should be lazy loaded.</p>
<p>In WordPress, lazy loading was typically enabled via plugin - as of WordPress 5.4, all images are lazy loaded by default.</p>
<p>If you're using a React framework like Next.js, and using the <code>next/image</code> package to load your hero/above the fold images, ensure you pass the <code>priority</code> prop. Next/image uses lazy loading by default, so you could be taking a hit on your First Contentful Paint without realising it.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Why VPS Plans are Cheaper than Shared Hosting]]></title>
            <link>https://onlineornot.com/why-vps-plans-cheaper-shared-hosting</link>
            <guid>https://onlineornot.com/why-vps-plans-cheaper-shared-hosting</guid>
            <pubDate>Sat, 24 Apr 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>You might be looking around for a new web host, or trying to get a better deal. You notice VPS plans are incredibly cheap - as low as $3-5 per month, while the cheapest shared hosting is around $10 per month.</p>
<p>You wonder, "How is this possible?!" - especially when people recommend moving to a VPS once a site becomes popular.</p>
<p>Let's discuss what these are first, before going into why cheaper doesn't necessarily mean better for you and your business.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#what-a-vps-provides">What a VPS provides</a></li>
<li><a href="#what-shared-hosting-provides">What Shared Hosting provides</a></li>
<li><a href="#why-is-a-vps-cheaper-than-shared-hosting">Why is a VPS cheaper than shared hosting</a></li>
<li><a href="#how-to-pick-vps-vs-shared-hosting">How to pick: VPS vs shared hosting</a></li>
</ul>
<h2>What a VPS provides</h2>
<p>VPS stands for Virtual Private Server. You typically get allocated a certain amount of CPU, RAM, Disk space, and bandwidth. Most providers let you pick the operating system, and you get admin access on that operating system.</p>
<p>Everything else, is up to you. This means:</p>
<ul>
<li>Operating System updates</li>
<li>Security (managing non-root users)</li>
<li>Configuring, and updating the software you want to run (for WordPress, this also means configuring the database)</li>
</ul>
<p>Access to your VPS is via SSH, and you use terminal commands to install what you need.</p>
<p>Essentially, you're only paying for exclusive access to the hardware, and bandwidth. The provider only ensures that your VPS stays up, and the network is working as expected.</p>
<h2>What Shared Hosting provides</h2>
<p>Shared Hosting provides you with pre-configured software, and some disk space for file storage. Depending on the provider, they'll also bundle in 24/7 Customer Support, CDN, SSL certificate, and domain management.</p>
<p>Access to your shared hosting is typically via some sort of admin panel, such as cPanel, and there's only so much configuration that's accessible to you as a user.</p>
<p>Hosting providers will rarely guarantee a certain amount of CPU/RAM. As a result, you'll find the more dodgy providers will attempt to run as many shared hosting plans on a single server as possible, while more reputable providers that have a reputation for fast and reliable shared hosting cap the number of sites running on a single server.</p>
<h2>Why is a VPS cheaper than shared hosting</h2>
<p>VPS plans offer limited customer support compared to shared hosting. As a result, a hosting provider is able to generate cash flow solely from keeping their hardware online.</p>
<p>In contrast, for shared hosting, a hosting provider also needs to pay for people to provide quality customer support, monitor the web servers and databases, and much more "keeping the lights on" type work. They typically pass on the additional costs, which is why shared hosting is more expensive than a VPS.</p>
<h2>How to pick: VPS vs shared hosting</h2>
<p>If you know your way around SSH, terminal commands, Linux in general, and don't want the provider's software to get in your way, chances are you'll want a VPS.</p>
<p>On the other hand, if you just want to get a website online, and are looking for a "set and forget" type solution, shared hosting is what you're looking for.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Understanding the Page Speed Metrics in Google Lighthouse]]></title>
            <link>https://onlineornot.com/understanding-page-speed-metrics-google-lighthouse</link>
            <guid>https://onlineornot.com/understanding-page-speed-metrics-google-lighthouse</guid>
            <pubDate>Sat, 17 Apr 2021 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>If your experience is anything like mine, you're probably pretty confused about all the acronyms involved in checking your page speed.</p>
<p>Chances are you probably also want to know which metrics matter, and if your score is good or bad.</p>
<p><strong>Table of Contents</strong></p>
<ul>
<li><a href="#metrics-in-a-google-lighthouse-report">Metrics in a Google Lighthouse report</a></li>
<li><a href="#what-do-all-of-these-metrics-even-measure">What do all of these metrics even measure?</a>
<ul>
<li><a href="#what-is-time-to-first-byte-ttfb">What is Time to First Byte (TTFB)?</a>
<ul>
<li><a href="#time-spent-on-the-server">Time spent on the server</a></li>
<li><a href="#time-spent-being-transmitted">Time spent being transmitted</a></li>
</ul>
</li>
<li><a href="#what-is-first-contentful-paint-fcp">What is First Contentful Paint (FCP)?</a></li>
<li><a href="#what-is-speed-index">What is Speed Index?</a></li>
<li><a href="#what-is-largest-contentful-paint-lcp">What is Largest Contentful Paint (LCP)?</a></li>
<li><a href="#what-is-time-to-interactive-tti">What is Time to Interactive (TTI)?</a></li>
</ul>
</li>
</ul>
<h2>Metrics in a Google Lighthouse report</h2>
<p>As an example, let's check out the performance section of a <a href="https://developers.google.com/web/tools/lighthouse">Google Lighthouse</a> v7.3.0 report:</p>
<p><img src="assets/understanding-page-speed-metrics-google-lighthouse/onlineornot-google-lighthouse-result.png" alt="Google Lighthouse Performance Result"></p>
<p>It helps to think of page speed metrics in a few categories:</p>
<ul>
<li>How long did the server take to start sending the page
<ul>
<li>Time to First Byte (TTFB), or Server Response Time falls under this category</li>
</ul>
</li>
<li>How long did it take for the page to show certain things on the page
<ul>
<li><a href="/how-to-improve-first-contentful-paint">First Contentful Paint (FCP)</a>, Speed Index, and Largest Contentful Paint (LCP) fall under this category</li>
</ul>
</li>
<li>How long did it take for the page to start responding to user interactions
<ul>
<li>Time to Interactive (TTI) and Total Blocking Time (TBT) fall under this category</li>
</ul>
</li>
</ul>
<h2>What do all of these metrics even measure?</h2>
<h3>What is Time to First Byte (TTFB)?</h3>
<p>Time to First Byte measures the time between requesting a page, and getting the first byte of data back from a server. Time to First Byte used to be featured as a main metric tracked by Google Lighthouse. In more recent versions, it's featured as an audit named "Reduce initial server response time".</p>
<p>Google considers values greater than 600ms to be bad.</p>
<p>Some people don't worry as much about this metric since they believe that they can't influence it, but that's not necessarily true.</p>
<p>There are two places you can fix your TTFB:</p>
<h4>Time spent on the server</h4>
<p>Your backend server <em>does</em> have some influence on TTFB, with the main contributors being server-side rendering, database queries, API calls, load balancing, your actual app's code, the server's load itself (particularly if you're using cheap shared hosting).</p>
<p>You can improve your TTFB on your server by:</p>
<ul>
<li>Upgrading your server's specs (more RAM, more CPU), or paying more for a better hosting provider</li>
<li>Upgrading your database's specs (assuming it's on a different server to your app server)</li>
<li>Optimizing your backend code</li>
</ul>
<h4>Time spent being transmitted</h4>
<p>The time spent transferring data from the server to the browser is often overlooked, but there are some easy wins here.</p>
<p>If you're not using a CDN, every request to your website (including images, CSS and JavaScript), gets routed across the world to your server.</p>
<p>As mentioned in <a href="ways-to-improve-page-speed#1---use-a-cdn">Ten Ways to Improve WordPress Page Speed</a>, using a CDN gives you hundreds of little servers across the world close to your users, that massively reduce the time it takes to load your page.</p>
<h3>What is First Contentful Paint (FCP)?</h3>
<p><a href="/how-to-improve-first-contentful-paint">First Contentful Paint</a> measures the time it takes the browser to load the first visible block of text, or image on your page.</p>
<p>Google considers FCP values of:</p>
<ul>
<li>less than 2 seconds as <strong>fast</strong></li>
<li>greater than 2 seconds, but less than 4 seconds as <strong>moderate</strong></li>
<li>greater than 4 seconds as <strong>slow</strong></li>
</ul>
<p>If you have a particularly high FCP value, it might be worth double checking how fonts load on your page. You want to ensure your text is visible while your fonts load to minimize FCP.</p>
<h3>What is Speed Index?</h3>
<p>Speed Index measures how quickly visible content loads over time. As a metric, it rewards pages that provide a good user experience (that is, show visible content as soon as possible), and punishes pages that download all of their JavaScript and CSS before displaying anything.</p>
<p>Essentially this metric penalizes websites that load just a header and footer, then take 10 seconds to load content, while rewarding sites that load all content gradually over the same period of time.</p>
<p>As a user, you typically visit a webpage because you want to see its content. Visiting a website that loads its content slowly is generally a frustrating experience, and we're more likely to abandon a page that takes ages to show anything.</p>
<p>Google considers Speed Index values of:</p>
<ul>
<li>less than 4.3 seconds as <strong>fast</strong></li>
<li>greater than 4.3 seconds, but less than 5.8 seconds as <strong>moderate</strong></li>
<li>greater than 5.8 seconds as <strong>slow</strong></li>
</ul>
<h3>What is Largest Contentful Paint (LCP)?</h3>
<p>In contrast with FCP, Largest Contentful Paint measures the time it takes the browser to load the largest visible block of text, or image on your page.</p>
<p>This metric can be dramatically different between mobile and desktop, depending on how much content you show when the page first loads.</p>
<p>Google considers LCP values of:</p>
<ul>
<li>less than 2.5 seconds as <strong>fast</strong></li>
<li>greater than 2.5 seconds, but less than 4 seconds as <strong>moderate</strong></li>
<li>greater than 4 seconds as <strong>slow</strong></li>
</ul>
<h3>What is Time to Interactive (TTI)?</h3>
<p>Time to Interactive measures the time it takes for a webpage to be fully interactive, which Google defines as:</p>
<ul>
<li>The page displays content</li>
<li>Event handlers are registered for most visible elements on the page</li>
<li>User interactions are responded to within 50ms</li>
</ul>
<p>TTI penalizes pages that optimize for content visibility at the expense of interactivity. That is, you can see the content, but you can't scroll or clicking on things has no effect.</p>
<p>Sites with slow TTI often use Server-Side Rendering (SSR) - a page is sent to the user with complete HTML and CSS, but effectively no scripts. It <em>looks</em> loaded, but actually isn't - not until we finish downloading jQuery, and 2.5MB of other JavaScript dependencies.</p>
<p>Google considers TTI values of:</p>
<ul>
<li>less than 3.8 seconds as <strong>fast</strong></li>
<li>greater than 3.9 seconds, but less than 7.3 seconds as <strong>moderate</strong></li>
<li>greater than 7.3 seconds as <strong>slow</strong></li>
</ul>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Ten Ways to Improve WordPress Page Speed]]></title>
            <link>https://onlineornot.com/ways-to-improve-page-speed</link>
            <guid>https://onlineornot.com/ways-to-improve-page-speed</guid>
            <pubDate>Tue, 13 Apr 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<h2>Introduction</h2>
<p>Page Speed is a pretty big deal these days. As of <a href="https://twitter.com/googlesearchc/status/1326192937164705797?s=20">May 2021</a>, Google will start combining <a href="https://web.dev/vitals/#core-web-vitals">Core Web Vitals</a> (how Google measures page speed) with other UX-related signals to rank your page. In other words, Page Speed impacts your SEO.</p>
<p>Since Google changed Googlebot's algorithm to highly favour fast, mobile-friendly websites, it has become more important to have a fast website. If that's not reason enough to start caring, users typically spend less time, and spend less money, the slower your website's experience is.</p>
<h2>What is Page Speed</h2>
<p>Page Speed is the amount of time it takes to completely load content on your webpage.</p>
<p>The metric we use at OnlineOrNot to measure page speed is mainly <a href="https://web.dev/lcp/">Largest Contentful Paint</a> (it's also one of Google's Core Web Vitals). It's a fancy way of saying "the amount of time it takes your page to show the largest image, or block of text, when loading the page".</p>
<p>There could be dozens of reasons, for any given user, for why your page is slow. Your users could be on the train, passing through a tunnel with a weak signal, or their internet could just be slow.</p>
<p>By following best practices, we can at least mitigate the issue on our end, by ensuring we've done the best job we can.</p>
<h2>10 Ways to Improve your Page Speed</h2>
<p>Now that you know what it is, I'm going to teach you what you need to look at to speed up your page.</p>
<p>Note: these are listed in order of difficulty. At some point, you will need a developer to help optimise your site.</p>
<p>Table of Contents</p>
<ul>
<li><a href="#introduction">Introduction</a></li>
<li><a href="#what-is-page-speed">What is Page Speed</a></li>
<li><a href="#10-ways-to-improve-your-page-speed">10 Ways to Improve your Page Speed</a>
<ul>
<li><a href="#1---use-a-cdn">#1 - Use a CDN</a></li>
<li><a href="#2---enable-gzip-compression">#2 - Enable GZIP compression</a></li>
<li><a href="#3---use-smaller-images">#3 - Use smaller images</a></li>
<li><a href="#4---reduce-the-number-of-requests-your-page-makes">#4 - Reduce the number of requests your page makes</a></li>
<li><a href="#5---avoid-redirects-where-possible">#5 - Avoid redirects where possible</a></li>
<li><a href="#6---reduce-time-to-first-byte">#6 - Reduce Time to First Byte</a>
<ul>
<li><a href="#time-spent-on-the-server">Time spent on the server</a></li>
<li><a href="#time-spent-sending-data">Time spent sending data</a></li>
</ul>
</li>
<li><a href="#7---reduce-and-remove-render-blocking-javascript">#7 - Reduce and remove render blocking JavaScript</a></li>
<li><a href="#8---minify-and-combine-your-css-and-js">#8 - Minify and combine your CSS and JS</a></li>
<li><a href="#9---remove-unused-css">#9 - Remove unused CSS</a></li>
<li><a href="#10---regularly-track-your-sites-speed">#10 - Regularly track your site's speed</a></li>
</ul>
</li>
</ul>
<h3>#1 - Use a CDN</h3>
<p>CDN stands for Content Delivery Network. Using a CDN effectively gives you access to hundreds of little servers across the world that host a copy of your site for you, massively reducing the time it takes to fetch your site. If you're not using a CDN, every request to your website (including images, CSS and JavaScript), gets routed across the world, slowly, to your server.</p>
<p>According to 468 million requests in the <a href="https://twitter.com/HTTPArchive">HTTPArchive</a>, 48% were not served from a CDN. That's more than 224 million requests that could have been more than 50% faster, if they spent a few minutes adding a CDN to their site.</p>
<p>Be sure to check you've configured your CDN correctly - high rates of cache misses in your CDN mean the CDN has to ask your origin server for the resource more often, which kind of defeats the purpose of using a CDN in the first place!</p>
<h3>#2 - Enable GZIP compression</h3>
<p>First of all, you'll want to enable GZIP compression on your CDN (if it isn't already enabled).</p>
<p>On some CDNs, GZIP compression will just be a checkbox labelled "enable compression". Enabling compression will roughly half the size of the files your users need to download to use your website, your users will love you for it.</p>
<p>Once your CDN has compression enabled, look into enabling compression on your server. Some servers will do this automatically, other need plugins (such as WordPress's <a href="https://wp-rocket.me/">WP Rocket</a>).</p>
<p>Assuming you have a CDN, enabling compression on your server will save you bandwidth costs between the CDN and your server, and make your site faster for requests that experience cache misses.</p>
<h3>#3 - Use smaller images</h3>
<p>This means both reducing the resolution (such as from 4000x3000 pixels your camera outputs to 1000x750 for the web), and reducing the size by compressing the file.</p>
<p>There are WordPress plugins that will do this automatically for you as you upload images. In particular, try these:</p>
<ul>
<li><a href="https://wordpress.org/plugins/autoptimize/">Autoptimize</a></li>
<li><a href="https://wordpress.org/plugins/imagify/">Imagify</a></li>
</ul>
<p>You don't necessarily need plugins, or WordPress to optimise your images though. I personally use <a href="https://tinyjpg.com/">TinyJPG</a> to compress images as I write blog posts, and upload the optimised images directly.</p>
<h3>#4 - Reduce the number of requests your page makes</h3>
<p>The goal is to reduce the number of requests necessary to load the top part of your page (known as "above the fold content").</p>
<p>You have a few options here:</p>
<ul>
<li>Reduce the number of requests on the page as a whole, by removing fancy animations, plugins, and images that don't improve the site's experience</li>
<li>Or, you can defer loading content that isn't a high priority through the use of <a href="https://developers.google.com/web/fundamentals/performance/lazy-loading-guidance/images-and-video">lazy loading</a></li>
</ul>
<h3>#5 - Avoid redirects where possible</h3>
<p>Redirects slow down your site considerably. Instead of having special subdomain for mobile users, use responsive CSS and serve your website from one domain.</p>
<p>Some redirects are unavoidable, such as www -> root domain or root domain -> www, but the majority of your traffic shouldn't be experiencing a redirect to view your site.</p>
<h3>#6 - Reduce Time to First Byte</h3>
<p>Time to First Byte or Server Response Time, is the amount of time your browser spends waiting after a request for a resource is made, to receive the first byte of data from the server.</p>
<p>There are two parts:</p>
<h4>Time spent on the server</h4>
<p>You can improve time spent on the server by optimising your server-side rendering, database queries, API calls, load balancing, your app's actual code, and the server's load itself (particularly if you're using cheap, shared web hosting - this <strong>will</strong> impact your site's performance).</p>
<p>Side note - using cheap hosting is a false economy. If your site handles any sort of eCommerce, using cheap, shared web hosting is costing you significantly in lost conversions. Any money you save (and realistically, it'd be $50 per month at most), you lose in missed sales from frustrated users of your site.</p>
<h4>Time spent sending data</h4>
<p>You can greatly reduce time spent sending data by following the earlier steps in this guide:</p>
<ul>
<li>Use a CDN</li>
<li>Enable compression</li>
<li>Send as little data as possible (smaller images, less JavaScript/CSS)</li>
</ul>
<h3>#7 - Reduce and remove render blocking JavaScript</h3>
<p>External scripts (particularly those used for marketing) will often be written poorly, and block your page from loading until it is finished running.</p>
<p>You can reduce this effect by marking external scripts async:</p>
<pre><code class="language-html">&#x3C;script async src="...">&#x3C;/script>
</code></pre>
<p>You can also delay the loading of your marketing scripts until your users start scrolling:</p>
<pre><code class="language-jsx">window.addEventListener(
  'scroll',
  () =>
    setTimeout(() => {
      //insert marketing snippets here
    }, 1000),
  { once: true }
);
</code></pre>
<p>In general, try to avoid using plugins and javascript snippets to add functionality to your site (such as pop-up banners, social media buttons, calendars, etc). In most cases, they'll slow your entire site down for little benefit.</p>
<p>Hire someone to add the required functionality to your theme, or use a different theme.</p>
<h3>#8 - Minify and combine your CSS and JS</h3>
<p>Minifying means using tools to remove spaces, newline characters, and shortening your variable names. Typically this would be done automatically as part of your theme's build process.</p>
<p>As you add plugins to your installation however, you start to accumulate JavaScript and CSS files, slowing your site down.</p>
<p>If you're using Wordpress, you can use a plugin to combine and minify your JavaScript and CSS files. Two worth checking out are <a href="https://wordpress.org/plugins/autoptimize/">Autoptimize</a> and <a href="https://wordpress.org/plugins/w3-total-cache/">W3 Total Cache</a>.</p>
<p>It's worth noting that these are band-aid fixes at best. The best time to make your theme fast was while it was under development.</p>
<h3>#9 - Remove unused CSS</h3>
<p>Since Chrome 59 (released in April 2017), it's been possible to <a href="https://developers.google.com/web/tools/chrome-devtools/coverage">see unused JS and CSS in Chrome DevTools</a>.</p>
<p>To see this, open the DevTools, show the console drawer (the annoying thing that appears when you hit Esc), click the three dots on the bottom left hand side, and open "Coverage".</p>
<p>Hitting the button with a reload icon will then refresh your page, and audit the CSS and JS for usage.</p>
<p>Here's what it looks like when you audit the starting page in Google Chrome:
<img src="assets/ways-to-improve-page-speed/cssunused.png" alt="Unused CSS in Chrome"></p>
<h3>#10 - Regularly track your site's speed</h3>
<p>It's much easier to fix problems with your site's speed within moments of slowing your site down. On top of that, if you make reviewing your site's speed a habit, it becomes a much smaller task to fix things that are slow.</p>
<p>There are free tools to check your website's current speed, two of the most popular being <a href="https://webpagetest.org/">WebPageTest</a> and <a href="https://developers.google.com/web/tools/lighthouse">Google Lighthouse</a>. The downside to these tools is that you need to remember to run them before and after you make a change.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Guidelines for picking where to send monitoring alerts]]></title>
            <link>https://onlineornot.com/guidelines-for-picking-where-send-monitoring-alerts</link>
            <guid>https://onlineornot.com/guidelines-for-picking-where-send-monitoring-alerts</guid>
            <pubDate>Sun, 14 Mar 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>If you've ever had to be on the receiving end of a monitoring system that uses email for alerts, you know how noisy things can get. Particularly if you're working in an agency or freelance-like environment, with dozens of client sites to maintain.</p>
<p>You get so many emails that you start looking into integrations with third-party services like Zapier, and coming up with more and more complex rules to try reduce the noise, such as:</p>
<ul>
<li>wait until a site has been down for 5 minutes before sending an email</li>
<li>wait until a site has been down for 30 minutes before forwarding the email to your helpdesk</li>
</ul>
<p>The trouble isn't so much with email - it's a pretty decent way to receive and store important information - but rather that the alerts are for humans, and humans tend to have limited attention spans. Every alert you get eats up more and more of your attention span, and the lower your attention span, the more likely you are to miss the real "huge business impact" alerts.</p>
<p>So let's take a look at what type of alerts we <strong>should</strong> care about, and where to send them.</p>
<h2>So what should you be alerting on?</h2>
<p>Basically, you should send an alert for anything that needs to be responded to <strong>immediately</strong>, such as your site going down (and your business's ability to make money with it). These alerts should be able to wake an on-call person up, so they can fix the issue ASAP.</p>
<p>If you can't trust that your site is <strong>actually down</strong> when receiving an alert, perhaps it's time to re-evaluate your monitoring service.</p>
<h2>Alert types, and where to send them</h2>
<p>When was the last time you woke up for an email? It's not particularly common, so here are some suggestions for where your alerts should go.</p>
<p>Mike Julian's <a href="https://www.oreilly.com/library/view/practical-monitoring/9781491957349/">Practical Monitoring</a> categorises alerts into three groups, depending on their use-case:</p>
<ul>
<li>Immediate action required
<ul>
<li>Events that should trigger these kinds of alerts: your site being unreachable, users being unable to pay for products, SSL certificate expired.</li>
<li>These are the alerts that should be going to Phone/SMS/Pager, and should wake people up (if you respond to outages overnight).</li>
</ul>
</li>
<li>Awareness needed, but immediate action not required
<ul>
<li>Events that should trigger these kinds of alerts: your database backup failed, warnings that your server is starting to run out of disk space</li>
<li>These are the types of alerts that should go to a team channel, such as Slack, Discord, or Microsoft Teams</li>
<li>You could also send these types of alerts to email, though it might rapidly get too noisy for a single person</li>
</ul>
</li>
<li>Record for historical/diagnostic purposes
<ul>
<li>Events that should trigger these kinds of alerts: your system returning a 5xx error for a request, server timeouts</li>
<li>These are the types of alerts your system should be sending to a logging service.</li>
</ul>
</li>
</ul>
<h2>Summary</h2>
<p>Don't use email as a "catch-all" for your monitoring system's alerts. You should split your alerts by use-case, and send them to different locations:</p>
<ul>
<li>Alerts that require immediate action should go to Phone/SMS/Pager</li>
<li>Alerts that just need awareness, but no immediate action should go to a team channel in Slack/Discord/Microsoft Teams</li>
<li>General system errors should be going to your logging service</li>
</ul>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Building a SaaS in one week: How I built OnlineOrNot]]></title>
            <link>https://onlineornot.com/building-saas-in-one-week-how-built-onlineornot</link>
            <guid>https://onlineornot.com/building-saas-in-one-week-how-built-onlineornot</guid>
            <pubDate>Sat, 06 Mar 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>When I first started building SaaS apps for side projects, it would take me a solid weekend solely to build the Stripe integration. These days, It's possible to build the whole SaaS app in one week.</p>
<p><strong>Update:</strong> I've been running this business for over three years now, if you're interested to see how it's going, check out <a href="https://maxrozen.com/lessons-from-my-third-year-running-a-saas">lessons from my third year running a SaaS</a>.</p>
<p><strong>Table of contents</strong></p>
<ul>
<li><a href="#background-info">Background Info</a></li>
<li><a href="#tech-stack">Tech Stack</a></li>
<li><a href="#approach">Approach</a></li>
<li><a href="#building-the-saas">Building the SaaS</a>
<ul>
<li><a href="#first-a-blog">First, a blog</a></li>
<li><a href="#uptime-checker-test">Uptime Checker Test</a></li>
<li><a href="#on-auth">On Auth</a></li>
<li><a href="#using-tailwind-for-ui">Using Tailwind for UI</a></li>
<li><a href="#on-ssr-and-emotion">On SSR and Emotion</a></li>
<li><a href="#forms">Forms!</a></li>
<li><a href="#latency-troubles">Latency troubles</a></li>
<li><a href="#docs">Docs</a></li>
<li><a href="#stripe-integration">Stripe integration</a></li>
</ul>
</li>
<li><a href="#shipped">Shipped!</a></li>
<li><a href="#where-to-next">Where to next?</a></li>
</ul>
<h2>Background Info</h2>
<p>Let's start with some background about the app: OnlineOrNot is a pretty standard uptime checker - it does have some fancy logic to ensure the page is <em>actually</em> down before notifying you, but essentially it's a login/sign-up page, a dashboard, a "new page" form, a checker service, a transactional email integration, and a payment integration. I built a similar tool in 2018 (under the same name) to snapshot test GraphQL queries, but that project never took off.</p>
<p>I recently <a href="https://maxrozen.com/walkthrough-migrating-maxrozen-com-gatsby-to-nextjs">migrated my blog from Gatsby to Next.js</a>, and that got me thinking - How much effort could it possibly be to use Next.js as the core of a SaaS?</p>
<p>I decided to do a quick test, and built a webpage to manually check if any URL is available, in a single afternoon (I later made it part of my landing page, before removing it entirely). I was impressed by how seamlessly Next.js blended together React, server-side rendering, and API routes - compared to my regular approach of using Gatsby.js as a landing page, create-react-app as the SaaS app, and a ton of AWS Lambda functions to handle the backend.</p>
<h2>Tech Stack</h2>
<p>Keeping in mind that I prefer to <a href="https://mcfunley.com/choose-boring-technology">choose boring technology</a> - I could have used DynamoDB/MongoDB to keep running costs as low as possible, but I opted for a relational data model.</p>
<p>Anyway, here's the tech stack:</p>
<ul>
<li>Core app: Next.js, hosted on <a href="https://vercel.com/">Vercel</a></li>
<li>UI/CSS framework: <a href="https://tailwindcss.com">Tailwind CSS</a></li>
<li>Auth: <a href="https://next-auth.js.org/">NextAuth.js</a> - a Next.js library</li>
<li>Backend: GraphQL + Node.js, written in TypeScript, sitting on a <a href="https://nextjs.org/docs/api-routes/introduction">Next.js API route</a></li>
<li>Database: Postgres</li>
<li>Worker functions: Node.js/TypeScript, hosted on AWS Lambda</li>
<li>Payment integration: <a href="https://stripe.com">Stripe</a></li>
<li>Transactional Emails: <a href="https://www.mailgun.com/">Mailgun</a> - I'd normally prefer to use <a href="https://postmarkapp.com/">Postmark</a>, but didn't want to blocked from shipping by having to wait around to get verified</li>
<li>Newsletter Emails: <a href="https://app.convertkit.com/referrals/l/4f6b4767-e189-42b0-b0eb-d171435c247d">Convertkit</a> (referral link)</li>
</ul>
<h2>Approach</h2>
<p>My work hours for my employer are 9am to 5pm, I tend to wake up at 7am, and go to bed around 11pm. If I gave 100% of my own time to this build, I would have had around 40 hours to work with.</p>
<p>In this particular week, I still either went to the gym, or ran each day, watched some Netflix, had dinner, and went out one night, so I had roughly 20 hours to work with.</p>
<h2>Building the SaaS</h2>
<h3>First, a blog</h3>
<p>To begin with, OnlineOrNot started its life as a blog. In particular, I cloned this example: <a href="https://github.com/vercel/next.js/tree/canary/examples/blog-starter-typescript">blog-starter-typescript</a> from the Next.js repo.</p>
<p>To get most of the blog functionality I wanted (sitemap, RSS feed, and code snippets in particular), I followed <a href="https://maxrozen.com/walkthrough-migrating-maxrozen-com-gatsby-to-nextjs#first-steps-with-nextjs">these steps</a>.</p>
<h3>Uptime Checker Test</h3>
<p>With the blog running smoothly, I built an additional page to manually check uptime on my landing page (since removed).</p>
<p>To get technical with you, the form would submit data to a page living under <code>status/[...url].tsx</code> in my Next.js pages folder. Once it got a request, it sent off a request to an <a href="https://nextjs.org/docs/api-routes/introduction">API route</a>, which then sent the request to one of ten serverless functions deployed around the world (I had to use AWS Lambda for this part).</p>
<p>Seeing app-like functionality work so seamlessly on a blog is what inspired me to see how long it would take me to build a whole SaaS around the AWS Lambda function I wrote.</p>
<h3>On Auth</h3>
<p>For auth I initially wanted to use Passport, and try out Max Stoiber's recently released <a href="https://github.com/mxstbr/passport-magic-login">passport-magic-login</a>, but I didn't enjoy the user experience of logging into an app with your email only, checking your email, clicking a link, then having to find the original tab to be able to use the app.</p>
<p>After a bit of searching, I came across <a href="https://next-auth.js.org/">NextAuth.js</a>, which boasted similar functionality (and heaps more Auth Providers supported), with a bit more polish (to be fair, it's been around longer). It only took me a couple of hours to have both email-only login, and Google Auth0 login working.</p>
<h3>Using Tailwind for UI</h3>
<p>My typical approach to building UI rapidly would still be to use Bootstrap. In React, I opt for the <a href="https://reactstrap.github.io/">Reactstrap library</a>. It's never going to win me any design awards, but it helps me get the job done fast.</p>
<p>Since rewriting my own blog from Gatsby to Next.js, I've started writing my CSS using Tailwind utility classes. I find that whereas with Reactstrap I would use the default Bootstrap styles, then manually override each component, Tailwind has me thinking "How can I make this a component I can re-use?" more often, and it makes the code feel cleaner as a result.</p>
<p>To make my use of Tailwind CSS even faster, I paid for <a href="https://tailwindui.com/">Tailwind UI</a>, which has removed the "oh crap, where do I even start?" step for me when working on the frontend. It's kind of like paying for a designer to tell you "here's what screens like this tend to look like" - I highly recommend it.</p>
<h3>On SSR and Emotion</h3>
<p>One pitfall I noticed with server-side rendering (SSR) and Emotion, is that some of the Tailwind classes use the <code>:first-child</code> selector, and Emotion's SSR support (a CSS in JS library) extracts its CSS to be the first sibling of your component.</p>
<p>As a result, your pages tend to flicker on first load. The fix is simple enough - in Next.js you need to extract the critical styles into the head of the page (I followed the instructions <a href="https://github.com/ben-rogerson/twin.examples/tree/master/next-emotion#extract-styling-on-server-optional">here</a> since I used <a href="https://github.com/ben-rogerson/twin.macro">twin.macro</a> to use Tailwind with Emotion).</p>
<h3>Forms!</h3>
<p>I used to hate building forms in React until I came across <a href="https://react-hook-form.com/">React Hook Form</a>.</p>
<p>Importing the hook, passing <code>ref={register({ required: true })}</code> to a few of my inputs, and implementing a submit handler is all it takes to build awesome forms in minutes these days (though it helps that I've spent a couple of years building forms in React).</p>
<p><img src="/assets/building-saas-in-one-week-how-built-onlineornot/onlineornot-add-page.jpeg" alt="OnlineOrNot Add Page Form"></p>
<p>Once the user submits the form, they get to see their page being monitored:</p>
<p><img src="/assets/building-saas-in-one-week-how-built-onlineornot/onlineornot-your-pages.jpg" alt="OnlineOrNot Your Pages List"></p>
<h3>Latency troubles</h3>
<p>Two days into building the SaaS with login/signup working, and it being possible to add pages to monitor, I decided to ship a v0.0.1 for friends to check out. I'm very glad I did, because I immediately noticed doing almost anything on the production build of the site took 3-6 seconds per click.</p>
<p>I eventually realised that the issue was caused by where I host my database. I typically host my Postgres database in AWS's ap-southeast-2 (Sydney, Australia) region, and Vercel hosts their functions somewhere on the US East Coast - probably in AWS's us-east-1 region.</p>
<p>So each request would hit their CDN, travel half-way across the world to the US, travel half-way across the world back to Sydney, Australia, and back to the US to finish the request.</p>
<p>Shutting down my Sydney database and spinning up a new one in AWS's us-east-1 region greatly reduced the latency issues.</p>
<h3>Docs</h3>
<p>Before launching, I knew I needed documentation - even though initially I was just building OnlineOrNot as a test to see how much effort is involved in building a SaaS in 2021, as a developer, nothing pisses me off more than crap/incomplete/missing documentation.</p>
<p>I followed the instructions at <a href="https://v2.docusaurus.io/">Docusaurus' homepage</a>, spun up a quick site in Vercel, and mounted it onto the /docs path of OnlineOrNot - a pretty nifty <a href="https://vercel.com/docs/configuration#routes">Vercel feature</a> that I loved from AWS CloudFront.</p>
<h3>Stripe integration</h3>
<p>Integrating Stripe's Checkout requires adding products in the Stripe dashboard, and following the <a href="https://stripe.com/docs/payments/checkout">docs</a> - there are also examples to follow.</p>
<p>Since I implemented my last Stripe integration, Stripe launched <a href="https://stripe.com/en-au/billing">Billing</a> - greatly reducing the amount of effort required to let users upgrade, change, or cancel their subscriptions.</p>
<p>One of my employer's core values is "Don't f**k the customer", and I <em>strongly</em> identify with it, carrying it with me even when building side-projects. So I knew I had to have Stripe Billing to let my customers cancel at any time.</p>
<p>Rather than build the whole integration myself, I paid for <a href="https://divjoy.com/?via=oon">Divjoy</a> (referral link), which conveniently includes a Stripe integration (with Checkout, Billing, and a webhook endpoint). While I only wanted the Billing portion of the codebase, it was money well-spent - it gave me ideas to refactor the integration I built by following the Stripe docs.</p>
<h2>Shipped!</h2>
<p>With the Stripe integration built and tested, I deployed everything, and here we are - one week after I started toying around with a simple manual uptime checker in Next.js, I have a SaaS with login/sign-up, an app that tracks page uptime, and a way to pay for it (there's also a <a href="/docs/free-tier">free tier</a>).</p>
<h2>Where to next?</h2>
<p>Glad you asked!</p>
<p>I keep a <a href="/changelog">changelog</a> and a <a href="/docs/roadmap">product roadmap</a> tracking what's on the radar, what's coming soon, and what's released - I intend to continue shipping features for OnlineOrNot, and continue running it as part of my side-business.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Migrating between WordPress hosts without downtime]]></title>
            <link>https://onlineornot.com/migrating-wordpress-hosts-without-downtime</link>
            <guid>https://onlineornot.com/migrating-wordpress-hosts-without-downtime</guid>
            <pubDate>Fri, 26 Feb 2021 00:52:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Sometimes you'll want to migrate WordPress hosts - maybe it's time for renewal, and you found a better deal elsewhere, or your hosting provider isn't as reliable as they promised.</p>
<p>Which is great for you, but your site's readers don't care that it's a better deal - they just want to see your content. So minimising downtime when transferring hosts is a pretty big deal.</p>
<p>Let's learn how to avoid downtime.</p>
<h2>Step 1: Transfer your content</h2>
<p>The first step is to transfer your content to your new host, <strong>without turning off the old host</strong>. Your site on your old hosting provider doesn't have to be shutdown for you to start standing up your site on the new hosting provider.</p>
<p>Whether that's via a plugin, and manually copying your files across using FileZilla, or by getting a developer to copy your database to the new host, you need to have two copies of your site running to avoid downtime.</p>
<p>At the end of this step you should have two WordPress sites:</p>
<ol>
<li>The old one, which your domain is still pointing to (as in, visiting https://yoursite.com still shows your WordPress site on your old hosting provider)</li>
<li>The new one, with a weird URL (depending on your new provider, this could be an IP address, or just a testing URL)</li>
</ol>
<h2>Step 2: Test everything on your new site</h2>
<p>The last thing you want is to realise you forgot to copy your images <strong>after</strong> your site's users start visiting the new site.</p>
<p>Click around your new WordPress site, ensure images work, pages and posts are all there with the URLs you expect.</p>
<p>Once you're satisfied the new site works as you expect, you're ready to update your DNS</p>
<h2>Step 3: Update your DNS settings</h2>
<p>This step will be different depending on your domain registrar (the company you bought the domain from), but the gist of it is that you're changing the address your domain points to from your old hosting provider, to the new one.</p>
<p>This step can take up to 72 hours for some old school providers, and can be as quick as a few minutes with AWS or CloudFlare. The time it takes is due to the DNS TTL values you use - you can read more about that <a href="https://ns1.com/resources/understanding-ttl-values-in-dns-records">here</a>. You can speed this step up by setting your TTL to a low value (like 300 seconds or 5 minutes) <strong>before you start the migration</strong>.</p>
<p>You can test whether the new hosting provider is live by creating a dummy blog post or page using the same URL you tested your new hosting provider with - if the post shows up at https://yoursite.com, your new site is live!</p>
<p>Once you're <strong>absolutely sure</strong> your domain is displaying the WordPress site from your new hosting provider, <em>then</em> you can turn off the WordPress site running at your old hosting provider.</p>
<h2>How does this prevent downtime?</h2>
<p>Essentially by not turning off either WordPress site, you entirely avoid downtime. When you switch the DNS records from the old hosting provider over to the new one, people visiting your site just start seeing the new site (since you don't turn off the old site until the migration is complete).</p>
]]></content:encoded>
        </item>
    </channel>
</rss>