Case study: WordPress stability issues during high traffic (fuckup night)

Case study: WordPress stability issues during high traffic (fuckup night)

Last verified: September 20, 2026
10 min read
Guide

Every engineering team that manages growing e-commerce stores or high-visibility publishers eventually meets a traffic spike that breaks their infrastructure. During our early WordUp community sessions, we presented an unvarnished post-mortem of a Black Friday outage on a popular lifestyle and retail portal. The site went from a calm baseline to complete infrastructure failure in less than four minutes.

The error logs showed a classic cascade: Nginx returned 502 Bad Gateway and 504 Gateway Timeout, PHP-FPM workers spiked to 100 percent CPU utilization, and MySQL ground to a halt with hundreds of sleeping connections and table locks. This case study details the forensic analysis of that crash, the architectural anti-patterns that caused it, and the caching strategy with Varnish and Redis that stabilized the platform for subsequent million-visit sales campaigns.

#1. Incident timeline: four minutes to total infrastructure lock

The promotion was scheduled to launch at 20:00 on Black Friday. The client ran an email campaign to 250,000 subscribers, paired with targeted social media ads pointing to a dedicated landing page and product category catalog.

19:58:00 - Baseline traffic: 45 req/min. Load average: 0.18. PHP memory usage: 14%.
20:00:15 - Marketing email lands in subscriber inboxes. Influx jumps to 620 req/min.
20:01:30 - Concurrent visitors reach 1,480 req/min. Server load average climbs from 0.40 to 6.80.
20:02:10 - PHP-FPM active processes hit pool limit (pm.max_children = 50). Queue depth begins expanding.
20:02:45 - Nginx worker connections begin timing out waiting for upstream FastCGI responses. Error 504 rates reach 42%.
20:03:30 - MySQL max_connections (150) saturated. Queries pile up in "Sending data" and "Locked" states.
20:04:00 - Server unresponsive to SSH commands. Complete outage across all catalog endpoints.

The immediate instinct of non-technical stakeholders in such situations is to blame “the server” and order larger cloud instances. However, doubling CPU cores and RAM on an unoptimized WordPress stack during an architectural collapse merely doubles the financial bill without solving the concurrency bottleneck.

#2. Root Cause 1: The admin-ajax.php Anti-Pattern on Unauthenticated Traffic

When the system was brought back online under maintenance mode, forensic examination of the Nginx access logs revealed a startling distribution of requests:

GET /black-friday-deals/ HTTP/2.0  -> 1,480 requests/min
POST /wp-admin/admin-ajax.php      -> 1,480 requests/min

For every single catalog page requested by an anonymous visitor, the browser was executing a synchronous background POST request to /wp-admin/admin-ajax.php.

The culprit was a third-party social proof plugin titled “Live Visitor Counter and Stock Alert”. The plugin script contained an inline jQuery snippet:

jQuery(document).ready(function($) {
    $.post('/wp-admin/admin-ajax.php', {
        action: 'record_product_view',
        product_id: 4821
    }, function(response) {
        $('.live-viewers-badge').text(response.active_viewers);
    });
});

#Why admin-ajax.php Breaks at Scale

In WordPress core, admin-ajax.php is located inside the admin directory, but it handles both backend administrative requests and frontend asynchronous callbacks. Crucially, calling admin-ajax.php executes the complete WordPress bootstrap lifecycle:

  1. Loads wp-config.php and initializes database connectivity.
  2. Loads wp-settings.php, parsing all core constants and classes.
  3. Loads and activates every single installed and active plugin in alphabetical order.
  4. Executes the active theme’s functions.php.
  5. Fires the core hook init, initializing rewrite rules and custom post types.
  6. Fires wp_loaded and resolves the specific AJAX action hook (wp_ajax_nopriv_record_product_view).

A single execution of this cycle took 140ms and consumed approximately 48MB of RAM on the server. At 1,480 requests per minute (roughly 25 requests per second), this single counter feature demanded 1.2GB of freshly allocated RAM every second and consumed all available PHP compute cycles simply setting up and tearing down the application runtime.

#The Engineering Remedy

Frontend tracking and telemetry must never trigger a full application framework bootstrap. For view counters and non-critical metrics:

  • Beacon API: Use navigator.sendBeacon('/api/telemetry', data) pointing to a lightweight standalone handler, an edge worker (Cloudflare Worker), or a high-throughput endpoint that logs to Redis or a streaming log queue (Kafka/Vector).
  • Decoupled API: If processing within PHP is strictly required, register a custom REST route via register_rest_route() that bypasses heavy admin hooks and returns lean JSON headers.
  • Client-Side Synthesis: For psychological social proof widgets (“14 people are viewing this right now”), calculate dynamic ranges client-side based on generalized traffic tiers rather than firing atomic database writes per visitor.

#3. Root cause 2: PHP-FPM process pool starvation and worker math

The production server was an 8-core VPS with 16GB of RAM running Ubuntu with Nginx and PHP-FPM 7.4 (and subsequently upgraded through 8.2 and 8.3). The default configuration of the PHP-FPM pool (/etc/php/8.x/fpm/pool.d/www.conf) had been left on generic package defaults:

pm = dynamic
pm.max_children = 50
pm.start_servers = 10
pm.min_spare_servers = 5
pm.max_spare_servers = 15

#Calculating safe worker capacity

The fundamental limit of concurrent request handling in PHP-FPM is physical RAM. When dynamic processes exceed physical memory and the operating system swaps to disk, request latency increases by a factor of 100x, causing immediate server death.

The mathematical formula for calculating pm.max_children is:

$$\text{Max Children} = \frac{\text{Total Available RAM} - \text{System Reserve (OS, Nginx, MySQL, Redis)}}{\text{Average Memory Usage per PHP Worker}}$$

In our audit, a running WordPress worker with WooCommerce and typical analytics plugins consumed approximately 75MB of RAM during uncached execution.

On a server where 6GB of RAM was reserved for MySQL buffer pools and system utilities, 10GB remained for PHP:

$$\text{Max Children} = \frac{10{,}240\text{ MB}}{75\text{ MB}} \approx 136\text{ workers}$$

Because the configuration had capped pm.max_children at 50, the server was artificially choked:

  • 50 workers running tasks that average 1.2 seconds each (slowed by database locking) can complete at most:

$$50 \div 1.2 \approx 41.6\text{ requests per second}$$

When incoming traffic hit 60 to 80 requests per second, requests filled the Linux socket listen queue (listen.backlog = 511). Once the backlog timed out, Nginx dropped the connections with 504 Gateway Timeout.

#Moving from Dynamic to Static Process Management

For dedicated production environments, pm = dynamic introduces unnecessary CPU overhead because the parent process repeatedly forks and kills worker processes. Switching to pm = static ensures that all workers are pre-allocated at service start:

pm = static
pm.max_children = 120
pm.max_requests = 1000
request_terminate_timeout = 30s
slowlog = /var/log/php/slow.log
request_slowlog_timeout = 3s

The directive pm.max_requests = 1000 instructs each worker to respawn after completing 1,000 requests, successfully clearing memory fragmentation without incurring runtime penalties during active user sessions.

#4. Root Cause 3: MySQL Lock Contention on wp_options and Transients

Without a persistent object cache, WordPress stores all transients (temporary cached query data, API responses, and rate counters) directly inside the wp_options database table.

During our incident, the visitor counter plugin was updating a transient on every request:

// Anti-pattern executed on every pageview:
set_transient( 'active_viewers_' . $product_id, $count, 300 );

Under the hood, set_transient() executes an UPDATE on wp_options where option_name = '_transient_active_viewers_4821'. If the transient does not exist, it executes an INSERT.

#The catastrophe of autoloaded options

In stock WordPress, options have an autoload flag set to 'yes' by default. On every regular page load, WordPress executes:

SELECT option_name, option_value FROM wp_options WHERE autoload = 'yes';

When 25 requests per second began updating transients simultaneously:

  1. Multiple concurrent transactions acquired row-level and table-level locks on wp_options.
  2. InnoDB row lock wait timeouts (innodb_lock_wait_timeout = 50) began triggering.
  3. Every subsequent user who arrived on the homepage could not complete the initial autoload query because the database engine was waiting on lock queues created by the transient updates.
  4. The MySQL connection pool saturated, reaching max_connections = 150. New connections were rejected with “Error establishing a database connection”.

#Solution: offloading state to Redis object cache

By deploying Redis and installing the open-source object-cache.php drop-in, WordPress completely changes its storage strategy:

  • Transients and cached database queries are stored directly in RAM as volatile Redis keys with native TTL expirations.
  • Zero disk I/O occurs for transient reads and writes.
  • Lock contention on wp_options disappears entirely.
  • Read latencies for cached options dropped from 12ms to 0.4ms.

#5. The structural fix: Varnish reverse proxy and Nginx edge caching

The golden rule of web scale is simple: dynamic code execution should only occur when a request is genuinely unique to an authenticated user.

Serving an identical HTML page to 1,500 anonymous visitors must never touch PHP or MySQL. The page should be stored in memory and served as a static snapshot.

We deployed Varnish Cache directly in front of Nginx. Varnish listens on port 80/443 (via TLS termination through Nginx or HAProxy) and inspects incoming HTTP cookies.

#Understanding the Varnish Logic (VCL)

The core configuration (default.vcl) inspects the request headers:

vcl 4.1;

backend default {
    .host = "127.0.0.1";
    .port = "8080";
    .first_byte_timeout = 30s;
}

sub vcl_recv {
    # Strip client cookies for static assets
    if (req.url ~ "\.(css|js|png|jpg|jpeg|gif|ico|webp|avif|woff2)$") {
        unset req.http.Cookie;
        return (hash);
    }

    # Pass administrative requests directly to PHP backend
    if (req.url ~ "^/(wp-admin|wp-login\.php)") {
        return (pass);
    }

    # Pass cart, checkout, and account endpoints in WooCommerce
    if (req.url ~ "^/(cart|checkout|my-account)/") {
        return (pass);
    }

    # If the user has a logged-in cookie or an active cart session, bypass cache
    if (req.http.Cookie ~ "(wordpress_logged_in_|woocommerce_items_in_cart|wp_woocommerce_session_)") {
        return (pass);
    }

    # Strip harmless tracking cookies (Google Analytics, Meta Pixel) that break cache hits
    if (req.http.Cookie) {
        set req.http.Cookie = regsuball(req.http.Cookie, "(^|(?<=; )) *__utm[^;]+;? *", "");
        set req.http.Cookie = regsuball(req.http.Cookie, "(^|(?<=; )) *_ga[^;]+;? *", "");
        set req.http.Cookie = regsuball(req.http.Cookie, "(^|(?<=; )) *_fbp[^;]+;? *", "");
        if (req.http.Cookie ~ "^ *$") {
            unset req.http.Cookie;
        }
    }

    # Serve from cache for anonymous visitors
    return (hash);
}

sub vcl_backend_response {
    # Store clean responses in Varnish RAM cache for 1 hour
    if (beresp.status == 200) {
        set beresp.ttl = 1h;
        set beresp.grace = 6h;
    }
}

#The post-Varnish performance metrics

Once Varnish was activated:

  • Cache Hit Ratio: 96.4% across all promotional catalog traffic.
  • Server Response Time (TTFB): Dropped from 1,850ms to 18ms for cached pages.
  • PHP-FPM Load: During the next promotional wave of 3,000 visitors per minute, PHP-FPM utilization remained below 8% CPU usage.
  • MySQL Query Volume: Reduced by 94% across the cluster.

#6. Architecture checklist for traffic peaks in 2026

To prevent similar outages, our deployment pipeline mandates this verification checklist prior to high-volume campaigns:

  1. Audit Frontend AJAX Calls: Search theme and plugin codebases for references to admin-ajax.php or rest_url() fired automatically on window.load or scroll. Eliminate or defer all telemetry.
  2. Implement Persistent Object Caching: Verify Redis is active and operational via wp redis status. Ensure the transient backend is verified as RAM-backed.
  3. Configure Static PHP Pools: Move away from default dynamic FPM sizing. Calculate your memory threshold and set pm = static.
  4. Deploy Page-Level Edge Caching: Protect origin servers with Varnish, Nginx FastCGI microcaching, or an Edge CDN rule (Cloudflare Cache Everything with Cookie Bypass).
  5. Set Session Cookie Boundaries: Ensure WooCommerce only writes session cookies (woocommerce_items_in_cart) when an item is actually placed in the cart, leaving browse-only visitors completely unauthenticated and cacheable.

When engineering teams respect the boundary between static content delivery and dynamic application compute, WordPress can easily sustain millions of pageviews on modest, cost-efficient infrastructure. If you are preparing a large-scale deployment or recovering from an infrastructure bottleneck, explore our advanced WordPress engineering and optimization services.

Next step

Turn the article into an actual implementation

This block strengthens internal linking and gives readers the most relevant next move instead of leaving them at a dead end.

Related cluster

Explore other WordPress services and knowledge base

Strengthen your business with professional technical support in key areas of the WordPress ecosystem.

What was the root cause of the crash in this case study?#
A visitor counter and stock notification script was dispatching an unauthenticated admin-ajax.php request on every pageview. Under 1500 concurrent visitors per minute, this spawned 25 PHP processes per second, completely saturating the PHP-FPM worker pool and locking the MySQL database with concurrent transient writes.
Why does admin-ajax.php perform poorly under high concurrency?#
admin-ajax.php initiates a complete WordPress core bootstrap on every execution. It parses wp-config.php, connects to the database, loads all active plugins, runs theme functions, and fires the init action hook. This consumes 30MB to 80MB of RAM and 80ms to 300ms of CPU time per request.
How did Varnish and Redis recover the site?#
Varnish was deployed as a reverse proxy in front of Nginx to serve static HTML snapshots of product catalog and blog pages to anonymous visitors directly from memory. Redis was deployed as a persistent object cache via an object-cache.php drop-in, eliminating 95 percent of MySQL queries and removing write lock contention on the wp_options table.
What is the optimal PHP-FPM process manager configuration for high-traffic WordPress sites?#
Production high-traffic servers perform best with pm = static rather than dynamic or ondemand. Static allocation prevents the CPU overhead of continuously spawning and killing worker processes. The number of workers is calculated by dividing available RAM allocated to PHP by the average memory footprint of a single worker.

Need an FAQ tailored to your industry and market? We can build one aligned with your business goals.

Let’s discuss

Related Articles