Where AI Compute Goes When the Router Does the Work

My last post ended on a small, quiet machine on the shelf beside the Wi-Fi router. Six months earlier, the IEEE 802.11 working group, which writes the Wi-Fi standard, had voted to start letting the router be that machine, or at least find it for you.

What the Standard Actually Covers

Qualcomm brought the proposal, now called AI Offload, in March, co-signed by Cisco, Meta, MediaTek, Huawei, Sony, Nokia and others. A device that can’t run an AI model (the proposal’s example is “a non-AI Wi-Fi home security camera”) sends the job to a router that can and gets the answer back. It can even tell the router where to download the model, so a camera maker can ship its own AI without an AI chip in the camera.

Routers would announce the service the way they announce their network name, so a device can see what AI is on offer (text, voice, images, video) before it connects.

By July the scope had grown. The draft charter covers AI work done by the router, by another device on your Wi-Fi, or by a device plugged into your network behind the router: the box on the shelf. MediaTek has proposed how that would work: AI boxes check in with the router, and the router passes your request to whichever one can take it.

The standard covers finding the service, connecting to it, and securing the connection. The AI work itself is out of scope; Cisco’s review of the early proposals calls the model “a black box” as far as Wi-Fi is concerned. That’s the right call. A Wi-Fi standard that tried to pin down AI models would be obsolete before it shipped.

When What happened
January 2026 Broadcom sends router makers early samples of the BCM4918, a Wi-Fi 8 router chip with built-in AI hardware
March 2026 Qualcomm announces the Dragonwing NPro A8 Elite, a router chip with its own AI processor; the Wi-Fi working group votes to start on AI Offload
May 2026 The AI Offload group meets for the first time, in Antwerp
September 2026 The working group approves the charter; final sign-off expected in November
July 2029 Target for the first formal vote on a draft standard
July 2030 Target for the finished standard

Why Your Camera Shouldn’t Need Its Own Brain

Today a smart home device gets its AI from a chip inside it or from its maker’s cloud.

Eufy’s $299.99 E40 lock does the first: it stores up to 50 faces, and “all facial data is stored locally.” That works, but every smart device then needs its own AI chip, built from scratch by every company.

Ring does the second. Its Familiar Faces feature tags people at your door, and Amazon says it “processes and stores biometrics collected by Ring cameras on its own servers.” It isn’t offered in Illinois, Texas, or Portland, Oregon, all of which have strict biometric privacy laws.

The cloud can also disappear. When Belkin shut down Wemo’s cloud on January 31, 2026, the app, remote access, and the Alexa and Google integrations went with it. Only the Wemo devices already set up in Apple’s HomeKit kept working.

The router is a third option. A $30 camera with no AI hardware sends its frames to the router, the router runs the model, and the faces stay in the house. If the camera company goes under, the model’s still on your router. The models are small, too: Frigate, a free, open-source recorder for home security cameras, added face and license plate recognition last year, and its small face model “runs locally on the CPU” of an ordinary computer.

The obvious example is the front door: the doorbell camera sees you coming, the router matches your face, and the lock opens before you find your keys. I wouldn’t do it on the face alone. A camera can’t always tell a face from a good photo of one, and a cheap camera won’t have the 3D sensor your phone uses for face unlock. But the router knows your phone just joined the network. Face plus phone beats either one alone (e.g. MFA for your front door).

Two-factor front door: the doorbell camera sends a face and your phone joins the Wi-Fi; the router checks that the face matches someone who lives here and that their phone is on the network, and only then unlocks the door. A photo held up to the camera passes the face check but fails the phone check, so the door stays locked.

Once a cheap device can borrow the router’s model, the list of things worth building gets long. A few:

Devices What the router’s model does
Driveway camera, garage door or gate Reads the plate and opens for the family car. Nobody needs a giant cloud AI to read a license plate.
Porch camera, phone “The FedEx driver left a box” instead of “motion detected,” and another alert if someone else walks off with it.
Cheap smart speaker The speaker only listens for its wake word, like “Alexa”; the router does the rest.
Smoke and CO alarms, lights, heating and cooling Any microphone in the house hears the alarm; the router turns on the lights and shuts off the furnace fan so it stops pushing smoke through the ducts.
Wi-Fi motion sensing, family alerts Spots a fall in a parent’s house from the way a body disturbs the Wi-Fi signal, with no camera in the bathroom. Wi-Fi sensing got its own standard, 802.11bf, in 2025.

Most of these exist today, each locked to one vendor’s app and cloud. The standard would let a camera from one company use a router from another, the way any Wi-Fi phone works with any Wi-Fi router.

Your Laptop Uses the Same Box

Cameras make the easy case, but most of my AI questions come from my laptop. Lenovo’s scope proposal lists laptops as offload clients too, to save battery and run models too big for them. The announcement that lets a camera find the router lets a laptop find the shelf box, so your AI apps can use it by default and go to the cloud only when the policy allows. The software is mostly there: Ollama, a free tool for running models at home, accepts the same requests as OpenAI’s service, so many apps can switch to it by changing one address.

The bigger change is what that box can see. MCP, an open standard for AI apps, lets you plug your calendar or files into a model so it can answer questions about them. At home, you plug in the house: every camera and sensor, what they’ve recorded, and what they can do. Home Assistant, the free home automation platform, already offers its devices to Claude and ChatGPT this way, and the charter’s supporting document describes a router that keeps a directory of the AI agents on the network.

Then your laptop can ask about the house, using data that never left it:

  • “Who came to the door while I was out?”
  • “Was the garage door open overnight this week?”
  • “Are we using more water this month than last?”

A cloud model could answer those too, if you sent it your doorbell footage and your water readings.

Sending Work to the Big Cloud Models Isn’t in the Standard

Some work still wants a frontier model, one of the big cloud AIs like ChatGPT or Claude: summarizing a contract, or anything that needs real reasoning. The standard doesn’t cover sending it there, but it did come up.

Lenovo’s proposal listed handing work to a cloud provider as one possible design, and Huawei argued that the biggest models “have to be deployed on the cloud.” The draft charter went the other way. It only counts a device “capable of fulfilling an AI offload request without further assistance from another device,” which, as I read it, rules out a router that just forwards your request to a cloud service. Cisco’s review said where the AI runs is something ordinary networking already handles.

That’s right for a Wi-Fi standard, and it doesn’t get in the way. The box behind the router can answer what it can and send the rest to the cloud. That’s the idea from my first post, sending each request to the cheapest model that can handle it, moved inside the house.

Home Assistant already works this way by default: its built-in assistant handles commands first, and “only questions or commands it can’t understand will be sent to the AI you’ve set up.” Apple Intelligence runs many models on the device, sends harder requests to “larger, server-based models” on Apple’s servers, and for ChatGPT, “users are asked before any questions are sent.”

Tier Where it runs Typical work Leaves the house?
Device The camera, speaker, or glasses Wake word, motion trigger No
Router AI chip in the router Face match, plate read, voice commands No
Shelf box Plugged into your home network Your laptop’s questions, the family’s mail and photos No
Cloud Someone else’s data center Hard reasoning, long research Yes, when the house allows it

The last column is why I want this. Today each device maker decides what goes to its cloud. If requests leave through the router, the house decides, in one policy. Nothing defines that policy yet; here’s roughly what I’d want mine to say:

One policy for the whole house: faces and license plates go to the router and never leave the house; mail, calendar and photos go to the shelf box and stay home; everything else goes to the big AI models in the cloud, but only after asking first.

The Four Years Before 2030

The chips are in router makers’ hands now. The standard lands in 2030 if the schedule holds, and Wi-Fi schedules tend to slip. Until then, every router maker will ship its own version, every camera maker will pair with whoever pays them, and you’ll get a doorbell that only talks to one brand of router. Smart home hubs have worked that way for a decade.

Memory is still the hard part. An AI chip that reads license plates is cheap; a router with enough memory to run a large AI model costs what I complained about last time. That’s why the clause I care most about is the box plugged into your home network. The router only has to know where the big model is, and the box can sit on the shelf and get cheaper on its own schedule.

If the standard ships close to today’s charter, every device in the house will find the box from my last post the same way it finds the Wi-Fi.

Climbing to PHPStan Level 10 on a Twelve-Year-Old Codebase

PHPStan has eleven rule levels, 0 through 10, and level 10 arrived with PHPStan 2.0. Running level 10 on a greenfield project is unremarkable. Running it on a codebase that started before PHP had scalar type hints is a different exercise, and the interesting part is where the return on effort falls off.

Start with a Baseline

The instinct is to run the tool, see thousands of errors, and start fixing. That stalls every time, because the work is unbounded and nothing improves until it is finished.

The baseline file inverts this. It records every existing error as accepted, so the reported count starts at zero and any error you add is yours:

vendor/bin/phpstan analyse --generate-baseline

From that point the tool is useful in CI immediately, on day one, at whatever level you eventually want to reach. New code is held to the standard and old code is a debt you pay down deliberately. The baseline shrinking over time is a much better signal than a number that started enormous and is still enormous.

The trap is letting the baseline get regenerated whenever it becomes inconvenient. Regenerating hides new errors along with old ones. Treat the file as append-only in practice, and make regenerating it a decision somebody has to justify in review.

What Each Level Actually Buys

The levels are not evenly spaced in either difficulty or value. On old code the distribution looked roughly like this for me.

Levels 0 through 5 are close to free and worth doing in one sitting. Unknown classes, unknown methods, wrong argument counts, obviously wrong argument types. These are real bugs in dead code paths, and they were the fastest defect-per-hour of the whole exercise.

Level 6 is the wall. It requires type hints everywhere, which on a legacy codebase means annotating thousands of parameters, properties, and returns that were never written down. It is also where the majority of the value is, because everything above it depends on the types being declared. Budget accordingly, and do it package by package rather than all at once.

Level 7 handles union types correctly, and level 8 reports calling methods on something that might be null. Level 8 found the largest number of genuine, user-visible bugs for me. Old PHP code is full of functions that return an object or false, and calling a method on the false path is a fatal error nobody had hit yet only because the input never went that way in production.

Level 9 treats explicit mixed strictly. Level 10 extends that to implicit mixed, meaning values that are untyped because nobody wrote a type rather than because someone chose mixed deliberately.

Where I Stopped, and Why

Nine was the point where the ratio turned. The errors at level 9 and above are overwhelmingly about proving to the analyser that something is what you already know it is, and the fix is usually an assertion or a narrowing check that exists for the tool rather than for the program.

That work is not worthless. Code paths handling data from outside the process, request parameters, decoded JSON, database rows, third-party API responses, are exactly where implicit mixed represents a real unverified assumption, and level 10 is right to flag them. Those are worth fixing properly.

Internal code where the types are obvious from three lines of context is a different story. There the analyser is asking you to restate something the reader can already see, and the resulting assertions add noise to every function. I run level 9 across the codebase and level 10 on the packages that parse external input, which PHPStan supports directly:

parameters:
    level: 9
    paths:
        - lib
        - bin
    ignoreErrors:
        - identifier: missingType.iterableValue
    tmpDir: build/phpstan

The missingType.iterableValue exclusion is the one I would flag as a genuine judgment call. Annotating every array with its value type via array<string, Thing> is real documentation and real safety, and on a large old codebase it is also an enormous amount of typing for arrays that get built and consumed twenty lines apart. Excluding it globally is a compromise. Excluding it only where the array does not cross a module boundary would be better and PHPStan cannot express that distinction.

The Part That Surprised Me

The bugs found were not the payoff. The payoff was that adding types forced decisions that had been deferred for years. A function returning Thing|false|null was not a typing problem, it was three different error conventions that had accumulated in one place, and writing the type down made that impossible to ignore.

Most of my level 6 work turned out to be design cleanup wearing a static analysis costume. That is a better argument for doing it than the defect count is.

HTTP/2 Multiplexing with libcurl in C++

We recently rewrote a piece of C++ that sends a few hundred small HTTPS requests at a time: DNS-over-HTTPS queries, binary POSTs of about 30 bytes, sent through HTTP proxies to a dozen public resolvers. The old code gave each request its own curl easy handle on a 50-thread pool. A batch of 630 requests opened 630 connections and took 10.3 seconds. The new code sends the same batch over 13 to 15 connections in 1.8 seconds.

The requests didn’t change. They now share HTTP/2 connections, which in libcurl takes a multi handle and one option that’s easy to miss.

Where the Time Actually Went

The old code looked like most libcurl code you’ll find:

CURL *c = curl_easy_init();

curl_easy_setopt(c, CURLOPT_URL, _url.c_str());
curl_easy_setopt(c, CURLOPT_HTTPHEADER, headers);
curl_easy_setopt(c, CURLOPT_POSTFIELDS, _body.data());
curl_easy_setopt(c, CURLOPT_POSTFIELDSIZE, static_cast<long>(_body.size()));

CURLcode res = curl_easy_perform(c);

curl_easy_cleanup(c);

That code already negotiated HTTP/2, since curl does by default over TLS. But the connection lives inside the easy handle, and curl_easy_cleanup() closes it, so every connection carried exactly one request.

For a 30-byte query, the request is the cheap part. Before it comes a CONNECT through the proxy, a TCP handshake and a full TLS handshake. With 50 threads firing at once, one server could see dozens of handshakes at the same moment, and median latency was 783ms for an exchange that takes about 25ms once a connection exists.

Left: a client opening six separate connections to a server, each with its own TCP and TLS handshake before one request. Right: the same six requests sent as streams 1 through 11 inside one TCP and TLS connection that was set up once

HTTP/2 fixes this. One connection carries many concurrent streams, and responses come back in whatever order they finish. You pay for the handshake once.

The Multi Handle Owns the Connections

To survive past one request, a connection has to outlive the easy handle. In libcurl that means the multi handle, which keeps a connection cache. Any transfer added to it can use a connection an earlier one opened, including one still busy with other streams.

CURLM *multi = curl_multi_init();

//
// streams on one connection; this has been the default since 7.62.0,
// but I'd rather the code say it
//
curl_multi_setopt(multi, CURLMOPT_PIPELINING, CURLPIPE_MULTIPLEX);

//
// at most 4 connections per host, and room in the cache for all of them
//
curl_multi_setopt(multi, CURLMOPT_MAX_HOST_CONNECTIONS, 4L);
curl_multi_setopt(multi, CURLMOPT_MAXCONNECTS, 64L);

2 and 4 connections per host performed about the same. I kept 4 as headroom in case a server limits concurrent streams per connection. MAXCONNECTS only needs room for every host’s connections, and we talk to about 12 hosts.

Easy handles are still created per request and freed afterwards. Pooling them buys nothing now that the connection lives in the multi handle.

CURL *c = curl_easy_init();

curl_easy_setopt(c, CURLOPT_URL, _req->url.c_str());
curl_easy_setopt(c, CURLOPT_HTTPHEADER, headers);
curl_easy_setopt(c, CURLOPT_POSTFIELDS, _req->body.data());
curl_easy_setopt(c, CURLOPT_POSTFIELDSIZE, static_cast<long>(_req->body.size()));
curl_easy_setopt(c, CURLOPT_TIMEOUT, 5L);
curl_easy_setopt(c, CURLOPT_NOSIGNAL, 1L);
curl_easy_setopt(c, CURLOPT_WRITEFUNCTION, write_cb);
curl_easy_setopt(c, CURLOPT_WRITEDATA, _req);

//
// hand the request back to us when the transfer completes
//
curl_easy_setopt(c, CURLOPT_PRIVATE, _req);

//
// if a connection to this host is still negotiating, wait for it to confirm h2
//
curl_easy_setopt(c, CURLOPT_PIPEWAIT, 1L);

curl_multi_add_handle(multi, c);

PIPEWAIT Is the Option That Matters

This one caught me. Without CURLOPT_PIPEWAIT, a burst of requests to a host with no open connection doesn’t multiplex. curl can’t know the server speaks HTTP/2 until the first TLS handshake finishes and ALPN says “h2”, so until then every new request sees no usable connection and opens its own. By the time the first one confirms h2, all fifty have one.

Two timelines of five requests to one host. Without PIPEWAIT, all five open their own TCP, TLS and ALPN connection, giving five connections. With PIPEWAIT, request 1 opens the connection while requests 2 through 5 wait for h2 to be confirmed, then all five run as streams on one connection

With PIPEWAIT set, curl holds those requests until the first connection confirms or denies multiplexing, then puts them on it as streams. To check, I ran the example program from the end of this post: 200 queries to Cloudflare’s DoH endpoint, 50 in flight, counting new connections with CURLINFO_NUM_CONNECTS.

PIPEWAIT MAX_HOST_CONNECTIONS New connections
off unlimited (0) 50
off 4 4
on unlimited (0) 1
on 4 1

Without it, multiplexing is enabled and never used. The host cap only limits the damage.

One Thread Talks to curl

A multi handle isn’t thread-safe, and I didn’t want to turn every caller into a callback. So one I/O thread owns the multi handle and nothing else touches curl. Worker threads keep their blocking code: push a request onto a queue, wake the I/O thread, wait on a condition variable.

Four worker threads push requests into a mutex-protected queue and call curl_multi_wakeup. A single I/O thread loops through adding queued handles, curl_multi_perform, curl_multi_info_read, completing and notifying, and curl_multi_poll. Its multi handle keeps one HTTP/2 connection per host, each carrying multiple streams to hosts A, B and C

The loop is curl_multi_perform() plus curl_multi_poll(). curl_multi_socket_action() with epoll scales to far more sockets, but this program has about 15 open. Poll is plenty.

for (;;)
{
    //
    // start anything the workers have queued; m_running is checked under
    // the same lock that submit() and shutdown use
    //
    {
        std::lock_guard<std::mutex> guard(this->m_queue_lock);

        if (this->m_running == false)
        {
            break;
        }

        while (this->m_queue.empty() == false)
        {
            this->_start(this->m_queue.front());
            this->m_queue.pop_front();
        }
    }

    int still_running = 0;
    curl_multi_perform(this->m_multi, &still_running);

    //
    // harvest completed transfers
    //
    int left = 0;
    CURLMsg *msg = nullptr;

    while ((msg = curl_multi_info_read(this->m_multi, &left)) != nullptr)
    {
        if (msg->msg != CURLMSG_DONE)
        {
            continue;
        }

        request_t *req = nullptr;

        curl_easy_getinfo(msg->easy_handle, CURLINFO_PRIVATE, &req);
        curl_easy_getinfo(msg->easy_handle, CURLINFO_RESPONSE_CODE, &req->http_code);
        req->result = msg->data.result;

        curl_multi_remove_handle(this->m_multi, msg->easy_handle);
        curl_easy_cleanup(msg->easy_handle);

        this->_complete(req);
    }

    //
    // sleep until there's socket activity, a timeout, or curl_multi_wakeup()
    //
    curl_multi_poll(this->m_multi, nullptr, 0, 1000, nullptr);
}

curl_multi_wakeup() is safe to call from any thread and interrupts a sleeping curl_multi_poll() immediately, so a new request joins a connection while other streams are still running on it. The producer side:

void cpool::submit(request_t *_req)
{
    {
        std::lock_guard<std::mutex> guard(this->m_queue_lock);

        if (this->m_running == false)
        {
            this->_fail(_req);
            return;
        }

        this->m_queue.push_back(_req);

        //
        // under the same lock shutdown takes before curl_multi_cleanup()
        //
        curl_multi_wakeup(this->m_multi);
    }

    std::unique_lock<std::mutex> lk(_req->lock);
    _req->cv.wait(lk, [_req] { return _req->done == true; });
}

The comment in there is about a real bug. A producer calling curl_multi_wakeup() while shutdown is inside curl_multi_cleanup() is a use-after-free, and taking the queue lock for both closes it. At shutdown, curl_multi_get_handles() (libcurl 8.4 and later) lists the transfers still in flight, so you can fail them and no worker waits forever.

The Numbers

Same machine, same batch, same HTTP CONNECT proxies, about 12 public DoH endpoints, at most 50 requests in flight:

Run Requests New connections p50 p95 Wall time
Easy handle per request, 50 threads 630 630 783ms 1058ms 10.3s
Multi handle, multiplexed (2 runs) 630 13 to 15 22 to 27ms 250 to 295ms 1.8s
Multi handle, multiplexed 1,260 16 27ms 279ms 3.4s

These are single runs from one environment, so treat the latencies as illustrative. The connection counts will hold anywhere: doubling the batch to 1,260 requests needed 16. I didn’t measure CPU or memory, though 615 fewer TLS handshakes can’t have hurt.

The p95 dropped far less than the median. The tail is mostly slow resolvers, and multiplexing can’t make a slow server answer faster.

One result surprised me. With a handshake per request, one public endpoint refused a share of them through one of our proxies: 13.3% of requests failed in one run and 7.7% in another, with SSL connect error and Send failure: Broken pipe. Multiplexed, 0 of 600 failed. Without the proxy, the per-request test had no failures either, so it depended on the path the traffic took. I don’t know the server’s rule, but dozens of simultaneous handshakes from one address is exactly the pattern that trips one.

When the Server Only Speaks HTTP/1.1

None of this code asks for HTTP/2. curl offers h2 and http/1.1 in ALPN and the server picks. If it picks HTTP/1.1, nothing errors: PIPEWAIT releases the waiting requests, and each one needs a connection of its own.

The multi handle still reuses HTTP/1.1 connections, so this beats the old code. But an HTTP/1.1 connection carries one request at a time, which turns the connection cap into a concurrency cap. At 4, only 4 requests to that host run at once and the rest queue inside curl.

Per the curl docs, a queued transfer is already counting down its CURLOPT_TIMEOUT. I pointed the example at a local HTTPS server that offers only HTTP/1.1 in ALPN, sent 200 requests with 50 in flight, and had it take 50ms or 500ms per response:

Server response time MAX_HOST_CONNECTIONS Connections Timed out
50ms 4 4 0 of 200
50ms unlimited (0) 50 0 of 200
500ms 4 51 117 of 200
500ms unlimited (0) 50 0 of 200

At 50ms, 4 connections keep up. At 500ms they manage 8 requests a second, and the 5-second timeout runs out on requests that never got a connection. Each timeout also closes the connection it was using, which is why that row opened 51. A cap sized for HTTP/2, where each connection carries up to 100 streams by default, is far too small for a server that falls back.

The cap covers every host on the multi handle, so you can’t raise it for one. You can find out which hosts fell back, though. CURLINFO_HTTP_VERSION reports what was negotiated:

long version = 0;

curl_easy_getinfo(msg->easy_handle, CURLINFO_HTTP_VERSION, &version);

//
// the server picked HTTP/1.1 in ALPN; this host can't multiplex
//
if ((req->result == CURLE_OK) && (version != CURL_HTTP_VERSION_2_0))
{
    fprintf(stderr, "%s negotiated HTTP/1.x, no multiplexing\n", req->url.c_str());
}

For an HTTP/1.1-only host you have to keep, give it its own multi handle with a higher cap, or a timeout that covers the time spent queued.

What Goes Wrong

Proxies split the connection pool. curl keys cached connections on the proxy as well as the host, since a tunnel through proxy A can’t carry a request for proxy B. Pick a proxy at random per request and requests to the same server land on different connections, and multiplexing quietly stops. We pin each destination host to one proxy and move it only after a connection-level failure.

Streams share fate. When a connection dies, every stream on it dies too. We saw Error in the HTTP2 framing layer and Failed sending data to the peer take out 6 to 8 requests at once. Per-request retries are no longer optional. Ours retries once after moving the host to another proxy, and in testing every retry succeeded.

Don’t force HTTP/1.1 to make an error go away. We tried it while debugging. One DoH server answered 505 HTTP Version Not Supported and two others refused the TLS negotiation. Plenty of DoH servers only speak h2.

The write callback has to be binary safe. A DNS response is full of NUL bytes, and a callback that treats the buffer as a C string truncates it without complaint:

static size_t write_cb(char *_ptr, size_t _size, size_t _nmemb, void *_data)
{
    //
    // append with a length; never strcat() or std::string(_ptr)
    //
    static_cast<request_t *>(_data)->response.append(_ptr, _size * _nmemb);

    return _size * _nmemb;
}

Idle connections expire. curl won’t reuse a connection idle longer than CURLOPT_MAXAGE_CONN, 118 seconds by default, so after a quiet stretch the next burst pays one handshake per host. Worth knowing if your traffic arrives in bursts minutes apart.

You need nghttp2. CURLPIPE_MULTIPLEX does nothing if libcurl was built without HTTP/2. If the curl tool uses the same library, curl --version should list HTTP2 under Features. curl_multi_poll() and curl_multi_wakeup() need 7.68.0 or later. Everything here ran on libcurl 8.18.0 with nghttp2 1.43.0.

A Complete Example

A stripped-down version you can compile and run. It sends A queries for example.com to Cloudflare’s DoH endpoint from a pool of worker threads and counts new connections. Build it with g++ -std=c++23 -O2 h2mux.cpp -o h2mux -lcurl, run ./h2mux 200 50, then comment out the PIPEWAIT line and run it again.

#include <stdio.h>
#include <stdlib.h>
#include <string>
#include <vector>
#include <deque>
#include <mutex>
#include <condition_variable>
#include <thread>
#include <atomic>
#include <curl/curl.h>

struct request_t
{
    std::string url;
    std::string body;
    std::string response;
    long http_code = 0;
    CURLcode result = CURLE_OK;
    bool done = false;
    std::mutex lock;
    std::condition_variable cv;
};

static std::mutex g_queue_lock;
static std::deque<request_t *> g_queue;
static bool g_running = true;
static long g_connects = 0;
static CURLM *g_multi = nullptr;
static curl_slist *g_headers = nullptr;

static size_t write_cb(char *_ptr, size_t _size, size_t _nmemb, void *_data)
{
    static_cast<request_t *>(_data)->response.append(_ptr, _size * _nmemb);
    return _size * _nmemb;
}

//
// a minimal DNS query in wire format: header, then QNAME, QTYPE A, QCLASS IN
//
static std::string dns_query(const std::string &_name)
{
    std::string q("\x00\x00\x01\x00\x00\x01\x00\x00\x00\x00\x00\x00", 12);
    size_t start = 0;

    while (start < _name.size())
    {
        size_t dot = _name.find('.', start);
        if (dot == std::string::npos)
        {
            dot = _name.size();
        }

        q += static_cast<char>(dot - start);
        q += _name.substr(start, dot - start);
        start = dot + 1;
    }

    q += std::string("\x00\x00\x01\x00\x01", 5);
    return q;
}

static void start(request_t *_req)
{
    CURL *c = curl_easy_init();

    curl_easy_setopt(c, CURLOPT_URL, _req->url.c_str());
    curl_easy_setopt(c, CURLOPT_HTTPHEADER, g_headers);
    curl_easy_setopt(c, CURLOPT_POSTFIELDS, _req->body.data());
    curl_easy_setopt(c, CURLOPT_POSTFIELDSIZE, static_cast<long>(_req->body.size()));
    curl_easy_setopt(c, CURLOPT_TIMEOUT, 5L);
    curl_easy_setopt(c, CURLOPT_NOSIGNAL, 1L);
    curl_easy_setopt(c, CURLOPT_WRITEFUNCTION, write_cb);
    curl_easy_setopt(c, CURLOPT_WRITEDATA, _req);
    curl_easy_setopt(c, CURLOPT_PRIVATE, _req);
    curl_easy_setopt(c, CURLOPT_PIPEWAIT, 1L);

    curl_multi_add_handle(g_multi, c);
}

static void io_thread()
{
    for (;;)
    {
        {
            std::lock_guard<std::mutex> guard(g_queue_lock);

            if (g_running == false)
            {
                break;
            }
            while (g_queue.empty() == false)
            {
                start(g_queue.front());
                g_queue.pop_front();
            }
        }

        int still_running = 0;
        curl_multi_perform(g_multi, &still_running);

        int left = 0;
        CURLMsg *msg = nullptr;

        while ((msg = curl_multi_info_read(g_multi, &left)) != nullptr)
        {
            if (msg->msg != CURLMSG_DONE)
            {
                continue;
            }

            request_t *req = nullptr;
            long connects = 0;

            curl_easy_getinfo(msg->easy_handle, CURLINFO_PRIVATE, &req);
            curl_easy_getinfo(msg->easy_handle, CURLINFO_RESPONSE_CODE, &req->http_code);
            curl_easy_getinfo(msg->easy_handle, CURLINFO_NUM_CONNECTS, &connects);
            req->result = msg->data.result;
            g_connects += connects;

            curl_multi_remove_handle(g_multi, msg->easy_handle);
            curl_easy_cleanup(msg->easy_handle);

            std::lock_guard<std::mutex> guard(req->lock);
            req->done = true;
            req->cv.notify_one();
        }

        curl_multi_poll(g_multi, nullptr, 0, 1000, nullptr);
    }
}

static void submit(request_t *_req)
{
    {
        std::lock_guard<std::mutex> guard(g_queue_lock);

        g_queue.push_back(_req);
        curl_multi_wakeup(g_multi);
    }

    std::unique_lock<std::mutex> lk(_req->lock);
    _req->cv.wait(lk, [_req] { return _req->done == true; });
}

int main(int _argc, char **_argv)
{
    int total = (_argc > 1) ? atoi(_argv[1]) : 100;
    int threads = (_argc > 2) ? atoi(_argv[2]) : 20;

    curl_global_init(CURL_GLOBAL_DEFAULT);

    g_multi = curl_multi_init();
    curl_multi_setopt(g_multi, CURLMOPT_PIPELINING, CURLPIPE_MULTIPLEX);
    curl_multi_setopt(g_multi, CURLMOPT_MAX_HOST_CONNECTIONS, 4L);
    curl_multi_setopt(g_multi, CURLMOPT_MAXCONNECTS, 64L);

    g_headers = curl_slist_append(g_headers, "Content-Type: application/dns-message");
    g_headers = curl_slist_append(g_headers, "Accept: application/dns-message");

    std::thread io(io_thread);

    std::vector<request_t> reqs(total);
    for (auto &r : reqs)
    {
        r.url = "https://cloudflare-dns.com/dns-query";
        r.body = dns_query("example.com");
    }

    //
    // workers block on one request at a time, exactly like the old code
    //
    std::atomic<int> next{0};
    std::vector<std::thread> workers;

    for (int t = 0; t < threads; t++)
    {
        workers.emplace_back([&]
        {
            int i;
            while ((i = next++) < total)
            {
                submit(&reqs[i]);
            }
        });
    }
    for (auto &w : workers)
    {
        w.join();
    }

    int ok = 0;
    for (auto &r : reqs)
    {
        if ((r.result == CURLE_OK) && (r.http_code == 200))
        {
            ok++;
        }
    }

    {
        std::lock_guard<std::mutex> guard(g_queue_lock);

        g_running = false;
        curl_multi_wakeup(g_multi);
    }
    io.join();

    curl_multi_cleanup(g_multi);
    curl_slist_free_all(g_headers);
    curl_global_cleanup();

    printf("%d/%d ok, connections opened: %ld\n", ok, total, g_connects);
    return 0;
}

Here it prints 200/200 ok, connections opened: 1. Without PIPEWAIT it opens 4, and with the host cap removed as well it opens 50.

PgBouncer, Prepared Statements, and Picking the Right Pool Mode

PgBouncer is the standard answer when an application opens more PostgreSQL connections than the server can usefully service. A process-per-request runtime and a serverless function that dials the database directly get there by completely different routes and land in the same place. The part that goes wrong is the pool mode, because the aggressive setting is the one everybody wants and it silently changes the semantics your code was written against.

Three Modes, One Real Decision

Session pooling assigns a server connection for the life of the client connection. It is safe and it buys you very little, since a worker process holding a connection for the length of a request is the situation you were trying to fix.

Transaction pooling assigns a server connection for the length of a transaction and returns it to the pool at commit. This is the mode that delivers the numbers people install PgBouncer for, and it is the mode with consequences.

Statement pooling returns the connection after every individual statement, which forbids multi-statement transactions entirely. It exists for specific workloads and is almost never what you want.

What Transaction Mode Takes Away

The rule is that anything living outside a transaction stops being reliable, because between two statements you may be on a different backend. That is a longer list than it first appears:

SET / RESET at session scope
LISTEN / NOTIFY
advisory locks taken outside a transaction
WITH HOLD cursors
temporary tables
session-scoped GUCs set by a connection hook

Session-level SET is the one that catches real applications. A framework bootstrap that sets search_path or timezone once on connect works perfectly in development against a direct connection, and in production the setting lands on whichever backend happened to serve that statement. Every later query gets a different backend without it. The failure is intermittent and nearly impossible to reproduce on demand.

Advisory locks are worse, because the failure is silent rather than noisy. A lock taken with pg_advisory_lock() outside a transaction belongs to a session you no longer control. The transaction-scoped variant, pg_advisory_xact_lock(), is released at commit and behaves correctly under transaction pooling. If you use advisory locks for job coordination, that one function name is the difference between working and quietly running the same job twice.

Prepared Statements Are Fixed, with Conditions

For years the answer to prepared statements under transaction pooling was that you could not use them. PgBouncer 1.21, released in October 2023, changed that. Set max_prepared_statements to a non-zero value and PgBouncer tracks named prepared statements itself, rewriting them to internal names and re-preparing them on whichever backend a client lands on:

[pgbouncer]
pool_mode = transaction
max_prepared_statements = 200

The value is the size of an LRU cache of statements kept on each server connection, so it wants to be at least as large as the number of distinct statements a typical request issues. Zero disables the feature and restores the old behavior.

The condition attached is important. This only works for protocol-level prepared statements, the ones sent through the extended query protocol. A literal PREPARE foo AS SELECT ... sent as a plain text query is invisible to PgBouncer and still breaks, because from the outside it is just another statement.

For PHP specifically, this means checking what PDO is actually doing. With ATTR_EMULATE_PREPARES left on, PDO interpolates parameters client-side and sends plain SQL, so there are no real prepared statements and nothing to break. Turn emulation off and you get genuine protocol-level prepares, which is what you want for both correctness and plan reuse, and which is exactly the case max_prepared_statements exists to handle.

$pdo = new PDO($dsn, $user, $pass, [
    PDO::ATTR_EMULATE_PREPARES => false,
    PDO::ATTR_ERRMODE          => PDO::ERRMODE_EXCEPTION,
]);

How I Choose

Start with transaction mode, then go looking for the session state your application depends on rather than waiting for it to surface. Grep for LISTEN, for pg_advisory_lock, for temporary tables, and for anything setting a GUC on connect. Move search_path out of a connection hook and into the connection string, where PgBouncer passes it through as part of the startup parameters and it survives correctly.

Keep a second pool in session mode on a different port for the small number of things that genuinely need a stable session, which is usually a migration runner and whatever handles LISTEN. Two pools with clear rules beats one pool with exceptions nobody remembers.

And size the pool against what PostgreSQL can actually do, since the point of pooling is to keep the database out of the region where it spends more time context switching than working. A pool larger than the database can service just moves the queue.

MTA-STS and TLS-RPT: Enforcing TLS on Inbound Mail

SMTP encryption is opportunistic by default, which means it is optional, which means it is strippable. A sending server connects, looks for STARTTLS in the EHLO response, and upgrades if it sees it. Remove that one line in transit and the sender falls back to plaintext without complaint, because it has no way to know TLS was ever supposed to happen. MTA-STS is how you tell it.

Why STARTTLS Alone Is Not Enough

The gap is that opportunistic TLS has no expectation to violate. A sender that fails to negotiate TLS has no basis to refuse delivery, since plenty of legitimate mail servers still do not offer it. Certificate validation is usually skipped for the same reason: a mail server presenting a self-signed or expired certificate is common enough that treating it as fatal would break real delivery.

So the default posture is an encrypted channel to an unverified party, downgradeable by anyone in the path. MTA-STS, specified in RFC 8461, lets a receiving domain publish a policy saying TLS is required, the certificate must validate, and here are the hostnames allowed to receive mail. Senders that implement it will refuse to deliver rather than fall back.

The Three Pieces

First, a DNS TXT record at the _mta-sts label telling senders a policy exists:

_mta-sts.example.com.  IN  TXT  "v=STSv1; id=20260916120000Z;"

The id is the whole mechanism for cache invalidation. Senders compare it against the one they saw last time, and they only re-fetch the policy when it changes. Bump it on every policy edit or your change will not be noticed until caches expire on their own.

Second, the policy itself, served over HTTPS from a specific host and path:

https://mta-sts.example.com/.well-known/mta-sts.txt
version: STSv1
mode: enforce
mx: mail.example.com
mx: mail2.example.com
max_age: 604800

mode takes enforce, testing, or none. Every MX that can receive mail for the domain needs a line, and a wildcard like *.example.com matches only the leftmost label, so it covers mail.example.com and not foo.bar.example.com. max_age is in seconds with a ceiling of 31557600, about a year.

Third, and this is the part people trip over: the mta-sts host serving that file needs a valid, publicly trusted certificate of its own. The policy is only as trustworthy as the HTTPS connection that delivered it, so a sender that cannot validate that certificate discards the policy entirely. You have now made your mail delivery depend on a certificate on a host that has nothing else to do with mail, which is exactly the kind of endpoint that quietly expires.

Start in Testing Mode

Publishing mode: enforce as the first step is how you find out about your forgotten backup MX by having mail to it rejected. mode: testing makes senders evaluate the policy, report failures, and deliver anyway.

Leave it there long enough to see a full cycle of your real traffic. A week is reasonable, longer if you have partners who send in bursts. The reports are the point of the exercise, which brings up the other half.

TLS-RPT Tells You What Broke

RFC 8460 defines a companion record that asks senders to report what happened. It is cheap to add and it is the only feedback channel you get:

_smtp._tls.example.com.  IN  TXT  "v=TLSRPTv1; rua=mailto:tlsrpt@example.com"

Reports arrive as JSON, daily, from each sending organization that implements it. The useful part is the failure detail:

{
  "organization-name": "Example Sender Inc",
  "date-range": { "start-datetime": "2026-09-16T00:00:00Z",
                  "end-datetime":   "2026-09-16T23:59:59Z" },
  "report-id": "2026-09-16T00:00:00Z_example.com",
  "policies": [{
    "policy": { "policy-type": "sts", "policy-domain": "example.com" },
    "summary": { "total-successful-session-count": 8214,
                 "total-failure-session-count": 17 },
    "failure-details": [{
      "result-type": "certificate-host-mismatch",
      "sending-mta-ip": "203.0.113.9",
      "receiving-mx-hostname": "mail2.example.com",
      "failed-session-count": 17
    }]
  }]
}

The result-type values are specific enough to act on directly. starttls-not-supported, certificate-expired, certificate-host-mismatch, certificate-not-trusted, validation-failure, and on the policy side sts-policy-fetch-error, sts-policy-invalid, and sts-webpki-invalid. That last group points at your policy host rather than your mail servers, which is a distinction worth internalizing before you start debugging the wrong machine.

The Ways This Bites

Caching cuts both ways. Once a sender has your policy with a week-long max_age, a mistake persists for that long even after you fix the file, unless the sender happens to re-check the id. Adding an MX means updating the policy and bumping the id before the new host starts receiving, and doing it in that order.

The certificate on the mta-sts web host is the failure nobody plans for. It is usually a static file on a server that gets less attention than the mail infrastructure, and when it expires, compliant senders stop trusting your policy. Whether that means mail is deferred or the policy is simply ignored depends on the sender.

Your MX certificates now genuinely have to be correct as well. Names have to match what the policy lists, chains have to be complete, and expiry stops being cosmetic. That is the point of turning it on, and it is worth knowing before you flip to enforce.

To check what you have published, the MTA-STS checker on mrdns.com fetches the record and the policy together and tells you whether they agree. For the mail server configuration underneath it, goodtls.com has per-application guides covering Postfix, Exim, Dovecot, and the rest.