OTA at scale: how Revopush serves billions of update checks

Every time a React Native app with Revopush starts or comes back from the background, it asks our servers whether there is a newer bundle. That's one small request. Across more than 300 million devices a month it adds up to one of the most read-heavy workloads we know of: 16.6 billion requests in the last 12 months, about 2 billion a month today, with peaks close to 4,000 update checks per second.

It didn't start like that. In January 2025 Revopush was a fork of Microsoft's open-source CodePush server, running in one region. This post walks through what we changed since then and why. It's written for engineers who run OTA updates in production (or plan to), and for teams who want to know what's behind the SDK they ship.

16.6B
requests in the last 12 months
300M+
monthly active devices
3,765
peak update checks per second
98.9%
update checks answered at the edge

Requests served per month

TL;DR

  • Plan for spikes, not the average. A push campaign or a release can add thousands of requests per second within a minute.
  • Devices talk to a small stateless service we call the acquisitor. They never hit the management API.
  • Update checks are cached at Cloudflare's edge, in memory in every process, and in Redis. Almost all of them are answered at the edge.
  • The cache key has no per-device data in it, and gradual rollouts are decided by a hash, so one cached response works for millions of devices.
  • On release we clear Redis first and the edge second.
  • Install reports and MAU are counted in memory and written in batches.
  • Cloudflare is the only way in, edge and origin talk over mutual TLS, and bundles are signed by the developer, not by us.

Where we started: a CodePush fork

When Microsoft retired App Center, it open-sourced the CodePush server as a standalone project. It's solid software, but it was built for a world where Azure absorbed the load behind the scenes. Once we ran it ourselves the problems showed up fast.

A single Express process did everything: the CLI API, the dashboard, release uploads, diff generation, and every device's update check. A big release upload competed for CPU with millions of phones asking if anything was new.

Cache misses were expensive. On a miss the server read a pointer from Azure Table, downloaded the deployment's whole release history (up to 50 releases) as one JSON blob, and scanned it. And misses came in bursts, because a release deleted the entire cached hash for that deployment and the next wave of devices all missed at once.

Device reports wrote to Redis on every request, including per-device hashes that grew with MAU. Bundles were served straight from Azure Blob, with Azure egress prices and no CDN in front.

None of this is wrong for a small install. At tens of millions of devices each item turned into either an outage or a large invoice.

Two kinds of traffic

The first thing we did was boring and turned out to matter most. We looked at who sends what.

Developers, CI pipelines and the dashboard send a few thousand requests a day. They create releases, promote, roll back, and they can wait a second or two. Devices send tens of millions of requests a day. Nearly all of them ask the same question, latency goes straight into app start time, and the answer only changes when someone ships a release.

That second workload is a textbook case for caching, as long as the answers can actually be shared between devices. So we kept the CodePush server as our control plane and in March 2025 started a separate service just for device traffic.

Spikes: from quiet to thousands of RPS in a minute

"2 billion requests a month" works out to roughly 770 per second. That sounds easy. The problem is that OTA traffic is never spread evenly.

Apps check for updates when they open, so our traffic follows what millions of people do with their phones, and people tend to do things at the same time. An app sends a push notification to its whole audience and a big share of them open it in the same minute. A team ships a release, every cache for that deployment is cleared on purpose, and every device that opens the app starts a download. On top of that, mornings and lunch breaks move across Asia, Europe and the Americas over the day.

Here's one from October 6. At 13:05 UTC update checks at our edge went from about 2,100 to 3,765 requests per second within a single minute. Requests that reached our own servers went from 32 to 39.

Update-check spike at the edge vs origin

Releases look different, because there the burst is in bytes. This is one release on October 7. Bundle downloads went up 10x in a minute, to about 430 MB/s (around 3.5 Gbps), and then slowly came down over the next hour as the rest of the users opened the app.

Bundle downloads after a release

Nobody warns us before these happen. If we sized the platform for the average, every big release would become an incident right when the customer is watching their rollout.

The acquisitor

The acquisitor is a small Node.js service built on Fastify. It speaks the CodePush device protocol and handles three kinds of requests. Update checks are answered from cache, and on a miss they're relayed to the control plane. Install and download reports are counted. Anything else is proxied to the control plane, so old SDK versions and unusual clients keep working.

It keeps no state of its own, which means we can run as many copies as we like in as many places as we like. It also never sends bundle bytes. The response contains a URL and the device downloads from object storage behind a CDN. We moved bundles from Azure Blob to Cloudflare R2 in February 2025.

Revopush OTA architecture

Caching

Two devices that send the same deployment key, the same app version and the same installed package get the same answer. We keep that answer in three places.

The first is Cloudflare's edge. Responses can stay there for up to an hour and are tagged with the deployment key, so most devices get their answer from the nearest Cloudflare location and never reach us. The second is an in-memory LRU cache inside every acquisitor process with a 30-second TTL. The third is Redis, which is shared by all processes. On a miss there, the acquisitor asks the control plane, which computes the answer and writes it to Redis.

The in-memory layer, added in June 2025, was one of the cheapest wins we've had. A handful of keys get read thousands of times a second, and holding them in RAM for 30 seconds removed almost all of that Redis traffic for a few megabytes of memory per process.

Making answers shareable

Caching only helps if lots of devices land on the same key. Two things made that work for us.

The SDK sends a unique client ID with every check. If it ends up in the cache key, each device gets its own entry and the hit rate is near zero. Our key is the request path plus the query string with the client ID removed and the parameters sorted. The control plane, the acquisitor and the edge purge code all build exactly the same string, and we treat it as a fixed contract between them.

Gradual rollouts ("give this release to 20% of users") look like they need per-device state. They don't. We cache both candidate packages and pick one per device with a stable hash:

ts
// same device + same release always gives the same answer
const inRollout = hash(`${clientUniqueId}-${releaseLabel}`) % 100 < rolloutPercent;

A device never flips between versions and we don't store who is in the rollout. The catch is that while a rollout is active the response differs per device, so those responses come from memory or Redis and skip the edge cache.

Invalidation order

When a release goes out, the old "no update" answer has to disappear from every layer quickly. The control plane deletes the deployment's hash in Redis first. Then it puts a purge for the deployment's cache tag on a queue, and a Worker purges the edge.

We learned to care about this order. If the edge is purged first, the next request can miss at the edge, find the stale entry still sitting in Redis, and put the old answer back at the edge for another hour. The in-memory layer doesn't need a purge at all because its 30-second TTL limits how stale it can get.

One small gotcha: Cloudflare matches cache tags case-insensitively, while deployment keys are case-sensitive. We use a hash of the key as the tag.

On a typical day 98.9% of update checks are answered by Cloudflare. Our servers see the rest: new app versions, rollouts in progress, and the first request after a release.

Geo-distribution

Our customers' apps are used all over the world, with a large share in Asia. A check that goes from Seoul to Amsterdam and back adds a few hundred milliseconds to app start, so we spread the read path out in two layers.

Cloudflare's network (300+ cities) terminates TLS, runs the WAF and answers cached checks close to the device. Behind it, Azure Traffic Manager uses geographic routing to send misses to the nearest acquisitor region: West Europe, Central US or Korea South. All regions run the same stateless service. Traffic Manager health-checks each one and stops sending traffic to a region that fails.

Since the acquisitor holds no data, adding a region is just another deployment. Production deploys go one region at a time. We promote the exact image that passed staging without rebuilding it, and if the first region fails the rollout stops there.

Writes without hurting reads

Devices also report when an update was downloaded, installed or failed, and we count monthly active devices. A database write per report would cost us more than all the reads put together.

So each acquisitor process counts in memory, and every few minutes it flushes the counters to Redis in one pipeline (HINCRBY per field). Moving them from Redis to durable storage is a separate step. An expiring marker key in Redis triggers it, and every instance listens for key-expiry events. One instance wins a short SET NX lock and writes up to 100 rows in one batch transaction. Afterwards it subtracts exactly what it wrote, so increments that arrived during the flush aren't lost. The lock stays until it expires and acts as a "done" marker, which stops a second instance from counting the same window twice.

The device request never waits for any of this. If analytics are slow, updates still go out on time.

For monthly active devices we don't keep a set of every device ID. We use HyperLogLog in Redis (PFADD / PFCOUNT), which takes about 12 KB per counter with roughly 0.8% error, no matter how many devices there are.

Edge caching broke this at first. If Cloudflare answers the check, our servers never see the device. We fixed it with a Worker at the edge that collects device IDs from cached hits and posts them to the acquisitor's /checkin endpoint in batches of up to 10,000.

Security

An OTA system can change the code running on millions of phones, so we treat it that way.

Our origin servers only accept traffic from Cloudflare's IP ranges and from the Traffic Manager health checker. On top of that, Cloudflare presents a client certificate signed by our private CA on every request to the origin, and the acquisitor checks it before doing anything else. A full check costs about 100 µs, so we remember certificates we've already verified and the repeat check takes under a microsecond.

The WAF rejects malformed deployment keys and only lets known API paths through. Everything else is blocked at the edge.

Bundles are signed by the developer's CLI and verified by the SDK on the device with a public key embedded in the app. The private key never touches our servers. Even if someone took over our servers, they couldn't ship a bundle your app would accept.

The whole platform runs on managed services, with zero virtual machines to patch.

We partnered with Vanta on SOC 2, and our controls, policies and subprocessors are listed on the Revopush Trust Center. We expect the official SOC 2 report from Vanta by the end of November 2026, and it will be available through the Trust Center. There's more on our security page.

Reliability

Some of the less exciting work that made the biggest difference:

  • Traffic Manager, App Service and our own process manager all watch every instance. A process stuck at high CPU gets restarted.
  • We hit a case in production where Redis client connections silently stopped answering. A watchdog now replaces them.
  • Two releases to the same deployment at the same moment could overwrite each other's history. History writes now carry an ETag and retry on conflict.
  • Metrics go to Grafana Cloud over OpenTelemetry with a capped set of labels. Per-request logging is off at this volume.
  • When there is no update the response is just isAvailable: false. When there is one, binary diffs often turn a 19 MB bundle into a few hundred KB.

Timeline

WhenWhat changed
Jan 2025Fork of Microsoft CodePush server; Redis over TLS
Feb 2025Bundles move from Azure Blob to Cloudflare R2 + CDN
Mar 2025Acquisitor launched: separate read path, metrics moved off the hot path
Jun 2025In-memory cache on top of Redis
Jul 2025Write-behind counters
Aug-Oct 2025Diff updates, signed binary diffs
Nov 2025HyperLogLog MAU
Apr 2026Multi-region acquisitor (EU, US, Korea)
Jul 2026Mutual TLS between edge and origin
Aug 2026Edge caching of update checks with cache-tag purge
Sep 2026Self-hosted (customer Azure) deployment option

What we learned

Look at the shape of your traffic before optimising anything. Splitting device traffic from management traffic did more for us than any tuning inside either one. And size for the spikes. The monthly average is the wrong number to plan around.

Treat the cache key as an API. Once three systems depended on ours, changing it needed a migration with a period of reading both formats.

Prefer determinism to state where you can. Hash-based rollouts let us cache per-device answers without storing a single assignment.

Get the invalidation order right early, because it's a correctness issue and it's painful to debug later.

Keep analytics off the request path. Counting in memory and writing in batches means slow analytics can't delay an update.

Keep the read path stateless. Adding a region then becomes an ordinary deploy.

Design security in from the start. mTLS, allowlists and client-side signature checks are much easier to build in early than to add later.

What's next

If you're moving off App Center CodePush, or your current OTA setup is struggling with load, book a call with us or start for free.