What the logs showed
On October 6 we audited the site for search. In 54 hours of the load balancer's logs, 40 page requests took more than 5 seconds, and the slowest 1% took 9.96 seconds or more. They weren't spread out: each was the first page after a few quiet minutes, and the pages after it came back in milliseconds.
That pattern has a name: a cold start. Cloud Run, where the site runs, can scale a service down to zero instances when nobody is using it, and starts one again on the next request. Nothing is billed while it sleeps. The first visitor pays instead, with their time.
Why we missed it
Whenever we opened the site, we had just deployed it or were working on it, so an instance was already running. Only the logs saw what a visitor arriving after a quiet spell saw.
Why it took so long
A cold container wasn't the whole story. The web server's container was ready in 1.6 seconds, and its first responses took 6 to 7. The longer wait was behind it:
- The API scaled to zero too, and its public catalog read took 11 to 21 seconds while it woke up.
- Every public page waits for that read, to decide whether to link our Work page, which appears only once a product is published.
- So the first visitor woke two services, one after the other, and waited for both.
One small read on every page had made the slowest service set the speed of the whole site.
What we changed
- 01One minimum instance for the web and one for the API (
minScale: "1"in each Cloud Run service). Neither scales to zero anymore. - 02A 2-second limit on the catalog read. When the API is slow, the page goes out without the Work link, the miss is cached for a few seconds, and the next request tries again.
The second change matters even with the first in place. A page shouldn't be as slow as the slowest thing it reads, and almost nothing on ours depends on that read, so it no longer waits for it more than 2 seconds.
What it costs
Cloud Run bills a minimum instance at its idle rate: $0.0000025 per vCPU-second and per GiB-second in Tier 1 regions, on Google's pricing page as of October 6, 2026. With 1 vCPU and 512 MiB, that's about $9.70 a month per service in a 30-day month, before the free tier. Both services, about $20.
On the other side: the first visitor after a quiet spell may be someone who just clicked a link in one of our posts. For a studio's own site, $20 a month to never make that person wait is an easy call.
What changed after
Over the following day (October 7, from 00:43 to 23:58 UTC), the logs recorded 445 page requests. The median took 0.12 seconds, the slowest 1% took 1.36 seconds, and the slowest of all 4.67. None took more than 5.
It's one day of a young site's traffic, so we'll keep watching. But the 25-second first page is gone.
What we'd tell anyone on Cloud Run
- Scaling to zero suits internal tools and background jobs. For a public site, or an API a page waits on, it moves the cost onto your first visitor.
- Measure from the load balancer's logs, not from your own browser: you are almost never the visitor who arrives cold.
- Give every read a page waits on a time limit, and decide what the page shows without it.
- Do the math before deciding. Here the fix cost about $20 a month.
If a product of yours has the same symptom, send us a brief. A person reads it and replies with questions or a first plan.