3

Next.jsonAWS: Part III
The Smoke Test
Smoke Test Side Effect
The smoke test has a check that verifies the real webhook secret reaches ECS β the test that would have caught the CloudFront query string stripping bug. To pass, it needs to make a real authenticated POST to /api/revalidate. The original payload was { _type: 'post', slug: 'smoke-test' }.
smoke-test is not a real post slug, but the route doesn't check that. It still deleted isr:/index β the homepage Redis key that warm-cache.mjs had just written moments earlier β created a CloudFront invalidation for /blog/smoke-test and /, and fired a warm request that re-rendered the homepage. Every single deploy was silently double-warming the homepage and burning a CloudFront invalidation on a path that doesn't exist.
The fix: a _type: 'ping' handler that returns { pong: true } immediately with no cache work. The smoke test sends ping, confirms query strings are forwarding, and never touches real documents.
EFS Health Check
The EACCES errors showed up in CloudWatch immediately. The fix was straightforward once the root path trap was understood. The fix was a /api/efs-health route β three lines of code that write and delete a probe file at /app/.next/cache/efs-probe from inside the running container. The smoke test calls it after every deploy. If it returns anything other than {ok: true}, the pipeline fails and the error message names the rootDirectory.path issue directly. The next time this misconfiguration occurs it surfaces as a red X on the first deploy.
// app/api/efs-health/route.ts
import { writeFileSync, unlinkSync } from 'fs';
import { join } from 'path';
const PROBE_PATH = join(process.cwd(), '.next', 'cache', 'efs-probe');
export async function GET() {
try {
writeFileSync(PROBE_PATH, 'ok');
unlinkSync(PROBE_PATH);
return Response.json({ ok: true });
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
return Response.json({ ok: false, error: message }, { status: 500 });
}
}
// scripts/smoke-test.mjs β Section 4
const efsRes = await fetch(`${SITE_URL}/api/efs-health`);
const efsJson = await efsRes.json();
if (efsRes.ok && efsJson?.ok) {
pass('/app/.next/cache is writable (EFS mount accepting writes)');
} else {
fail(`/app/.next/cache write failed β ${efsJson?.error}`);
// message points at rootDirectory.path fix
}Gallery Cache Check
Slow gallery lightbox loads have the same structural problem as the EFS incident: there's no feedback loop. A user clicks a photo, it takes two seconds to appear, and nothing in the deploy pipeline told you Pass 3 was broken or that the quality param was mismatched. The failure is invisible until a human notices it.
Section 5 closes that loop. After every deploy it queries Sanity for the first post with a photoGallery block, replicates the exact computeLightboxWidth() math warm-cache uses, builds the /_next/image URL with the correct w= and q= params, and checks the response for x-cache: Hit from cloudfront. If the header is missing, the pipeline fails and the error message names the URL it tested and the source post β so the first place you look is warm-cache Pass 3, not prod logs.
// scripts/smoke-test.mjs β Section 5
const galleryPhoto = await getOneGalleryPhoto();
if (!galleryPhoto) {
warn('No gallery photo found β skipping gallery cache check');
} else {
const url = buildLightboxImageUrl(galleryPhoto.photo);
const res = await fetch(url);
const cacheHeader = res.headers.get('x-cache') || '';
if (cacheHeader.toLowerCase().includes('hit')) {
pass(`Gallery lightbox image is cached (${galleryPhoto.slug})`);
} else {
fail(
`Gallery lightbox image cache MISS\n` +
` URL tested: ${url}\n` +
` Source post: ${galleryPhoto.slug}\n` +
` Hint: check warm-cache Pass 3 and that quality={80} is set on PhotoSlot and PhotoFullscreen`
);
}
}Route Coverage + Chunk Integrity
Two gaps. First: CloudFront serving stale HTML could pass the cache hit check even with origin completely dead. Second: the Oct 2026 BuildKit incident β the container's HTML referenced chunk hashes that didn't exist. ECS healthy. GH Actions green. Smoke test passed. Site broken.
Section 6a: HEAD every static page and blog slug, fail on any non-200. Dead origin, unreachable route, renamed path β any of these fails immediately. Section 6b: fetch the homepage HTML, extract every /_next/static/ URL, HEAD each one. A 404 is the same ChunkLoadError a browser would get.
Dev Mode Drain
With the health check drain fixed, Upstash usage crept back toward the 10 GB monthly limit again β with no unusual deploy session or debugging activity to explain it. Bandwidth charts showed a consistent 4 GB/day on weekdays, dropping sharply to under 600 MB on weekends. The math didn't add up. CloudFront caching was confirmed. At 24-hour ISR intervals, routine revalidation should generate ~14 MB/day. Something was generating 4 GB/day.
The culprit was next dev running with production credentials. In dev mode, Next.js marks all pages immediately stale β every page load triggers a Redis GET + full re-render + Redis SET against the production database. Hundreds of round-trips a day just from building features. The weekday/weekend pattern was the tell.
// cache-handler.js
// Never connect to Redis in dev mode β Next.js marks all pages stale immediately,
// so every page request generates a GET + SET, burning production bandwidth.
const redis =
process.env.NODE_ENV !== 'development' &&
process.env.UPSTASH_REDIS_REST_URL && process.env.UPSTASH_REDIS_REST_TOKEN
? new Redis({
url: process.env.UPSTASH_REDIS_REST_URL,
token: process.env.UPSTASH_REDIS_REST_TOKEN,
})
: null;When UPSTASH_REDIS_REST_URL is unset, every get() and set() in the cache handler returns immediately. Zero Upstash traffic in dev mode; production is unchanged.
Health Check Drain
After the dev mode fix was deployed, bandwidth was still climbing at roughly 250 MB/hour. The dev guard was confirmed in place. No local dev server was running. Tailing the ECS CloudWatch logs showed the same line, every 30 seconds, in bursts of two or three:
aws logs tail /aws/ecs/default/<your-log-group> --since 30m --region us-east-2
# 18:29:32 [cache-handler] GET hit: /index (477 KB)
# 18:29:32 [cache-handler] GET hit: /index (477 KB)
# 18:29:33 [cache-handler] GET hit: /index (477 KB)
# 18:30:02 [cache-handler] GET hit: /index (477 KB)
# 18:30:03 [cache-handler] GET hit: /index (477 KB)
# 18:30:32 [cache-handler] GET hit: /index (477 KB)
# 18:30:33 [cache-handler] GET hit: /index (477 KB)
# ...
# Exact 30-second cadence. Only /index. Machine-generated.The exact 30-second interval and the fact that it only ever hit /index pointed immediately to a health check rather than real traffic. Checking the ALB target groups confirmed it: both were configured to check / every 30 seconds.
aws elbv2 describe-target-groups --region us-east-2 \
--query 'TargetGroups[*].{Name:TargetGroupName,HealthCheckPath:HealthCheckPath,Interval:HealthCheckIntervalSeconds}' \
--output table
# HealthCheckPath Interval Name
# / 30 ecs-gateway-tg-<id-a>
# / 30 ecs-gateway-tg-<id-b>Both target groups were checking / every 30 seconds. Each ping went through the full ISR pipeline and hit Redis. At 477 KB per read, that's ~5 reads/minute β 3.3 GB/day. The flat 0.1 hits/second in the Upstash chart wasn't user traffic. It was the health check.
A dedicated /api/health route returning 200 ok with zero Redis, Sanity, or EFS involvement. Both ALB target groups now check /api/health. Saved 3.3 GB of Upstash bandwidth per day.
// app/api/health/route.ts
export async function GET() {
return new Response('ok', { status: 200 });
}
// AWS CLI β update both target groups
aws elbv2 modify-target-group --region us-east-2 \
--target-group-arn <tg-1-arn> \
--health-check-path /api/health
aws elbv2 modify-target-group --region us-east-2 \
--target-group-arn <tg-2-arn> \
--health-check-path /api/healthEFS Root Path Trap
Weeks after the EFS access point was added, CloudWatch logs still showed hundreds of EACCES: permission denied lines per minute from the image optimizer. The EFS filesystem had 6 KB stored β essentially untouched. No images had ever been cached. The original fix had never actually worked.
posixUser was correctly set (uid 1001, gid 1001) and the access point showed as available. The problem was rootDirectory.path set to "/". AWS only applies CreationInfo β the part that sets ownership β when the specified path doesn't exist. The EFS root always exists. CreationInfo was silently skipped from day one.
New access point with rootDirectory.path=/nextcache β a path that doesn't exist. AWS creates it on first mount and applies CreationInfo correctly. Container immediately has write access. Belt-and-suspenders: all Dockerfile COPYs got --chown=nextjs:nodejs and mkdir/chown moved to after the COPYs.
AWS doesn't report that CreationInfo was ignored. The access point shows as available either way. The only signal is operational: EACCES errors in CloudWatch.
Chunk Errors Mid-Deploy
On a new deploy, chunk filenames change. Users still on old HTML navigate client-side and the browser 404s. A component in the root layout catches ChunkLoadError and reloads. Brief reload, not a broken page.
The worse case: Redis persists stale HTML across deploys. Warm-cache hits ECS, ECS hits Redis, Redis returns old HTML with old chunk hashes, CloudFront caches it. Every visitor gets 404s on static assets. A reload doesn't help β it fetches the same stale HTML. Fix: flush all isr:* keys before CloudFront invalidation on every deploy.
'use client';
import { useEffect } from 'react';
export default function ChunkErrorReloader() {
useEffect(() => {
const handler = (event) => {
const err = event.reason ?? event.error;
if (err?.name === 'ChunkLoadError' || err?.message?.includes('Loading chunk')) {
window.location.reload();
}
};
window.addEventListener('unhandledrejection', handler);
window.addEventListener('error', handler);
return () => {
window.removeEventListener('unhandledrejection', handler);
window.removeEventListener('error', handler);
};
}, []);
return null;
}3
Draft Mode Gotcha
With the cache stack in place, it would have been easy to ship and move on β the revalidate exports were set, CloudFront was in front of everything, the setup looked right. Instead: two curl hits per URL, the first to let CloudFront prime, the second to check for a HIT. The homepage and image requests were behaving correctly. Blog posts weren't.
# Two hits per URL β first primes CloudFront, second reveals whether it cached.
# Run this against the live site, not localhost. Dev mode never caches anything.
# --- BEFORE (draftMode + Math.random + missing generateStaticParams) ---
curl -sI https://www.emilhewitt.com/ | grep -i "cache-control\|x-cache"
# cache-control: private, no-cache, no-store, max-age=0, must-revalidate
# x-cache: Miss from cloudfront
curl -sI https://www.emilhewitt.com/ | grep -i "cache-control\|x-cache"
# cache-control: private, no-cache, no-store, max-age=0, must-revalidate
# x-cache: Miss from cloudfront β never caches, always hits origin
curl -sI https://www.emilhewitt.com/blog/site-migration | grep -i "cache-control\|x-cache"
# cache-control: private, no-cache, no-store, max-age=0, must-revalidate
# x-cache: Miss from cloudfront
curl -sI https://www.emilhewitt.com/blog/site-migration | grep -i "cache-control\|x-cache"
# cache-control: private, no-cache, no-store, max-age=0, must-revalidate
# x-cache: Miss from cloudfront
# --- AFTER (all fixes applied) ---
curl -sI https://www.emilhewitt.com/ | grep -i "cache-control\|x-cache"
# cache-control: s-maxage=1000, stale-while-revalidate=31535000
# x-cache: Miss from cloudfront β miss on first hit (CloudFront cold)
curl -sI https://www.emilhewitt.com/ | grep -i "cache-control\|x-cache"
# cache-control: s-maxage=1000, stale-while-revalidate=31535000
# x-cache: Hit from cloudfront β cached
curl -sI https://www.emilhewitt.com/blog/site-migration | grep -i "cache-control\|x-cache"
# cache-control: s-maxage=3000, stale-while-revalidate=31533000
# x-cache: Miss from cloudfront
curl -sI https://www.emilhewitt.com/blog/site-migration | grep -i "cache-control\|x-cache"
# cache-control: s-maxage=3000, stale-while-revalidate=31533000
# x-cache: Hit from cloudfront β blog posts finally cachingBlog posts were returning private, no-cache, no-store on every request β CloudFront was passing everything straight to the origin. The contrast with the homepage made the cause obvious: something in the blog post rendering path was opting the route into dynamic rendering. Calling draftMode() in a Server Component opts that entire route into dynamic rendering β regardless of whether draft mode is actually active. The call itself is the trigger, not the return value.
The shared layout had a draftMode() call β used to show an 'exit drafts' link in the navbar when Sanity preview was active. One call in the layout meant every page on the site was dynamic. All the revalidate exports were being silently ignored. The fix: ISR pages use a plain readClient with no draftMode() call, and a separate /preview/blog/[slug] route handles draft content explicitly β dynamic on purpose, isolated.
Next.js never caches in dev mode, so every route looks dynamic anyway. Only catchable by checking Cache-Control headers on the deployed site.
4
Random Value Gotcha
After fixing draftMode(), curling the pages again showed the static routes caching correctly β homepage, about, vinyl, all returning s-maxage and X-Cache: Hit on the second request. Blog posts were still broken. A wider audit of the codebase turned up another dynamic rendering trigger: Math.random() in the shared layout.
The footer shows a random motto from Sanity, picked with Math.random() on the server. Any non-deterministic expression in a Server Component tells Next.js the output can't be cached β so the route goes dynamic. In the shared layout, that meant every page on the site was affected again.
The fix was moving Math.random() into InnerWrapper (a client component) using useMemo. The layout passes the full mottos array as a prop; the client picks one on mount β consistent server render, random result per client.
5
Static Params Gotcha
Math.random() fixed the static routes but blog posts were still returning private, no-cache, no-store on every request. The next step was checking Redis directly β listing every ISR cache key the server had written.
node -e "
const { Redis } = require('@upstash/redis');
require('dotenv').config({ path: '.env.local' });
const redis = new Redis({
url: process.env.UPSTASH_REDIS_REST_URL,
token: process.env.UPSTASH_REDIS_REST_TOKEN,
});
redis.keys('isr:*').then(keys => keys.forEach(k => console.log(k)));
"The output: isr:/about, isr:/vinyl, isr:/technology, isr:/agreement β every static route. Zero blog post entries. Next.js was rendering blog posts on every request without ever writing to the ISR cache. The revalidate export was being completely ignored.
Every page that was caching correctly had either no generateMetadata or a static string. Blog posts were the only route using await params inside generateMetadata. In Next.js 16, a dynamic route that resolves params at request time inside generateMetadata is treated as always-dynamic if generateStaticParams isn't present β no ISR cache write happens, regardless of what revalidate says.
generateStaticParams enumerates known slugs at build time. Next.js pre-renders those slugs and writes ISR cache entries for them. New posts published after build still render on-demand and get cached on first request.
import { cache } from 'react';
export const revalidate = 3000;
// Enumerate known slugs at build time so Next.js writes ISR cache entries.
// Without this, a dynamic route with generateMetadata that awaits params
// is always-dynamic in Next.js 16 β revalidate is silently ignored.
export async function generateStaticParams() {
const slugs = await readClient.fetch(
`*[_type == "post" && defined(slug.current)].slug.current`
);
return (slugs || []).map((slug) => ({ slug }));
}
// React.cache() deduplicates this fetch across generateMetadata + page render.
const getPostData = cache(async (slug) => {
return readClient.fetch(BLOG_POST_PAGE_QUERY, { slug });
});
export async function generateMetadata({ params }) {
const { slug } = await params;
const result = await getPostData(slug);
return { title: result?.post?.title || 'Blog post' };
}
const BlogPostPage = async ({ params }) => {
const { slug } = await params;
const result = await getPostData(slug);
// ...
};What I'd Do Differently
Curl the live site before calling anything done β not localhost. revalidate exports and Cache-Control headers don't tell you caching is working. X-Cache: Hit on the second request does. If it's missing, list the Redis ISR keys. If your route isn't in there, Next.js never wrote a cache entry β something in the render path is opting out quietly.
Test volume mount permissions before shipping β a single touch /app/.next/cache/test in a running container catches the EFS bug immediately. Treat the pipeline as the only deployment path; aws ecs update-service directly skips the HTTP sync and cache warm. Don't assume the infrastructure is doing what you configured. Check.
September Bill
First full month on AWS. Expected ~$15β20/mo based on estimates going in. Actual bill: $64.24. Breakdown: ECS Fargate (1 vCPU / 2 GB) β $33.75. ALB β $15.28, billed hourly just for existing. Public IPv4 (ALB + ECS task) β $13.63. Data transfer (AZ hops from canary deploys) β $1.50. ECR: $0.08. EFS, CloudFront, CloudWatch: $0.00. Total: $64.24.
AWS began charging $0.005/hr per public IPv4 in February 2024 β including IPs on running services, not just elastic IPs. ALB + ECS task = $13.63/mo that wasn't in the original cost model.
None of this is reducible by changing the app. The ALB charges $0.0225/hr whether the site gets zero requests or a million. The IPv4 charge is per address per hour, not per request. The only lever is replacing the infrastructure.
Considering a Move
The obvious alternative would have been AWS App Runner β no ALB, no public IPv4, auto-deploy from ECR, ~$29/mo. But App Runner closed to new customers on April 30, 2026 and entered maintenance mode. This account was created after that date, so it's not an option.
The more interesting option is Fly.io β a Docker-native PaaS that would run the same image, keep CloudFront in place (just update the origin), and drop the ALB and IPv4 charges entirely. Estimated cost: ~$9/mo at 1 GB RAM, potentially ~$5/mo after right-sizing. The deployment pipeline would simplify to: push to main β flyctl deploy β CloudFront invalidation β cache warm. No task definitions, no load balancer listener sync, no stabilization polling.
The architecture is solid and the caching stack is working. The cost question is purely infrastructure β whether the ALB and IPv4 overhead is worth paying for what this site actually needs. That's a migration for another day, but the plan is mapped out.
