7


Next.js on AWS: Part I

Next.jsonAWS: Part I

The Architecture

Scroll down

This site runs on AWS β€” Docker, ECS Fargate, CloudFront, Redis. This is a walkthrough of how the stack is set up, what broke along the way, and one lesson that kept coming up: you can have the right setup, the right config, the right framework exports β€” and caching still isn't working. The only way to know is to curl the live site and check.

Moving Off Vercel

This site started on Vercel, which is great for most Next.js deployments. The friction point was the Chatbot Mixer β€” it was using Vercel's Workflow Dev Kit, which ties you to Vercel's infrastructure for durable executions. That, plus wanting to own the full pipeline, made self-hosting the right move.

The migration was done in two phases: move everything except the Chatbot Mixer first (Phase 1), then migrate the Chatbot Mixer separately once the main site was stable on AWS (Phase 2). Phase 2 also dropped the Workflow Dev Kit entirely β€” the Chatbot Mixer is now a standard stateless streaming endpoint using the AI SDK, no Vercel deps. The site is fully self-hosted. Sanity CMS is unchanged and doesn't care where the frontend lives.

Standard Next.js runs identically in Docker β€” Sanity CMS, API routes, server actions, ISR, next/image, Studio. No Vercel-specific dependencies remain. The Chatbot Mixer moved to a plain AI SDK streaming route.

Infrastructure

The app runs in Docker. The Dockerfile has three stages: install deps, run next build, then copy only the standalone output into a lean runtime image β€” no source, no dev tools, no build artifacts. Final image is around 150MB. next.config.js uses output: 'standalone' which produces a self-contained Node server that doesn't need node_modules at runtime.

Three-stage Dockerfile
# Stage 1: install deps
FROM node:22-alpine AS deps
WORKDIR /app
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile

# Stage 2: build
FROM node:22-alpine AS builder
WORKDIR /app
COPY --from=deps /app/node_modules ./node_modules
COPY . .
ARG NEXT_PUBLIC_SANITY_PROJECT_ID
ENV NEXT_PUBLIC_SANITY_PROJECT_ID=$NEXT_PUBLIC_SANITY_PROJECT_ID
RUN yarn build

# Stage 3: lean runtime image
FROM node:22-alpine AS runner
WORKDIR /app
ENV NODE_ENV=production
ENV HOSTNAME=0.0.0.0
RUN addgroup --system --gid 1001 nodejs
RUN adduser --system --uid 1001 nextjs
RUN mkdir -p /app/.next/cache && chown -R nextjs:nodejs /app/.next
USER nextjs
COPY --from=builder /app/.next/standalone ./
COPY --from=builder /app/.next/static ./.next/static
COPY --from=builder /app/public ./public
COPY --from=builder /app/cache-handler.js ./cache-handler.js
CMD ["node", "server.js"]

The container runs as a non-root user (nextjs, uid 1001). NEXT_PUBLIC_* vars are passed as build args β€” Next.js bakes them into the bundle at compile time, so they can't be injected at runtime.

AWS Stack

Images go to Amazon ECR and run on ECS Fargate Express Mode β€” managed containers, no EC2 to think about. Express Mode does canary deploys by shifting traffic between two target groups, so there's no downtime during a deploy. An ALB sits in front of ECS, CloudFront sits in front of the ALB.

CloudFront β†’ ALB over HTTP

CloudFront talks to the ALB over HTTP. The TLS cert is issued for the public domain, not the ALB hostname β€” HTTPS to the ALB would fail. CloudFront handles TLS for visitors; the internal hop doesn't need it.

CI/CD Pipeline

Every push to main triggers GitHub Actions. It builds for linux/amd64 (dev machine is Apple Silicon, ECS is x86), pushes to ECR, registers a new task definition, updates the service, waits for stabilization, then syncs the HTTP listener weights to match HTTPS after the canary swap. That last step is easy to forget β€” more on that in the incidents section.

1Push to main
1

Push to main

git push to main triggers the deploy workflow.

2

Docker build

GH Actions builds for linux/amd64. Layer caching active.

3

Push to ECR

New image pushed to ECR, tagged with commit SHA and :latest.

4

New task definition

New task definition revision registered.

5

ECS canary deploy

ECS Express Mode shifts traffic to the new task. Old task drains. Zero downtime.

6

HTTP listener sync

HTTP:80 listener weights synced to match HTTPS:443. Skip this and CloudFront 503s.

7

Flush Redis ISR cache

All isr:* keys flushed from Redis. Stale HTML has old chunk hashes β€” must clear before CloudFront invalidation.

8

CloudFront invalidation

CloudFront invalidation submitted. Pipeline waits for full propagation before continuing.

9

Cache warmer runs

Three passes: page routes β†’ Redis + CloudFront; image URLs from HTML β†’ EFS; gallery photos from Sanity β†’ EFS. Gallery images never appear in SSR HTML so Pass 3 is required.

10

Smoke test

smoke-test.mjs runs. Red X if anything is wrong.

Cache Stack

There are three cache layers, each solving a different problem at a different point in the request path. They're independent β€” CloudFront has no clue whether your server used Redis or EFS to build a response. It just caches the HTTP response it gets back.

1

Redis β€” Page Cache

Next.js's default ISR cache dies with the container β€” every redeploy means cold-rendering every page. cache-handler.js swaps that out for Upstash Redis so the cache survives. Server checks Redis first on every request. Hit means no Sanity fetch. 24-hour TTL, with the webhook handling real-time invalidation.

Buffer round-trip fix
// Next.js RSC payloads include Buffer objects (rscData).
// JSON.stringify converts them to { type:'Buffer', data:[...] }
// but they aren't revived to real Buffers on parse.
// Custom replacer/reviver so Buffers survive the Redis round-trip.

function toStorable(value) {
  return JSON.parse(
    JSON.stringify(value, (_k, v) =>
      Buffer.isBuffer(v) ? { __b64__: v.toString('base64') } : v
    )
  );
}

function fromStorable(value) {
  return JSON.parse(
    JSON.stringify(value),
    (_k, v) =>
      v && typeof v === 'object' && typeof v.__b64__ === 'string'
        ? Buffer.from(v.__b64__, 'base64')
        : v
  );
}
cacheMaxMemorySize: 0

next.config.js sets cacheMaxMemorySize: 0. Without it, Next.js keeps its own in-memory cache in front of the custom handler β€” Redis gets bypassed until memory fills up.

2

EFS β€” Image Cache

next/image processes images on first request β€” fetches the original from Sanity, resizes, converts to WebP or AVIF, saves to disk. Fast on every subsequent request. In a container, that processed cache lives on ephemeral disk and gets wiped on every deploy. An EFS volume mounted at /app/.next/cache keeps it persistent across restarts and redeployments.

At scale this matters more. Optimize once, write to shared persistent storage, serve from the edge. The container can be replaced or redeployed without losing any cached work. This also makes hover-preloading viable. When a user hovers a blog card, a /_next/image request fires before they click. EFS reads from disk in milliseconds β€” fast enough to prime the browser before navigation completes. Without it, the preload would still be fetching when the page opened.

Preloading is layered: warm-cache runs at deploy time, listing cards preload the header image on hover, gallery cells preload lightbox-size images on hover. Any one layer is enough; all three together make a cold cache hit nearly impossible.

EFS mounts as owned by root, but the container runs as uid 1001. An access point with posixUser: { uid: 1001, gid: 1001 } fixes this β€” but there's a trap: setting rootDirectory.path to "/" silently skips the ownership setup because the root always exists. Use a non-root path like /nextcache so AWS creates it fresh with the right ownership.

3

CloudFront β€” Edge Cache

CloudFront caches the full HTTP response at edge locations globally. Once it has a page or image, subsequent requests never touch the server. Images are cached for 30 days. If one drops out of CloudFront's cache after 30 days, the next request re-primes it from EFS β€” no reprocessing, just a fast disk read.

β†’User requests a page
β†’

User requests a page

Request hits the nearest CloudFront edge node.

βœ“

CloudFront hit

If CloudFront has the response cached: done. Served from the edge, never reaches your server. Fastest possible.

↓

CloudFront miss β†’ origin

Cache miss (e.g. first request after deploy). Request forwarded to ALB β†’ ECS container.

βœ“

Redis hit

Server checks Redis for the page. On hit, serves cached HTML β€” no Sanity fetch.

↓

Redis miss β†’ Sanity

Cache miss (TTL expired or first request). Server fetches from Sanity, renders page, writes result to Redis.

β†’

Image request

Browser requests /_next/image?url=... Server checks EFS. On hit: serves WebP from disk. On miss: fetches original from Sanity CDN, processes, writes to EFS, serves.

βœ“

CloudFront primed

The response flows back through CloudFront which caches it for next time. Subsequent requests for this URL never reach the server.

Cache Warming

Redis and EFS survive deploys, but CloudFront is cold for anything that hasn't been requested recently. Without warming, the first real visitor after a deploy eats the full slow path: CloudFront miss, hit ECS, Redis/EFS serves it fast, CloudFront caches it. Every visitor after that is fine β€” but that first one shouldn't have to pay the price.

The warmer is the last step in the deploy workflow. Pass 1 hits every page route to seed Redis and CloudFront. Pass 2 scrapes /_next/image URLs from the rendered HTML and requests each one to populate EFS. Pass 3 queries Sanity directly for gallery photo URLs and warms lightbox sizes. Gallery images never appear in SSR HTML β€” Pass 2 alone would miss them entirely.

Three-pass warm strategy
async function main() {
  const blogPaths = await getBlogSlugs();
  const allPaths = [...STATIC_PATHS, ...blogPaths];

  // Pass 1: hit every page β†’ populates Redis ISR cache
  console.log(`── Pass 1: ISR cache β€” warming ${allPaths.length} pages ──`);
  const allImageUrls = new Set();
  for (const path of allPaths) {
    const html = await warmPage(path);
    if (html) {
      // collect /_next/image URLs embedded in the SSR HTML
      extractImageUrls(html).forEach(u => allImageUrls.add(u));
    }
  }

  // Pass 2: request each image URL β†’ triggers optimizer, writes to EFS, primes CloudFront
  console.log(`── Pass 2: SSR image cache β€” warming ${allImageUrls.size} images ──`);
  for (const imageUrl of allImageUrls) {
    await warmImage(imageUrl);
  }

  // Pass 3: gallery images β€” lightbox and fullscreen variants
  // Gallery images are never in SSR HTML (they render inside client-side portals triggered
  // by user clicks). Pass 2 would never find them. Without this pass, every first click
  // on a gallery photo hits cold EFS.
  const galleryImageUrls = await buildGalleryImageUrls();
  console.log(`── Pass 3: Gallery image cache β€” warming ${galleryImageUrls.length} images ──`);
  for (const imageUrl of galleryImageUrls) {
    await warmImage(imageUrl);
  }
}

By the time a real user arrives, every page is in Redis, every SSR image is in EFS and CloudFront, and every gallery photo is pre-warmed. The warmer takes the cold-cache hit so no one else has to.

One caveat: the warmer only seeds the CloudFront edge node (POP) nearest to wherever GitHub Actions runs. CloudFront has hundreds of global edge locations β€” the first visitor from a POP that hasn't been warmed yet still takes the full origin hit. In practice this is fine for a personal site. Traffic is low enough that most POPs will see their first visitor well after they've been naturally warmed by someone else. We looked into CloudFront Origin Shield, which would fix this by adding a single shared caching layer all edge nodes pull from β€” effectively letting one warm run cover every POP. The distribution is on the Free plan which blocks Origin Shield, and the upgrade path wasn't clear enough to justify. At this traffic level, the current setup holds up.