3


Next.js on AWS: Part II

Next.jsonAWS: Part II

The Incidents

Incidents

1

EFS Permissions

The symptom was slow image loads. CloudWatch logs showed hundreds of EACCES: permission denied, mkdir '/app/.next/cache/images' per minute β€” one for every image request. The image optimizer was running, but it couldn't write the result anywhere. Every image was being reprocessed from scratch on every single request.

The Dockerfile had RUN chown at build time β€” sets ownership inside the image. But EFS mounts over /app/.next/cache at container startup, replacing that directory with a root-owned filesystem. Build-time chown undone at runtime. Fix: EFS access point with posixUser: { uid: 1001, gid: 1001 } so the mount presents as uid 1001.

Creating the EFS access point
# Wrong β€” Path=/ always exists, so CreationInfo is silently skipped.
# The EFS root stays owned by root. uid 1001 gets EACCES.
aws efs create-access-point \
  --file-system-id fs-<your-filesystem-id> \
  --posix-user Uid=1001,Gid=1001 \
  --root-directory "Path=/,CreationInfo={OwnerUid=1001,OwnerGid=1001,Permissions=755}" \
  --region us-east-2

# Correct β€” /nextcache doesn't exist, so AWS creates it with uid 1001 ownership.
aws efs create-access-point \
  --file-system-id fs-<your-filesystem-id> \
  --posix-user Uid=1001,Gid=1001 \
  --root-directory "Path=/nextcache,CreationInfo={OwnerUid=1001,OwnerGid=1001,Permissions=755}" \
  --region us-east-2
Volume mounts overwrite the image

When ECS mounts a volume at a path, it replaces whatever the image had there. Build-time chown, chmod, and mkdir are gone the moment the container starts.

2

HTTP Listener Sync

The EFS fix required deploying a new ECS task definition revision. That deployment was triggered manually via the AWS CLI rather than through GitHub Actions. The site immediately went 503 on every page.

ECS manages the HTTPS:443 listener. The HTTP:80 listener is manual β€” a workflow step mirrors the weights after each deploy. Bypassed the workflow once. ECS swapped HTTPS to the new target group; HTTP stayed on the old empty one. CloudFront routes via HTTP β€” 503.

Running aws ecs update-service directly skips the HTTP listener sync, cache warm, and smoke test. For out-of-band fixes, use workflow_dispatch in the GitHub Actions UI β€” all the steps still run.

3

Redis Bandwidth

The day after the caching stack was finally confirmed working, an email arrived from Upstash: the database had used 7 GB of its 10 GB monthly free-tier bandwidth limit in under 24 hours. Redis was working correctly β€” the issue was that it was working correctly and the system was configured in a way that consumed enormous bandwidth doing it.

The previous day involved 8 debug deployments, each of which ran the full cache warm cycle: CloudFront invalidation, wait for propagation, hit every page and image URL. That was the obvious suspect. But to know whether it was the deployment session or a structural problem with how much data moved on every request, the right move was to look at the Redis entries directly.

Diagnosing Redis entry sizes
# List all ISR keys in Redis via Upstash REST API
curl -s "$UPSTASH_REDIS_REST_URL/keys/isr:*" \
  -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" | jq '.result[]'

# isr:/_not-found
# isr:/about
# isr:/blog/site-migration
# isr:/vinyl
# ... 33 keys total

# Memory usage (bytes) for a specific key
curl -s "$UPSTASH_REDIS_REST_URL" \
  -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" \
  -H "Content-Type: application/json" \
  -d '["MEMORY", "USAGE", "isr:/vinyl"]' | jq '.result'
# 6504388  (6.2 MB)

# Wire size β€” what Upstash actually bills per GET
curl -s "$UPSTASH_REDIS_REST_URL" \
  -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" \
  -H "Content-Type: application/json" \
  -d '["GET", "isr:/vinyl"]' | wc -c
# 7,323,412 bytes β€” 6.99 MB transferred on every read

The vinyl page's Redis entry was 7 MB β€” every record, full tracklist, serialized into the RSC payload. Every ISR revalidation transferred 7 MB from Upstash. No user triggered it. Just the cycle, every 83 minutes. The webhook was also silently broken. revalidateTag() on the cache handler was a no-op β€” Redis keys were never deleted. Sanity publishes weren't making it to the live site.

Three fixes. A size cap in cache-handler.js β€” entries over the threshold skip the write and re-render from Sanity on demand instead. All revalidate intervals raised to 24 hours β€” short intervals were burning Redis reads with no user benefit since the webhook handles real-time updates. And the webhook now calls redis.del() directly, bypassing the broken revalidateTag() path.

Entry size cap
// cache-handler.js β€” entry size cap
async set(key, data, ctx) {
  if (!redis) return;
  const ttl = typeof ctx?.revalidate === 'number' ? ctx.revalidate : DEFAULT_TTL;
  const entry = { value: data, lastModified: Date.now() };
  const storable = toStorable(entry);
  const serialized = JSON.stringify(storable);
  const byteSize = Buffer.byteLength(serialized, 'utf8');

  if (byteSize > MAX_ENTRY_BYTES) {
    console.warn(`[cache-handler] SET skipped (too large): ${key} (${Math.round(byteSize/1024)} KB)`);
    return;
  }

  console.log(`[cache-handler] SET ${key} (${Math.round(byteSize/1024)} KB, ttl=${ttl}s)`);
  await redis.set(`isr:${key}`, storable, { ex: ttl });
}

The webhook only touches affected pages. Publishing a post deletes isr:/blog/new-slug and isr:/index β€” nothing else. Everything else stays cached and undisturbed.

Calibrating the Cap

The initial 512 KB cap was a guess β€” set before any real post data existed. The first smoke test after the stack was running showed the actual picture: most posts around 79 KB, heavier posts with images and long body blocks in the 200–500 KB range, and four posts between 500 KB and 1.3 MB. Three of those four were being skipped by the cache handler. Every CloudFront miss on those posts was a full Sanity re-render.

Raising to 900 KB covers three of the four. Redis only reads these on CloudFront expiry β€” at most once a day. Three entries averaging ~650 KB is ~2 MB/day additional, nothing against a 10 GB monthly limit. The 1.3 MB outlier stays outside the cap β€” one slow render per day, then 24 hours of edge cache. Cap: 900 KB.

The smoke test's Redis key size check exists precisely for this: after the first real deploy you have actual entry sizes from real content. Calibrate the cap from those numbers, not estimates.

After the health check drain fix, daily bandwidth dropped under 100 MB. CloudWatch logs showed most posts 79–869 KB, the 1.3 MB outlier still skipped. At that baseline, one extra read per day at 1.3 MB is ~39 MB/month β€” trivial. Cap raised to 1400 KB.

4

Vinyl Webhook

With the cap in place, the vinyl page's Redis entry is skipped entirely β€” too large to store. CloudFront still caches it for 24 hours, so users are fine. But the webhook was still firing a CloudFront invalidation and a full warm request every time any vinyl record changed in Sanity. That's where the logic broke down. 700+ records, one route β€” /vinyl. No per-record slugs. Any change to any record stales the whole page. Every webhook trigger fetched all 700+ records from Sanity to re-render the entire collection.

Blog posts get per-post granularity because each has its own URL. Publishing one post only invalidates that slug and the homepage. Vinyl renders everything at one address β€” any change invalidates the whole page. The data model determines cache granularity. Removed vinylRecord from the webhook. The page had revalidate = 86400 already β€” that's now the only mechanism. New records show up within 24 hours, no webhook required. The handler returns 400 on a vinylRecord event; a test pins that so it can't quietly regress.

Stale Keys After Cap Changes

Changing the cap doesn't clean up what's already in Redis. The cap only blocks new writes β€” anything written before the change sits there until its TTL expires, still being read on every ISR cycle. The smoke test caught this after the first deploy: five keys over the new threshold, all written before the cap existed. Vinyl at 6.3 MB. Four blog posts between 519 KB and 1.3 MB. Had to delete them manually.

Any time you raise MAX_ENTRY_BYTES in cache-handler.js, run the smoke test immediately after deploy and manually delete any keys it flags as over cap. They won't self-correct until their TTL expires.

Schema vs. Query Audit

The vinyl payload was also bloated by two fields that were never populated: coverImage and recordImage. They were added with the idea of displaying per-record artwork, but the component that actually does that uses a separate schema entirely. Just a naming coincidence.

With 700+ records, the GROQ query was returning two null fields per document β€” 1,400+ nulls serialized into the RSC payload on every render. A migration script confirmed zero records had ever had either field populated. Dead weight from day one. Removing them from the schema and query trimmed the payload without touching a single record.

One unused field on 700 documents is 700 null values in every query result, RSC payload, and Redis entry. When a page's payload is unexpectedly large, audit the Sanity query first.