$ cat ./blog/rethinking-aws-serverless-api-latency

How AI helped me rethink an AWS serverless API and cut latency by 63%

2026-10-01

Warm API calls from Pakistan were taking ~700 ms. Taking CloudFront, Lambda@Edge and a Function URL out of the request path brought them down to ~260 ms. AI's part was questioning the architecture, not writing the code.

I recently worked on a serverless application where warm API requests from Pakistan were taking around 700 ms. The stack was AWS Lambda behind CloudFront, Lambda@Edge and a Lambda Function URL, with PostgreSQL on Neon.

The admin dashboard made the problem obvious. It fires around 14 API calls in parallel, so every page load paid that 700 ms, and the page wasn’t ready until the slowest call came back.

I could have reached straight for more infrastructure or put caching everywhere. Before that I wanted an answer to one question: where were those 700 ms actually going?

The legacy flow

Every API request took this path:

Client → CloudFront → Lambda@Edge → Lambda Function URL → AWS Lambda → Neon PostgreSQL

Diagram of the legacy invocation flow, about 700 ms per warm request across 5 hops. 1: a user's browser in Pakistan calls the nearest CloudFront edge location. 2: Lambda@Edge processes the request at the edge. 3: a long network hop carries it from the edge to a Lambda Function URL in the US region. 4: the Function URL invokes a warm AWS Lambda instance. 5: Lambda queries PostgreSQL on Neon and responds, and the response returns through the same 5 hops.
fig. legacy flow · 5 hops · ~700 ms warm

The obvious suspects didn’t hold up:

  • Cold starts. The Lambda instances were already warm, so they weren’t behind the baseline latency.
  • Lambda@Edge. It was surprisingly light, with a median execution time of about 1.5 ms.
  • The database. Neon query latency wasn’t where the time was going either.

That left the network itself, specifically the hop between the CloudFront edge and the regional origin in the US.

Using AI as an engineering partner

This is where AI was useful. I didn’t ask it to generate code; I used it to challenge the architecture:

  • What happens to latency when the client is geographically far from the Lambda region?
  • Is CloudFront actually helping this API workload?
  • What are the trade-offs of moving to an API Gateway HTTP API?
  • How would connection reuse affect warm requests?
  • What happens to authentication, CORS, client IP handling and webhooks?
  • How can the migration go in without disrupting production?
Two-step diagnosis. Step A asks where the 700 ms is going: Lambda cold starts were checked (instances were already warm), Lambda@Edge overhead was checked, Neon PostgreSQL query latency was checked, and the CloudFront network path, extra network latency in front of the API, is marked as the primary cause. Step B weighs two options. Keeping CloudFront at the edge keeps edge capabilities but the network overhead stays. A regional API Gateway calling Lambda directly gives fewer hops and connection reuse but gives up edge features; it is marked as chosen. The criteria were latency, connection reuse, cost, complexity, auth and CORS, and client IP.
fig. diagnosis · where the 700 ms went

We compared the options on latency, connection reuse, cost, complexity, authentication, CORS and client IP handling. Keeping CloudFront in front of the API kept the edge features, but the network overhead stayed with it. A regional API Gateway HTTP API → Lambda path meant fewer hops and reused connections, at the cost of those edge features. That’s the one I tested.

The optimized flow

The new request path:

Client → API Gateway HTTP API → AWS Lambda → Neon PostgreSQL

Diagram of the optimized invocation flow, about 260 ms per warm request across 3 hops. 1: the user's browser in Pakistan calls a regional API Gateway HTTP API in the US directly, with no edge detour. 2: API Gateway invokes a warm AWS Lambda instance in 12 to 26 ms. 3: Lambda queries PostgreSQL on Neon and responds. CloudFront, Lambda@Edge and the Function URL are shown struck through as removed from the request path.
fig. optimized flow · 3 hops · ~260 ms warm

CloudFront, Lambda@Edge and the Function URL are all gone from the API request path. In my measurements, the API Gateway → Lambda integration itself took only 12–26 ms. Requests needed some adapting along the way, and the client IP is now read from a trusted source.

Adding API Gateway wasn’t the important part. Removing a network path this API didn’t need was.

Rolling out without gambling with production

I didn’t switch everything at once. The migration went in step by step:

  1. Add API Gateway alongside the existing setup, without touching the CloudFront path.
  2. Configure SST stages and DNS, sequencing deploys so no DNS record is ever deleted.
  3. Validate behavior: authentication, CORS, webhooks, client IP handling and the main application flows.
  4. Move development traffic, but only once validation passed.
  5. Keep production isolated until the new path was proven, with a rollback planned.
Table of the five-step incremental migration, showing whether dev and prod traffic go through the CloudFront path or the API Gateway path. Step 1, add API Gateway alongside with no disruption to the existing CloudFront setup: dev and prod both on CloudFront. Step 2, configure SST stages and DNS, sequencing deploys so no DNS record is deleted: both still on CloudFront. Step 3, validate auth, CORS, webhooks, client IP and app flows: dev under test on API Gateway, prod on CloudFront. Step 4, switch dev traffic only after verification passes: dev on API Gateway, prod on CloudFront. Step 5, keep production isolated until the new path is ready, with rollback planned: dev on API Gateway, prod on CloudFront.
fig. rollout · dev first, production isolated

Because the old path stayed in place the whole time, there was always a clean way back if anything unexpected showed up.

The result

  • Before: ~700 ms per warm request
  • After: ~260 ms per warm request
  • Reduction: ~63%
Before and after comparison. Before: 5 hops, user to CloudFront to Lambda@Edge to Function URL to Lambda to PostgreSQL, about 700 ms. After: 3 hops, user to API Gateway to Lambda to PostgreSQL, about 260 ms, a 63% reduction. Below: the API Gateway to Lambda integration takes 12 to 26 ms; the admin dashboard's roughly 14 parallel calls are noticeably faster; new connections are slower via API Gateway, so connection reuse is what wins; cold starts remain a separate optimization challenge.
fig. before vs. after · −63%

The admin dashboard is where the difference shows most, because it makes so many requests in parallel.

There are trade-offs. New connections can be slower, because the request now goes straight to the AWS region instead of terminating at a nearby CloudFront edge. The gain comes from reusing connections. Cold starts are untouched and remain a separate optimization problem.

So this isn’t a case of one AWS service being better than another. It’s about matching the architecture to the workload, then measuring the result.

What I learned

The biggest lesson wasn’t about CloudFront or API Gateway. It was about how I used AI.

AI didn’t make the architectural decision for me. It helped me:

  • question assumptions
  • explore alternatives
  • identify potential failure cases
  • structure the migration
  • validate the implementation
  • get from investigation to a working solution faster

That’s where I see AI becoming genuinely valuable in software engineering. Not just “write this code”, but “help me reason through this system”.

The goal isn’t to add more technology. It’s to understand the system well enough to know what should be removed, what should stay, and what actually needs to change.

cd ../blog