Skip to main content

Command Palette

Search for a command to run...

"FastAPI Is Fast." So Why Is Your API Slow?

A fast framework doesn't guarantee a fast app. Here's what actually decides your API's performance, plus a 30-second async quiz.

Updated
โ€ข4 min readโ€ขView as Markdown
"FastAPI Is Fast." So Why Is Your API Slow?

Quick question before we start:

If FastAPI is one of the fastest Python frameworks, why do so many FastAPI apps still feel slow in production?

Take a second to think about it. I'll wait. โ˜•


The uncomfortable truth

"FastAPI is fast" is a statement about the framework. "My API is fast" is a statement about your application.

They're not the same thing.

Fast framework + poor code = potentially poor application

Your app's speed depends on everything it touches: the database, the network, external APIs, CPU, memory, and (if you're building AI products) your LLM calls. The framework can't fix any of that for you.

Pop quiz ๐Ÿง 

Both endpoints below return the same response. Which one handles many users better?

import asyncio, time
from fastapi import FastAPI

app = FastAPI()

@app.get("/a")
async def endpoint_a():
    time.sleep(2)            # ๐Ÿ˜ด
    return {"ok": True}

@app.get("/b")
async def endpoint_b():
    await asyncio.sleep(2)   # ๐Ÿ™‚
    return {"ok": True}

Answer: B.

time.sleep() inside an async def blocks the event loop, so every other request waits. await asyncio.sleep() hands control back, so the server can work on other requests in the meantime.

Same framework. Same output. Completely different behavior under load. (Bonus: a plain def endpoint runs in a threadpool, so blocking there hurts less. But that's a story for another post.)

Did you get it right? Tell me in the comments. ๐Ÿ‘‡

Syntax is the easy part

Most of us start here:

@app.get("/")
async def root():
    return {"message": "Hello"}

It works. But it hides the questions that decide whether your app survives production:

  • What actually happens when a request hits your server?

  • What's running between Uvicorn and your endpoint?

  • Who validates the data, and when?

  • How are dependencies resolved?

  • Why is one implementation faster than another?

Once you understand what your code is doing and why, you write better code.

What's under the hood?

Your Application
      โ†“
   FastAPI      โ† developer-friendly API layer
      โ†“
  Starlette     โ† web / ASGI toolkit underneath
      โ†“
    ASGI        โ† the interface between server and app
      โ†“
  Uvicorn       โ† the ASGI server
      โ†“
     OS

FastAPI isn't magic. It's a clever layer over solid pieces, and knowing those pieces is how you debug the "why is this slow?" moments.

Demo code vs production code

A demo is simple:

Request โ†’ Endpoint โ†’ Logic โ†’ Response

Production is not:

Request โ†’ Routing โ†’ Auth โ†’ Validation โ†’ Dependency resolution
โ†’ Business logic โ†’ DB / AI / External API โ†’ Error handling
โ†’ Serialization โ†’ Response โ†’ Logging & Monitoring

Every arrow is a place where performance can quietly leak away: unnecessary external calls, inefficient queries, repeated processing.

"But I'm building AI apps, not backends"

Plot twist: AI engineering doesn't replace backend engineering. It depends on it.

Here's what a simple chat endpoint really does:

POST /chat
 โ”œโ”€โ”€ Authenticate the user
 โ”œโ”€โ”€ Validate the request
 โ”œโ”€โ”€ Load the conversation
 โ”œโ”€โ”€ Retrieve context from a vector DB (RAG)
 โ”œโ”€โ”€ Call the LLM
 โ”œโ”€โ”€ Save the conversation
 โ””โ”€โ”€ Return the response

A model alone isn't a product. Login, auth, storage, and error handling are what turn it into one.

Why FastAPI for this?

  • Built for APIs: one backend can serve React, mobile apps, Vue, or other services

  • Performance foundation: async + ASGI from the start

  • Types and validation: Python type hints and Pydantic do the heavy lifting

  • Auto docs: OpenAPI comes for free

  • Dependency injection: reusable, testable building blocks

  • One language: the same Python ecosystem for ML and the API around it

It's a strong pick, but not the only pick. Flask is minimal and flexible. Django is batteries-included. Choose for your problem, not for the hype.

Bonus: one habit that saves headaches

Use isolated environments for every project. With uv:

uv venv
uv pip install fastapi

And treat the environment as disposable: recreate it from your dependency files (pyproject.toml / uv.lock) instead of copying .venv around.

The formula

Good framework
+ Good architecture
+ Good code
+ Correct infrastructure
= Good production system

FastAPI gives you the first one for free. The other three are on you.

Keep digging


Your turn ๐Ÿ’ฌ

What's the slowest thing you've ever shipped in a FastAPI (or any Python) app, and what finally fixed it?

Drop it in the comments. I'm collecting war stories for the next post. ๐Ÿ‘€

If this helped, a โค๏ธ and a follow means a lot. More "under the hood" posts are on the way.

A

The async sleep comparison makes the event-loop problem easy to see. Moving blocking work to a plain def endpoint is useful, but I would measure thread-pool saturation too: it changes where requests wait rather than making the blocking operation disappear.

For the chat pipeline, a load test should separate connection-pool wait, model latency and response serialization, while recording event-loop lag. A bounded inference queue with an explicit overload response would keep many slow model calls from occupying the application indefinitely. That gives the framework benchmark a clear boundary from the end-to-end service behaviour you are describing.

Ship It: Full-Stack Development with Generative AI

Part 1 of 1

A hands-on series on building full-stack applications powered by generative AI. We go from a blank repo to a production-ready app: React frontend, FastAPI backend, PostgreSQL, LLM integration, RAG, and deployment on AWS. No toy demos. Each post covers the why behind the design decisions, not just the code.