HTTP & APIs¶
Rate limiting, ETags, gzip compression, request correlation IDs, OpenAPI schema generation, a GraphQL endpoint, and idempotency keys.
rate_limit ¶
Rate limiting: an in-memory sliding-window limiter by default, a
per-route throttle() guard, and an opt-in blanket middleware.
The in-memory backend is process-local by design — correct for a single
worker, and a real (if common) limitation once you run multiple processes
or machines, where each would count independently. That's a documented
trade-off, not an oversight: :class:RedisRateLimiter is the opt-in,
shared alternative for when you need distributed limits.
RateLimiter ¶
Bases: ABC
Counts hits against a key within a rolling time window.
hit
abstractmethod
async
¶
Record a hit for key and report whether it's within limit per window seconds.
InMemoryRateLimiter ¶
InMemoryRateLimiter(
*,
clock: Callable[[], float] = time.monotonic,
max_tracked_keys: int = 100000,
)
Bases: RateLimiter
A sliding-window-log limiter, correct and simple, scoped to one process.
Each key keeps a deque of hit timestamps; on every call, timestamps older than the window are dropped before counting. A single lock serializes hits — fine here since each check is a handful of in-memory operations, not I/O.
Bounded to max_tracked_keys distinct keys (default 100,000). An
idle key's now-empty bucket is only ever cleaned up by another
:meth:hit call for that same key -- without a bound, a key space
that includes request-supplied data (:func:throttle's own docstring
suggests f"login:{email}" as an example key, precisely so distinct
accounts get separate limits) would let a flood of distinct
emails/IPs grow this process's memory without limit, one abandoned
entry per attempt. Past the cap, the least-recently-hit key is
evicted to make room for a new one -- the same "may reset a
rarely-used key's window a bit early under sustained load" trade-off
any bounded-memory rate limiter accepts, and a far safer failure mode
than unbounded growth.
Source code in src/zeython/rate_limit.py
RedisRateLimiter ¶
Bases: RateLimiter
A Redis-backed :class:RateLimiter, shared across every process/machine
pointed at the same Redis — the limitation :class:InMemoryRateLimiter's
docstring names. Requires the redis extra (pip install zeython[redis]).
Fixed-window (INCR + EXPIRE), not sliding-window-log like
InMemoryRateLimiter — the standard, well-known Redis rate-limiting
pattern, and its standard trade-off: a client can get up to 2x limit
requests through across a window boundary (e.g. a burst just before the
window resets, then another just after). Accept that, or implement a
sliding-window version yourself if you need the tighter guarantee.
All keys are namespaced under prefix (default "zeython:ratelimit:")
— safe to point at a Redis instance shared with other subsystems.
Source code in src/zeython/rate_limit.py
RateLimitMiddleware ¶
Pure ASGI middleware applying a blanket per-IP limit to every request.
Keyed under a blanket: prefix -- deliberately distinct from
:func:throttle's own default key (f"ip:{client_ip(request)}"),
even though both are per-IP by default. Without that distinction, a
request to a route also protected by :func:throttle (called with
no explicit key=, so it uses that same default) would hit the
exact same counter this middleware just hit for the same request --
double-counting a single request against one shared budget, and
letting ordinary traffic to other routes eat into a route's own
dedicated allowance (e.g. a login endpoint's brute-force protection
silently sharing its counter with plain page views from the same
IP).
Source code in src/zeython/rate_limit.py
RateLimitHeadersMiddleware ¶
Pure ASGI middleware that stamps the standard X-RateLimit-Limit/
X-RateLimit-Remaining/X-RateLimit-Reset headers (as GitHub's,
Stripe's, and Laravel's APIs do) onto every response for which a rate
limiter actually ran -- whether that was a per-route :func:throttle
call or the blanket :class:RateLimitMiddleware.
Both of those store their :class:RateLimitResult on request.state
(backed by scope["state"], read here directly since a plain ASGI
middleware has no Request of its own); a handler that never calls
either leaves no result to report, so no headers are added. Registered
automatically -- and unconditionally, since :func:throttle works
without the blanket middleware being enabled -- by
:class:RateLimitServiceProvider.
Source code in src/zeython/rate_limit.py
RateLimitServiceProvider ¶
Bases: ServiceProvider
Binds a :class:RateLimiter into the container. :class:InMemoryRateLimiter
(process-local) by default.
The limiter is always available for :func:throttle calls in your own
handlers. The blanket, all-routes middleware is opt-in via .env:
RATE_LIMIT_ENABLED— defaultfalseRATE_LIMIT_MAX_REQUESTS— default60RATE_LIMIT_WINDOW_SECONDS— default60
For a shared, distributed limiter, bind :class:RedisRateLimiter
instead of registering this provider::
app.container.singleton(RateLimiter, lambda: RedisRateLimiter(config.get("redis.url")))
Source code in src/zeython/providers.py
client_ip ¶
The connecting client's IP, or "unknown" if the ASGI server didn't report one.
throttle
async
¶
Raise :class:~zeython.exceptions.TooManyRequestsException past limit hits per window seconds.
Defaults to limiting per client IP; pass key to scope it differently
(e.g. f"login:{client_ip(request)}" to namespace it separately from
other throttled endpoints, or f"login:{email}" to limit attempts per
account regardless of source IP). Call at the top of any handler you
want to protect::
async def login(self, request):
await throttle(request, limit=5, window=60) # 5 attempts/minute per IP
Every response -- allowed or rejected -- carries the standard
X-RateLimit-Limit/X-RateLimit-Remaining/X-RateLimit-Reset
headers, added by :class:RateLimitHeadersMiddleware (registered
automatically by :class:RateLimitServiceProvider).
Source code in src/zeython/rate_limit.py
etag ¶
Conditional GETs: an ETag on a cacheable response, and a 304 Not
Modified (empty body) when the client's If-None-Match already
matches it -- for a GET/HEAD response that doesn't change on every
request (a list endpoint between writes, a lookup table), this saves the
client (and the network) the cost of re-downloading a body it already
has. See docs/api-standards.md.
ETagMiddleware ¶
Pure ASGI middleware. Buffers a GET/HEAD response's full
body (so it necessarily holds the whole thing in memory -- fine for a
typical JSON API response, a poor fit in front of a large file
download or a streaming response; scope this middleware to specific
routes, or use :class:~zeython.storage.Storage for downloads
instead of running it in front of everything, if that's a concern),
computes a strong ETag (a SHA-256 hash of the body) for any 200
response at or above minimum_size bytes, and short-circuits to
304 Not Modified when the request's If-None-Match already
matches it.
Source code in src/zeython/etag.py
ETagServiceProvider ¶
Bases: ServiceProvider
Registers :class:ETagMiddleware -- not registered by default::
# main.py
app.register(ETagServiceProvider)
Configurable via .env:
ETAG_MINIMUM_SIZE-- default0(every 200 response gets an ETag). Raise it to skip the hashing cost for tiny responses that aren't worth caching.
Source code in src/zeython/providers.py
gzip ¶
Response compression: wires Starlette's own GZipMiddleware in as an
opt-in provider, configurable via .env instead of a bare
app.add_middleware(GZipMiddleware, ...) call. See docs/api-standards.md.
GzipServiceProvider ¶
Bases: ServiceProvider
Compresses any response at or above GZIP_MINIMUM_SIZE bytes
whose client sent Accept-Encoding: gzip -- not registered by
default::
# main.py
app.register(GzipServiceProvider)
Configurable via .env:
GZIP_MINIMUM_SIZE-- default500. Responses smaller than this aren't worth the CPU cost of compressing.GZIP_COMPRESS_LEVEL-- default9(Starlette's own default; 1 is fastest/least compression, 9 is slowest/most).
Compressing a response over TLS that reflects both attacker-controlled input and a secret in the same body (a page that echoes a query parameter next to a CSRF token or session-bound value, say) opens a BREACH-style side channel -- an attacker who can trigger many requests and observe response sizes can recover the secret byte by byte, since compression ratio leaks whether a guessed byte matched. This requires that specific reflection pattern to be exploitable, not something gzip alone causes, but it's worth knowing before turning this on for a response that might ever carry both.
Source code in src/zeython/providers.py
request_id ¶
Correlation IDs for tracing one request across a busy log stream.
A single request can produce several log lines -- received, a DB query,
maybe a background job dispatched, the response sent -- and without
something tying them together, reconstructing "what happened for this one
request" from concurrent traffic is guesswork. RequestIdMiddleware
stamps every request/response pair with an ID (the caller's own
X-Request-ID if it sent one, so a request can be traced across
multiple services that all honor the same header, otherwise a fresh
UUID4) and makes it available to every log line emitted while handling
that request.
RequestIdLogFilter ¶
Bases: Filter
Adds the current request's ID to every LogRecord as request_id.
Installed on the root logger's handlers automatically by
:class:RequestIdServiceProvider, so %(request_id)s is available to
any formatter -- "-" for a log line emitted outside of a request
(startup, a scheduled job) rather than a missing attribute. Deliberately
attached to each handler, not the logger itself: a Logger.filter
only runs for records that logger originates directly, never for ones
propagated up from a child logger (uvicorn.error, your own
__name__ loggers) -- which is most of them.
RequestIdMiddleware ¶
Pure ASGI middleware: stamps every request/response with a correlation ID.
Sets the ID as a contextvar for the duration of the request -- readable
via :func:request_id from any code running while handling it,
including a logging filter -- and echoes it back as a response header
so a client can correlate its own logs with the server's.
Source code in src/zeython/request_id.py
RequestIdServiceProvider ¶
Bases: ServiceProvider
Registers :class:RequestIdMiddleware and wires up its logging context.
Registered by default in a generated project -- unlike
:class:~zeython.security_headers.SecurityHeadersServiceProvider, there
is no application-specific default to get wrong here: it only adds a
response header and a piece of log context, never changes what a
request is allowed to do.
REQUEST_ID_HEADER-- defaultX-Request-ID.
Source code in src/zeython/providers.py
openapi ¶
OpenAPI spec generation and an interactive Swagger UI, built from the app's actually-registered routes.
This is deliberately not FastAPI-style automatic request/response
validation from type hints -- Zeython's handlers take a plain request
and call request.json() themselves, and models validate via __rules__
(see docs/validation.md), not typed request models. Rearchitecting that to
get automatic schema inference would be a much bigger, separate change,
and would fight the framework's existing conventions rather than describe
them. What this module does instead: read the routes actually registered
on the router (the same technique zeython.mcp.introspect.describe_routes
uses) and build an OpenAPI document from them, enriched by an optional
@describe(...) decorator for handlers that want a real summary, tags, or
request/response schema instead of a generic placeholder.
OpenApiServiceProvider ¶
OpenApiServiceProvider(
app: Application,
*,
title: str,
version: str = "1.0.0",
description: str | None = None,
)
Bases: ServiceProvider
Serves a generated OpenAPI document and an interactive Swagger UI.
Not registered by default — opt in once your routes exist (register
after RouteServiceProvider, order otherwise doesn't matter)::
app.register(RouteServiceProvider(app, modules=("routes.web",)))
app.register(OpenApiServiceProvider(app, title="My API"))
.env:
OPENAPI_ENABLED— defaulttrueonce registered; setfalseto register the provider (sogenerate_openapiis still usable programmatically) without exposing the two routes, e.g. in production.OPENAPI_JSON_PATH— default/openapi.jsonOPENAPI_DOCS_PATH— default/docs
Swagger UI loads from a CDN — unlike the Tailwind Play CDN (see docs/frontend.md), this is a static asset bundle, not something recompiled per request, so there's no dev-only caveat for using it in production too.
Source code in src/zeython/openapi.py
describe ¶
describe(
*,
summary: str | None = None,
tags: list[str] | None = None,
request_body: dict[str, Any] | None = None,
responses: dict[int, dict[str, Any]] | None = None,
) -> Callable[[Endpoint], Endpoint]
Attach OpenAPI metadata to a route handler, read by :func:generate_openapi.
Without this, a route still appears in the generated spec — just with a generic "Successful response" 200 and no summary/tags::
@describe(
summary="List posts",
tags=["posts"],
responses={200: {"description": "The post list", "content": {
"application/json": {"schema": {"type": "array", "items": model_schema(Post)}},
}}},
)
async def index(self, request): ...
Works on both function-based routes and Controller methods —
attaches to the underlying function, which bound-method attribute
lookup falls through to automatically.
Source code in src/zeython/openapi.py
model_schema ¶
A JSON Schema object for model_cls's mapped columns, for use in
@describe(request_body=...)/responses=....
Excludes model_cls.__hidden__ (e.g. password_hash) automatically
— the same fields :meth:~zeython.db.Model.to_dict never serializes
shouldn't show up as "here's the shape of this response" either. This
is documentation, not enforcement: nothing checks an actual response
body against it.
Source code in src/zeython/openapi.py
generate_openapi ¶
generate_openapi(
app: Application,
*,
title: str,
version: str = "1.0.0",
description: str | None = None,
) -> dict[str, Any]
Build an OpenAPI 3.0 document from every route currently on app.router.
Reads the router directly (the same technique
zeython.mcp.introspect.describe_routes uses), not a separately
maintained route list — the spec always reflects what's actually
registered, not what the source looks like it registers.
Routes mounted via Router.mount()/.include() (StaticFiles,
a nested sub-router) aren't recursed into and don't appear — this
covers directly-registered routes only, a v1 scope limitation, not an
oversight.
Source code in src/zeython/openapi.py
graphql ¶
A GraphQL endpoint built on graphql-core, the pure-Python reference GraphQL implementation.
Zeython owns the transport: the /graphql endpoint, request parsing,
error formatting, and an optional interactive GraphiQL UI in debug mode.
Your app owns the schema -- built however graphql-core lets you build
one (programmatically with GraphQLObjectType, or from SDL via
graphql.build_schema() with resolvers attached afterward). There's no
FastAPI/Strawberry-style automatic type-hint-to-schema generation here,
for the same reason zeython.openapi doesn't auto-generate request
models: it would fight the framework's existing conventions rather than
describe them.
graphql-core is an optional dependency (pip install
zeython[graphql]), not a default one -- most apps don't need a GraphQL
API alongside their REST one.
GraphQLServiceProvider ¶
GraphQLServiceProvider(
app: Application,
*,
schema: GraphQLSchema,
path: str = "/graphql",
graphiql: bool | None = None,
max_depth: int | None = DEFAULT_MAX_QUERY_DEPTH,
max_tokens: int | None = DEFAULT_MAX_QUERY_TOKENS,
)
Bases: ServiceProvider
Serves a GraphQL endpoint for a schema you provide.
# main.py
from zeython import Application, GraphQLServiceProvider
from app.graphql.schema import schema # a graphql.GraphQLSchema you build
app = Application()
app.register(GraphQLServiceProvider(app, schema=schema))
A single route (default /graphql) handles both verbs: POST
executes a query/mutation from a JSON body (query, optional
variables/operationName); GET serves an interactive
GraphiQL UI, when enabled.
Every resolver's info.context is a dict with request (the
current starlette.requests.Request) and container (the app's
:class:~zeython.container.Container, for pulling out anything else
a resolver needs -- a service, current_user(), the database
session) -- the same shape a request handler already has, just handed
to resolvers instead of read off the request directly.
graphiql defaults to self.config.debug -- on locally, off in
production, the same default zeython.openapi's Swagger UI and the
HTML debug error page use, since it exposes your full schema and lets
anyone run arbitrary queries against it.
Every query is rejected before execution if it nests deeper than
max_depth (default :data:DEFAULT_MAX_QUERY_DEPTH) or parses to
more than max_tokens (default :data:DEFAULT_MAX_QUERY_TOKENS) --
graphql-core enforces neither on its own, and without them a single
deeply-nested or highly-aliased query can make the server do
exponentially more work than the request itself costs to send. Pass
max_depth=None/max_tokens=None to disable either check.
Source code in src/zeython/graphql.py
execute_graphql
async
¶
execute_graphql(
schema: GraphQLSchema,
*,
query: str,
variables: dict[str, Any] | None = None,
operation_name: str | None = None,
context_value: Any = None,
max_depth: int | None = DEFAULT_MAX_QUERY_DEPTH,
max_tokens: int | None = DEFAULT_MAX_QUERY_TOKENS,
) -> dict[str, Any]
Execute a single GraphQL query/mutation against schema.
Returns a {"data": ..., "errors": [...]} body per the GraphQL-over-HTTP
spec -- errors is only present when at least one occurred.
max_depth rejects a query nested deeper than that (None disables
the check) and max_tokens bounds how much the parser will read
before giving up -- both guard against the query-complexity denial-of
-service shape GraphQL is notorious for, since graphql-core itself
enforces neither. Either rejection is returned as a normal GraphQL
error response (never a raised exception), matching how any other
validation failure is reported.
Raises ImportError with a pip install zeython[graphql] hint if
graphql-core isn't installed.
Source code in src/zeython/graphql.py
idempotency ¶
Idempotency keys: replay a mutating request's first response instead of
running it again, for a client that must safely retry a POST/PUT/
PATCH/DELETE it can't tell succeeded or not -- a dropped
connection, a timeout -- without double-charging a card, double-sending
an email, or creating a duplicate row.
Opt-in per request, the same way Stripe's API works: a request without an
Idempotency-Key header is never touched. Built on
:class:zeython.cache.Cache (InMemoryCache by default, RedisCache
for a shared store across processes/machines) -- the same abstraction
application code already uses for caching, not a bespoke storage backend.
IdempotencyMiddleware ¶
IdempotencyMiddleware(
app: Any,
*,
cache: Cache,
methods: Iterable[str] = DEFAULT_METHODS,
ttl: float = DEFAULT_TTL,
header: str = DEFAULT_HEADER,
)
Pure ASGI middleware. See :class:IdempotencyServiceProvider for the
usual way to register this.
A request without the configured header (default Idempotency-Key),
or whose method isn't in methods (default POST/PUT/
PATCH/DELETE -- the ones that aren't already naturally
idempotent), passes straight through untouched.
A first-seen key runs the request normally and stores its response
(status, headers, body) under that key, scoped to this method, path,
and caller (see :func:_identity_fingerprint -- whatever
Authorization/Cookie header the request carries, so two
different callers reusing the same key value can never see each
other's cached response) -- unless that response is a 5xx, which
is never cached: a transient server error isn't a completed
operation, and locking a retry into replaying the same failure until
ttl expires would defeat the entire point of retrying. A
2xx/4xx response is
exactly as deterministic-and-final as an idempotency key is meant to
protect, so those are cached and replayed. A repeated key within
ttl replays the stored response verbatim instead of running the
request again, adding an Idempotency-Replayed: true response
header so a client (or your own logs) can tell the two cases apart.
A repeated key whose request body doesn't match the first request's
raises :class:~zeython.exceptions.ConflictException (409) rather
than silently returning a stale response for what might be a
different operation that reused the same key by mistake.
A repeated key that arrives while the first request with that key is
still being processed waits for it to finish, then replays its
result, instead of running the operation a second time in parallel --
correct within one worker process. Across multiple processes or
machines, only the stored result is shared (via
:class:~zeython.cache.RedisCache); two processes racing on a
brand-new key can both start processing it before either finishes --
the same in-process-only limitation :class:~zeython.rate_limit.RateLimiter
and :class:~zeython.cache.Cache already document for their default
backends, not something new here.
Source code in src/zeython/idempotency.py
IdempotencyServiceProvider ¶
IdempotencyServiceProvider(
app: Any,
*,
cache: Cache | None = None,
methods: Iterable[str] = DEFAULT_METHODS,
ttl: float = DEFAULT_TTL,
header: str = DEFAULT_HEADER,
)
Bases: ServiceProvider
Registers :class:IdempotencyMiddleware::
app.register(IdempotencyServiceProvider(app))
Uses its own process-local :class:~zeython.cache.InMemoryCache by
default -- pass cache= to share a :class:~zeython.cache.RedisCache
with the rest of the app instead, so a replay works across every
process/machine, not just the one that handled the original request.