Reference
Current constraints and practical limits. These reflect the current implementation, not
fundamental architectural boundaries.
Active-text mutation admission
The engine checks each text document a write indexes before the write commits. When one document’s final text exceeds a bound, the write fails with HTTP 400 and codeactive_text_mutation_limit_exceeded, and none of its graph changes commit. The
diagnostic reports the measured resource, the document’s usage and its allowance.
Each document is checked on its own. The index worker later swaps an entity’s indexed
document for its newest one in a single publication, so each document may use at most
half of that publication’s row budget. Only the newest document is analyzed, so it gets
the whole analysis budget:
In practice a document holds up to about 16,000 unique terms, or fewer when terms or
tenant values are long, and about 200,000 short tokens.
Separately, the queued work one write stages for one index is capped at 8 MiB. That
bounds the combined text of every document the write indexes there; a larger write fails
with
index_operation_batch_too_large and should be split.
Index builds apply the same per-document allowances, text analysis included, to existing
rows. A longer row blocks the build with oversized_entity. Shorten or delete the row,
then retry the build.
These are runtime policy limits, not storage-format maxima; embedded configurations
can supply a different validated policy. Retrying the same write does not help, and
there is no time-based Retry-After: shorten the document instead. See
Error handling for the response
contract. HTTP 429 rate limiting below is a separate mechanism.
Helix Cloud request rate limits
Helix Cloud applies a distributed token bucket toPOST /v2/query. The bucket is
scoped to the authenticated Cloud database, so reads and writes from every API
key, application instance, and gateway replica draw from the same allowance.
Database-specific overrides can change the sustained rate, burst capacity, and query attempt budget.
The values assigned to your database take precedence over its plan limits. Each
admitted query request costs one token from one shared bucket per Cloud database.
Token-bucket behavior
- A full bucket can admit requests up to its burst capacity. Tokens then refill continuously at the plan’s sustained rate, up to that capacity. This is not a fixed one-second window.
- One incoming request consumes one token whether it is a read, write, or cache warming request. Warming fanout and gateway retries do not consume additional tokens.
- Requests rejected during authentication, gateway header validation, or outer request JSON decoding do not consume a token. Query AST and planner validation happen after admission, so those later validation failures consume one token.
- The bucket is shared across API keys and gateway replicas. Rotating keys or distributing calls across connections does not create more capacity.
- Burst capacity controls short-term admission, not the number of queries that can execute concurrently. Bound client concurrency separately.
Rate-limit responses
When no token is available, the gateway rejects the request before database execution:Retry-After is a whole number of seconds. Wait at least that long before
retrying, and add jitter when many workers share the same database. The response
does not currently include RateLimit-* or X-RateLimit-* limit, remaining, or
reset headers.
Current official SDK error objects expose the HTTP status, stable code, and
diagnostic, but not response headers. Use direct HTTP or an application
transport that retains headers when the exact Retry-After value is required;
otherwise use a configured, bounded status/code-aware delay with jitter.
If the gateway cannot make a safe distributed rate-limit decision, it fails
closed before database execution:
rate_limit_unavailable with bounded exponential backoff and jitter. A
402 tenant_disabled response is an account-credit gate, and a 408
query_timeout response is an execution deadline; neither means the request
bucket was exhausted.
Application guidance
- Coordinate admission across workers that target the same Cloud database.
- Honor
Retry-Afteronrate_limitedinstead of retrying immediately. - Use bounded exponential backoff with jitter for transient
503responses. - Bound in-flight concurrency as well as request rate to avoid local queues and latency spikes.
- Apply a separate per-user or per-workspace limiter when multiple application tenants share one Cloud database; the Helix bucket does not distinguish those application tenants.