Skip to main content
Reference
Current constraints and practical limits. These reflect the current implementation, not fundamental architectural boundaries.

Active-text mutation admission

The engine checks each text document a write indexes before the write commits. When one document’s final text exceeds a bound, the write fails with HTTP 400 and code active_text_mutation_limit_exceeded, and none of its graph changes commit. The diagnostic reports the measured resource, the document’s usage and its allowance. Each document is checked on its own. The index worker later swaps an entity’s indexed document for its newest one in a single publication, so each document may use at most half of that publication’s row budget. Only the newest document is analyzed, so it gets the whole analysis budget: In practice a document holds up to about 16,000 unique terms, or fewer when terms or tenant values are long, and about 200,000 short tokens. Separately, the queued work one write stages for one index is capped at 8 MiB. That bounds the combined text of every document the write indexes there; a larger write fails with index_operation_batch_too_large and should be split. Index builds apply the same per-document allowances, text analysis included, to existing rows. A longer row blocks the build with oversized_entity. Shorten or delete the row, then retry the build. These are runtime policy limits, not storage-format maxima; embedded configurations can supply a different validated policy. Retrying the same write does not help, and there is no time-based Retry-After: shorten the document instead. See Error handling for the response contract. HTTP 429 rate limiting below is a separate mechanism.

Helix Cloud request rate limits

Helix Cloud applies a distributed token bucket to POST /v2/query. The bucket is scoped to the authenticated Cloud database, so reads and writes from every API key, application instance, and gateway replica draw from the same allowance. Database-specific overrides can change the sustained rate, burst capacity, and query attempt budget. The values assigned to your database take precedence over its plan limits. Each admitted query request costs one token from one shared bucket per Cloud database.

Token-bucket behavior

  • A full bucket can admit requests up to its burst capacity. Tokens then refill continuously at the plan’s sustained rate, up to that capacity. This is not a fixed one-second window.
  • One incoming request consumes one token whether it is a read, write, or cache warming request. Warming fanout and gateway retries do not consume additional tokens.
  • Requests rejected during authentication, gateway header validation, or outer request JSON decoding do not consume a token. Query AST and planner validation happen after admission, so those later validation failures consume one token.
  • The bucket is shared across API keys and gateway replicas. Rotating keys or distributing calls across connections does not create more capacity.
  • Burst capacity controls short-term admission, not the number of queries that can execute concurrently. Bound client concurrency separately.

Rate-limit responses

When no token is available, the gateway rejects the request before database execution:
Retry-After is a whole number of seconds. Wait at least that long before retrying, and add jitter when many workers share the same database. The response does not currently include RateLimit-* or X-RateLimit-* limit, remaining, or reset headers. Current official SDK error objects expose the HTTP status, stable code, and diagnostic, but not response headers. Use direct HTTP or an application transport that retains headers when the exact Retry-After value is required; otherwise use a configured, bounded status/code-aware delay with jitter. If the gateway cannot make a safe distributed rate-limit decision, it fails closed before database execution:
Retry rate_limit_unavailable with bounded exponential backoff and jitter. A 402 tenant_disabled response is an account-credit gate, and a 408 query_timeout response is an execution deadline; neither means the request bucket was exhausted.

Application guidance

  1. Coordinate admission across workers that target the same Cloud database.
  2. Honor Retry-After on rate_limited instead of retrying immediately.
  3. Use bounded exponential backoff with jitter for transient 503 responses.
  4. Bound in-flight concurrency as well as request rate to avoid local queues and latency spikes.
  5. Apply a separate per-user or per-workspace limiter when multiple application tenants share one Cloud database; the Helix bucket does not distinguish those application tenants.

Data model

Vector indexes

Text indexes

Secondary indexes

Queries

Index lifecycle

Embedded runtime