Skip to main content
POST
The single retrieval endpoint for everything. Use it any time you need to feed an LLM with grounded context, surface user preferences, or fetch chunks ranked by relevance. Three independent dimensions control behavior:
  • type picks what to query: "knowledge", "memory", or "all" (both, merged and re-ranked together).
  • query_by picks how to match: "hybrid" (semantic + BM25, the default) or "text" (BM25 only; pair with operator).
  • mode picks how to rank results: "fast" (single-pass, low-latency), "thinking" (expands query, reranks, and can include forceful-relation context), or "auto" (scores the query and routes to "fast" or "thinking" automatically, defaulting to "thinking" when the signal is inconclusive; the default if mode is omitted).
Recommended configurations maps common use cases to these three settings.

Querying multiple collections

Use collections when one query should fan out across multiple user, workspace, or team scopes. The field accepts either a list or a weighted object:
Equal weighting
Weighted ranking
A list gives every collection equal normalized weight. An object treats values as positive relative ranking weights with at most one decimal place and normalizes them server-side. You can send at most 100 collections. When max_results is omitted, HydraDB uses up to 10 results per collection, capped at 1000 fanout candidates before the final ranked response is shaped. When max_results is set, it is the final global response cap across the merged fanout result set. Every listed collection must exist, or the call returns 400 INVALID_INPUT naming the missing ones.
Caching tip: collections list order is not semantically significant for fanout selection. Sort list values before constructing cache keys; for weighted objects, sort keys and keep weights at the documented one-decimal precision so equivalent calls share the same cache entry.

Transforming the response into LLM context

Use build_string / buildString from the SDK. It takes any POST /query result and returns a formatted plain string.

Common use-cases and their configurations

HydraDB scores the query before retrieval and routes it to "fast" or "thinking": a query naming several distinct entities like this one is likely to route to "thinking". Use "auto" for traffic where query complexity varies call-to-call and you don’t want to hand-pick per request. This is also the default: an omitted mode field behaves exactly like mode: "auto". Set mode to "fast" or "thinking" explicitly if you want a deterministic pipeline instead.

Request body

Tuning heuristics:
  • alpha: start at 0.8. Lower toward 0.3 to 0.5 when the query contains literal tokens (error codes, SKUs, product names). Raise toward 0.9 for conceptual questions.
  • recency_bias: set to 0 for static reference material. Set 0.2 to 0.4 for mixed content, 0.6 to 0.8 for changelogs, news, or status updates.
  • max_results: start at 10. Drop to 5 for tight context windows; raise to 20 if you rerank downstream.

Decision matrix

For query_by: "hybrid":"auto"’s resolved pipeline isn’t reported back in the response, so budget latency as thinking-level in the worst case. Set mode to "fast" or "thinking" when you need a predictable pipeline.
metadata_filters are hard constraints applied before ranking and re-checked after hydration. The shape combines two filter scopes:
Separate keys are ANDed. Each metadata (top-level) key takes an operator object naming the comparison:
Adding values to contains_any widens the result set. There is no ALL/AND operator within a single key, and range and fuzzy operators are not supported; run multiple queries or post-process client-side for those cases.Operators apply to metadata (top-level keys) only. Inside additional_metadata, use a bare scalar for an exact match or a bare array to match any listed value.
An operator used inside additional_metadata is not rejected. It is read as an exact-match filter against a stored object, so on a normal field it matches nothing and the request returns 200 with an empty result rather than an error.
The bare forms still work and are unchanged, but are deprecated in favour of the operators, because the comparison they perform is inferred from the JSON shape rather than stated. A bare scalar behaves as equals, a bare array as contains_any, and a bare single-element array as contains, so {"emails": "a@x"} and {"emails": ["a@x"]} differ by one character and return different results.
contains, contains_any and lists are supported on VARCHAR fields only: any of them passed for a declared field of another type is rejected with 400 VALIDATION_ERROR. equals works on every declared type, so {"priority": {"equals": 7}} is valid on an INT64 field.A known operator given the wrong operand type, or several operators in one object, is rejected with 400 INVALID_INPUT, and so is a null filter value. A misspelled operator is not: {"contian": "x"} is indistinguishable from a filter for a stored object with that key, so it is left alone and matches nothing.contains, contains_any and equals are reserved key names: an object built only from them is read as an operator and can no longer exact-match a stored object, and an object whose keys are ALL operator names is rejected with 400. Mixing an operator name with any other key ({"contains": "a", "other": 1}) is unaffected. A caller needing the reserved shape must rename the nested key or the field.
A zero-result query returns empty arrays/maps rather than an error, as shown in the Zero results tab.

Behavior notes

Important Considerations & Common Mistakes
  • query_forceful_relations requires mode to resolve to "thinking": In fast mode the flag is silently skipped. The server does not error or warn; your additional_context will simply be empty. Under mode: "auto" this depends on that request’s routing decision, not on what you asked for.
  • Relation timestamp is a Unix epoch float here: In the graph_context slice returned by /query (and in the passthrough relations returned by List Documents with include_fields: ["relations"]), each relation’s timestamp is a Unix epoch value in seconds (a float, e.g. 1778573640.0). The dedicated Context Relations endpoint returns the same field as an ISO-8601 string instead. Normalize before comparing relation timestamps across endpoints.
  • Use the right metadata namespace: Top-level metadata_filters keys match metadata; free-form per-document fields must be nested under additional_metadata (document_metadata is only a legacy alias). Declare hot top-level filter fields in database_metadata_schema.
  • Check indexing first: New documents are not searchable until Ingestion Status reports them indexed.
  • Collection scope: If you omit collection, HydraDB queries the default collection, and a query that names a collection does not read the default one. Use List Collections to discover available IDs.

Errors

Common codes: 400 INVALID_INPUT (empty query, operator: "and" or "phrase" without query_by: "text", a malformed filter operator, or a listed collection that does not exist), 400 VALIDATION_ERROR (a metadata filter that does not fit the declared field type), 404 DATABASE_NOT_FOUND, 422 TENANT_INFRA_NOT_READY (the database is still provisioning; poll Database Status until ready_for_ingestion is true), 500 INTERNAL_ERROR. See Error Responses for the full list. 400 also covers oversized filters: a metadata_filters list above 500 values, or a metadata_filters object above 64 KiB of compact JSON. The message names the offending key or reports the actual byte count. See Filter size limits.
Related Resources

Authorizations

Authorization
string
header
required

API key sent as a Bearer token: "Bearer prefix.secret"

Body

application/json

Unified query request

acl
string[]

ACL scopes retrieval to documents the given principals may access (PRO-1684 document ACLs): a document matches when its stored ACL is empty (unrestricted, pre-RBAC content and connectors without permission support), contains public, or intersects these principals. Entries are bare emails or prefixed principals (user_email:/group:/domain:). Omitted, empty, or ["*"] disables ACL filtering entirely, today's behavior. Like IDs, the resulting clause survives the metadata zero-result retry. An entry that is not a known principal fails CLOSED: it matches only public and unrestricted documents, never restricted.

additional_context
string

Optional context string prepended to the query to improve retrieval relevance.

Example:

"The user is a senior engineer onboarding to the platform."

alpha
default:0.8

Semantic weight in hybrid search: 0 is BM25 only and 1 is semantic only. The default is 0.8; "auto" also resolves to 0.8.

Required range: 0 <= x <= 1
collection
string

Collection scope. Defaults to the default collection when omitted. Formerly sub_tenant_id; the sub_tenant_id alias is still accepted (deprecated).

Example:

"team_docs"

collections

A list of existing collections (equal weights), or an object mapping collection IDs to positive ranking weights with at most one decimal place. Maximum 100 collections. A missing collection returns 400. Use instead of a single collection selector.

Required array length: 1 - 100 elements
Example:
database
string

Database to query. The deprecated tenant_id alias is still accepted.

Example:

"acme_corp"

graph_context
boolean
default:true

Include the graph slice. False takes effect only in fast mode; thinking always includes graph context.

Example:

true

graph_vector_prune
boolean

GraphVectorPrune switches the graph-connected-chunks lane from "fetch graph-selected chunks and let the fusion reranker sort them out" to "fetch a wider graph-selected candidate pool, then rank that pool by Milvus vector similarity, fully replacing the final chunk list." Works in either fast or thinking mode. Default false preserves existing behavior. Also gated server-side by a repo-level config flag (SearchService's graphVectorPruneEnabled) — if that flag is off, this is forced to false regardless of what the request sets, so a deployment can disable the mechanism without any client-side change.

Example:

true

graph_vector_prune_spacy_entities
boolean

GraphVectorPruneSpacyEntities: when GraphVectorPrune is also set, swaps the graph lane's entity-extraction source from the default LLM-based extractor to a local spaCy subprocess (faster, no network round trip, but a narrower/mismatched entity vocabulary versus the graph's own LLM-extracted node names). No-op if GraphVectorPrune is false (including when forced false by the server-level flag) or no spaCy extractor was configured at startup.

Example:

true

ids
string[]

IDs optionally scopes retrieval to specific source ids. The v2 wire field is ids (matching /context/list); empty means search the whole corpus. Applied as a Milvus source_id in [...] pre-filter that is preserved across the metadata zero-result retry, so a source-scoped search that matches nothing returns nothing rather than silently widening to the whole corpus.

Example:
max_results
integer
default:10

Maximum number of chunks to return.

Required range: 1 <= x <= 250
Example:

10

metadata_filters
object

Filters results by source metadata. Top-level keys target tenant metadata (for example department, priority, active, or tags). Nested additional_metadata keys target document metadata. Separate keys are ANDed. A scalar value is an exact match; an array means match ANY one of the listed values (OR) - there is no ALL/AND operator within a single key. Arrays are supported on VARCHAR fields only: an array passed for a declared field of any other type is rejected with 400 VALIDATION_ERROR. Size limits: each list may hold at most 500 values, and the whole metadata_filters object is capped at 64 KiB measured on its compact JSON encoding in UTF-8 bytes (keys and punctuation count). Exceeding either returns 400 naming the offending key or the actual byte count.

Example:
mode
enum<string>
default:auto

fast: one retrieval pass; thinking: expanded retrieval and reranking; auto (default): select fast or thinking per query.

Available options:
fast,
thinking,
auto
Example:

"thinking"

Number of adjacent chunks to pull alongside each matched chunk for additional context.

Example:

3

operator
enum<string>
default:or

BM25 term operator. "and" and "phrase" require query_by: "text"; otherwise the request returns 400.

Available options:
or,
and,
phrase
Example:

"and"

query
string

Natural-language search query.

Example:

"Which mode does the user prefer?"

query_apps
boolean
default:true

Adds app-aware retrieval to knowledge hybrid queries without excluding other knowledge. Send false to disable it.

Example:

true

query_by
enum<string>
default:hybrid

Retrieval method to use for the query.

Available options:
hybrid,
text
Example:

"hybrid"

query_forceful_relations
boolean
default:true

Include declared related knowledge sources in additional_context when mode resolves to thinking. Ignored in fast mode.

Example:

true

recency_bias
number
default:0.4

Recency boost applied to ranking. 0 disables it; higher values favour more recent sources.

Required range: 0 <= x <= 1
Example:

0.2

sub_tenant_id
string
deprecated

Deprecated for /query (since 2.0.1). Use collection for a single scope or collections for multiple. Backwards-compatible and will be removed in a future version. Do not send together with a multi-scope selector.

Example:

"sub_tenant_4567"

sub_tenant_ids
deprecated

Deprecated for /query (since 2.0.1). Use collections instead; it accepts the same list or weighted-object shape. Backwards-compatible and will be removed in a future version. Do not send together with collections.

Required array length: 1 - 100 elements
Example:
temporal_intent
object

TemporalIntent (EXPERIMENTAL) lets the caller supply the classification (mode/window/phrases) directly, bypassing the regex classifier — for agents whose own LLM already understands the query, and for non-English queries. Invalid overrides fall back to the classifier.

Example:
temporal_now
string

TemporalNow optionally anchors "now" for temporal reasoning (ISO-8601). Callers replaying past conversations (or backfilling) must supply it or to-now durations and recency windows resolve against the server's wall clock (LongMemEval measured 0 exact to-now durations from this alone).

temporal_reasoning
boolean

TemporalReasoning activates the temporal read path: the query is classified into a temporal mode (current/as-of/range/upcoming...), matching edge-level temporal facts are resolved from the edge_temporal store and ride back on the response (temporal_facts / temporal_duration / temporal_filter). CONTRACT: chunk ranking is NEVER altered — ON returns the same chunks as OFF; the layer is additive payload + computed answers only (rank shaping measured net-negative on BEAM/LongMemEval/TEMPO; see temporal_filters.go). Optional; ON by default — pass temporal_reasoning:false to disable. Resolved by GetTemporalReasoningOrDefault (ownership rule).

Example:

true

tenant_id
string
deprecated

deprecated: use database

Example:

"tenant_1234"

type
enum<string>
default:knowledge

Store to query: knowledge (default), memory, or all. Both stores use the same collection scope.

Available options:
knowledge,
memory,
all

Response

OK

data
object
Example:
error
object | null

Null on success; an object with code and message on failure.

Example:

null

meta
object
Example:
success
boolean

Whether the request succeeded.

Example:

true