Skip to main content
POST /query returns matching passages (chunks) and source details, not a generated answer. Format these results as context (information supplied alongside the question), then pass it to your language model. The SDK helper below handles the formatting; you do not need to interpret the graph fields to get started.

1. What the response looks like

The SDKs return the full response envelope from POST /query; the retrieval payload is on its data field (e.g. result.data). Raw HTTP responses use the same envelope, with the payload under data. The build_string / buildString helper accepts either the full envelope or just the data payload, so you can pass the SDK return value to it directly. The data payload has the same core shape regardless of type or query_by. Expand this reference when you need individual fields; continue to Section 2 to format a result.
Four things matter for prompt construction:
  • chunks: passages from matching documents or memories. Ranked by relevance; preserve the order HydraDB returns.
  • graph_context.query_paths: chains of relationships connecting named people, services, or topics relevant to the query. See Context Graphs.
  • graph_context.chunk_relations + chunk_id_to_group_ids: per-chunk graph relations grouped by group_id, so you can attach the right triplets to each chunk.
  • additional_context: a map keyed by chunk_uuid. When a chunk includes extra_context_ids, use those IDs to look up related chunks here.
The helper in Section 2 formats passages, source details, and any relationships together.

2. Transforming the response into LLM context

Use build_string / buildString from the SDK. It takes any POST /query result and returns a formatted plain string.

3. Feeding the context into your LLM


4. Combining Knowledge and Memories

The simplest path is one POST /query with type: "all". HydraDB queries both stores in parallel and returns one merged, ranked result set. type: "all" reads both stores from the same scope, so this example assumes the shared documents are in user_123 too. If they live in their own collection, send collections: ["company_docs", "user_123"] instead of collection.
If you need to format knowledge and memories in separate labeled sections, call POST /query twice in parallel:
If the memory query fails or times out, fall back to the knowledge-only prompt.

5. Practical guidance

  • Preserve server order: Don’t re-sort chunks client-side.
  • Start small on chunks: max_results: 10 is a reasonable default. Drop to 5 if you hit token limits, raise to 20 if you rerank downstream.
  • Use graph context selectively: It improves relational queries and bloats simple lookups. See Context Graphs.
  • Give the model a grounding instruction: Use a system prompt like “answer only from the provided context” to reduce unsupported answers when retrieval is thin.
  • Format consistently: Whatever section delimiters you choose (=== CONTEXT ===, Chunk N, Source:), keep them stable across calls so the model learns the structure.

6. Common mistakes


  • Query: request parameters and response shape
  • Memories: what POST /query with type: "memory" queries
  • Knowledge: what POST /query with type: "knowledge" queries
  • Context Graphs: what graph_context fields contain

If you are not using the SDK, use the helper below to call POST /query, format the response, and pass the result to your LLM.