Korely

Context & facts

Get context

Assemble a prompt-ready Markdown block, the end user's known facts plus query-relevant memories, under a token budget. Deterministic retrieval and formatting only; no LLM generation.

GET /v1/context

SDK: korely.get_context(query=..., ...). This is the one call you drop into a prompt. Korely cosine-ranks the end user's typed facts, searches the query-relevant memories, packs both under your token budget, and hands back a Markdown string you can paste straight into the model's context. Nothing is generated, the assembly is deterministic. The block also comes in two parts: stable, the reader note, the same text from one call to the next, and volatile, which changes with the question.

Authentication

HTTP header, required: Authorization: Bearer kor_live_.... The key must carry the memories:read scope.

Query parameters

ParameterTypeRequiredDescription
querystringRequiredThe query the context block is retrieved and ranked against, facts are cosine-ranked and memories are searched against it. min_length=1, max_length=2000.
user_idstringOptionalThe end user to build context for. Max 255 characters. When omitted, facts are not filtered by end user and the memory search runs across all end users of the key's project. Default null.
agent_idstringOptionalAgent namespace filter: narrows the facts and the memories to one agent, as on search and list. Default null (every agent).
token_budgetintegerOptionalTotal token budget for the assembled block. The reader note and the facts together are capped at 50% of it; memories fill the rest. ge=50, le=8000. Default 800; we recommend 4000 for an agent that answers questions about the past (see the note on the budget below).

Example request

Terminal window
curl -G https://api.korely.ai/v1/context \
-H "Authorization: Bearer kor_live_..." \
--data-urlencode "query=What does Giulia want for the weekly sync?" \
--data-urlencode "user_id=customer-giulia-4812" \
--data-urlencode "agent_id=support-bot" \
--data-urlencode "token_budget=800"

Response

200 OK. A prompt-ready Markdown block, the same block in its stable and volatile parts, its estimated token count, and the ordered source ids that contributed to it. The reader note is shortened to [...] in the example; the real one is about 180 tokens.

{
"context": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. [...] do not abstain when it is present._\n\n## Known facts\n- The user prefers async standups (since 2026-05-12)\n- The user works_at Acme Corp (since 2026-04-01)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-05-12] Giulia asked to move the weekly sync to Mondays and keep it under 30 minutes.",
"stable": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. [...] do not abstain when it is present._",
"volatile": "## Known facts\n- The user prefers async standups (since 2026-05-12)\n- The user works_at Acme Corp (since 2026-04-01)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-05-12] Giulia asked to move the weekly sync to Mondays and keep it under 30 minutes.",
"stable_hash": "a345b30abc62923c0ac2bd5f87a6682ef3843a8c787ef3a3a96aec91015f7bbe",
"tokens": 235,
"sources": ["fct_a1", "fct_a2", "mem_8f2c1a"],
"degraded": false,
"degraded_parts": []
}
FieldTypeDescription
contextstringThe Markdown block, stable and volatile joined by a blank line: the reader-trust note, then ## Known facts (at most 10 facts, ranked by relevance to the query weighted by how recently each was last confirmed, within 50% of the budget shared with the note), then ## Relevant memories with one line on how to read their dates, each memory under the day it was said ([YYYY-MM-DD], the write's timestamp, else when it was stored) and a relative time followed by the date it resolves to (next month [June 2023]), and last a NOTE: when part of the retrieval failed. Empty string when nothing fits.
stablestringThe head of context that does not depend on the question: the reader-trust note, the same text from one call to the next. It holds no date, no echo of the query and no count that depends on it. Empty when the block is empty or the budget is under 400 tokens. Put it in your system prompt.
volatilestringThe rest of context: ## Known facts, ## Relevant memories and the closing NOTE:, if any. Put it in the user turn.
stable_hashstringSHA-256 (hex) of stable. The same value as on the previous call means the same prefix, so the system prompt your provider cached is still valid.
tokensintegerEstimated token count of the context string (a char / 4 heuristic).
sourcesarray<string>Ordered list of the source public ids that contributed to the block: fact ids (fct_) first, then memory ids (mem_).
degradedbooleantrue when part of the block could not be retrieved the normal way. A partial context that looks complete is worse than an error: check it before trusting an empty answer.
degraded_partsarray<string>Which part: facts (the facts are the most recently confirmed ones rather than the most relevant) and memories (the memories could not be retrieved). If the question cannot be embedded, both: facts by recency and no memories. If only the memory search fails, ["memories"]; if only the fact ranking fails, ["facts"]. When the block has any content it closes with a NOTE: saying the same to the model, counted in token_budget.

Errors

StatusCodeCause
401invalid_keyThe Authorization: Bearer credentials are missing, or the kor_live_ key does not resolve to a live API key. Response carries WWW-Authenticate: Bearer.
403forbiddenThe API key does not carry the memories:read scope.
429rate_limit_exceededPer-tier fixed-window minute/hour/day rate limit exceeded. Response carries Retry-After, X-RateLimit-Limit, and X-RateLimit-Remaining headers.
429quota_exceededMonthly query quota reached (tier queries_per_month plus a 10% grace). Upgrade for more.
422invalid_requestRequest validation failed, query is missing or violates min_length=1/max_length=2000, token_budget falls outside ge=50/le=8000, or user_id is over 255 characters.

Notes

  • Read-only. A plain GET, no soft-delete or pagination semantics on this endpoint. It is deterministic retrieval and formatting only; there is no LLM generation.
  • Budget split. The reader note and the facts together are capped at 50% of token_budget (at the default 800, about 215 tokens remain for facts), at most 10 facts, ranked by relevance to the query weighted by how recently each was last confirmed. Memories (the verbatim source) fill the remainder; the first memory that does not fit is cut at a sentence boundary and packing stops. Headings and line breaks count toward the budget.
  • Which budget. The default 800 keeps the block short for a chat that only needs the latest facts. For questions about what happened and when, use token_budget=4000: on 60 LongMemEval questions the evidence a question needs was in the block 94% of the time at 4,000 tokens and 80% at 2,000 (October 2026), and the block shows its first memories whole. The cost is on your side: about 4,000 more input tokens per call for your model.
  • Budget is clamped. token_budget is clamped server-side to [50, 8000] even though the query parameter already enforces the same bounds.
  • Reader-trust note. The leading guidance note (about 180 tokens, counted in the budget) opens stable whenever the block has any content, facts or memories, and the budget is at least 400 tokens. Below 400 stable is empty and the block starts at ## Known facts.
  • Prompt caching. Model providers cache only an identical prompt prefix. Put stable at the start of your system prompt and volatile in the user turn: from one question to the next the system prompt stays byte for byte the same, and the provider can reuse it within its own caching rules. When stable_hash matches the previous call, the prefix is unchanged.
  • Graceful degradation, said out loud. When the question cannot be embedded, the facts fall back to the most recently confirmed ones and no memories are included; if only the memory search is unavailable, the block holds facts only. Either way the call succeeds, degraded is true, degraded_parts names the parts, and a block with any content closes with a NOTE: for the model, inside volatile.
  • End-user scope. user_id is the end-user scope; omitting it builds context across all end users of the key's project.
  • Counts as a read. Every call that passes the key and scope checks counts one query against your monthly quota. A call refused with 422 for a malformed parameter does not: its query is given back. It still counts against the per-minute, per-hour and per-day rate limits.

Related