Context & facts
Get context
Assemble a prompt-ready Markdown block, the end user's known facts plus query-relevant memories, under a token budget. Deterministic retrieval and formatting only; no LLM generation.
/v1/context
SDK: korely.get_context(query=..., ...). This is the one call you
drop into a prompt. Korely cosine-ranks the end user's typed facts, searches
the query-relevant memories, packs both under your token budget, and hands
back a Markdown string you can paste straight into the model's context.
Nothing is generated, the assembly is deterministic. The block also comes in
two parts: stable, the reader note, the same text from one call
to the next, and volatile, which changes with the question.
Authentication
HTTP header, required: Authorization: Bearer kor_live_.... The key must carry the memories:read scope.
Query parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
query | string | Required | The query the context block is retrieved and ranked against, facts are cosine-ranked and memories are searched against it. min_length=1, max_length=2000. |
user_id | string | Optional | The end user to build context for. Max 255 characters. When omitted, facts are not filtered by end user and the memory search runs across all end users of the key's project. Default null. |
agent_id | string | Optional | Agent namespace filter: narrows the facts and the memories to one agent, as on search and list. Default null (every agent). |
token_budget | integer | Optional | Total token budget for the assembled block. The reader note and the facts together are capped at 50% of it; memories fill the rest. ge=50, le=8000. Default 800; we recommend 4000 for an agent that answers questions about the past (see the note on the budget below). |
Example request
curl -G https://api.korely.ai/v1/context \ -H "Authorization: Bearer kor_live_..." \ --data-urlencode "query=What does Giulia want for the weekly sync?" \ --data-urlencode "user_id=customer-giulia-4812" \ --data-urlencode "agent_id=support-bot" \ --data-urlencode "token_budget=800"Response
200 OK. A prompt-ready Markdown block, the same block in its
stable and volatile parts, its estimated token count, and the ordered
source ids that contributed to it. The reader note is shortened to
[...] in the example; the real one is about 180 tokens.
{ "context": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. [...] do not abstain when it is present._\n\n## Known facts\n- The user prefers async standups (since 2026-05-12)\n- The user works_at Acme Corp (since 2026-04-01)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-05-12] Giulia asked to move the weekly sync to Mondays and keep it under 30 minutes.", "stable": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. [...] do not abstain when it is present._", "volatile": "## Known facts\n- The user prefers async standups (since 2026-05-12)\n- The user works_at Acme Corp (since 2026-04-01)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-05-12] Giulia asked to move the weekly sync to Mondays and keep it under 30 minutes.", "stable_hash": "a345b30abc62923c0ac2bd5f87a6682ef3843a8c787ef3a3a96aec91015f7bbe", "tokens": 235, "sources": ["fct_a1", "fct_a2", "mem_8f2c1a"], "degraded": false, "degraded_parts": []}| Field | Type | Description |
|---|---|---|
context | string | The Markdown block, stable and volatile joined by a blank line: the reader-trust note, then ## Known facts (at most 10 facts, ranked by relevance to the query weighted by how recently each was last confirmed, within 50% of the budget shared with the note), then ## Relevant memories with one line on how to read their dates, each memory under the day it was said ([YYYY-MM-DD], the write's timestamp, else when it was stored) and a relative time followed by the date it resolves to (next month [June 2023]), and last a NOTE: when part of the retrieval failed. Empty string when nothing fits. |
stable | string | The head of context that does not depend on the question: the reader-trust note, the same text from one call to the next. It holds no date, no echo of the query and no count that depends on it. Empty when the block is empty or the budget is under 400 tokens. Put it in your system prompt. |
volatile | string | The rest of context: ## Known facts, ## Relevant memories and the closing NOTE:, if any. Put it in the user turn. |
stable_hash | string | SHA-256 (hex) of stable. The same value as on the previous call means the same prefix, so the system prompt your provider cached is still valid. |
tokens | integer | Estimated token count of the context string (a char / 4 heuristic). |
sources | array<string> | Ordered list of the source public ids that contributed to the block: fact ids (fct_) first, then memory ids (mem_). |
degraded | boolean | true when part of the block could not be retrieved the normal way. A partial context that looks complete is worse than an error: check it before trusting an empty answer. |
degraded_parts | array<string> | Which part: facts (the facts are the most recently confirmed ones rather than the most relevant) and memories (the memories could not be retrieved). If the question cannot be embedded, both: facts by recency and no memories. If only the memory search fails, ["memories"]; if only the fact ranking fails, ["facts"]. When the block has any content it closes with a NOTE: saying the same to the model, counted in token_budget. |
Errors
| Status | Code | Cause |
|---|---|---|
401 | invalid_key | The Authorization: Bearer credentials are missing, or the kor_live_ key does not resolve to a live API key. Response carries WWW-Authenticate: Bearer. |
403 | forbidden | The API key does not carry the memories:read scope. |
429 | rate_limit_exceeded | Per-tier fixed-window minute/hour/day rate limit exceeded. Response carries Retry-After, X-RateLimit-Limit, and X-RateLimit-Remaining headers. |
429 | quota_exceeded | Monthly query quota reached (tier queries_per_month plus a 10% grace). Upgrade for more. |
422 | invalid_request | Request validation failed, query is missing or violates min_length=1/max_length=2000, token_budget falls outside ge=50/le=8000, or user_id is over 255 characters. |
Notes
- Read-only. A plain GET, no soft-delete or pagination semantics on this endpoint. It is deterministic retrieval and formatting only; there is no LLM generation.
- Budget split. The reader note and the facts together are capped at 50% of
token_budget(at the default 800, about 215 tokens remain for facts), at most 10 facts, ranked by relevance to the query weighted by how recently each was last confirmed. Memories (the verbatim source) fill the remainder; the first memory that does not fit is cut at a sentence boundary and packing stops. Headings and line breaks count toward the budget. - Which budget. The default 800 keeps the block short for a chat that only needs the latest facts. For questions about what happened and when, use
token_budget=4000: on 60 LongMemEval questions the evidence a question needs was in the block 94% of the time at 4,000 tokens and 80% at 2,000 (October 2026), and the block shows its first memories whole. The cost is on your side: about 4,000 more input tokens per call for your model. - Budget is clamped.
token_budgetis clamped server-side to[50, 8000]even though the query parameter already enforces the same bounds. - Reader-trust note. The leading guidance note (about 180 tokens, counted in the budget) opens
stablewhenever the block has any content, facts or memories, and the budget is at least 400 tokens. Below 400stableis empty and the block starts at## Known facts. - Prompt caching. Model providers cache only an identical prompt prefix. Put
stableat the start of your system prompt andvolatilein the user turn: from one question to the next the system prompt stays byte for byte the same, and the provider can reuse it within its own caching rules. Whenstable_hashmatches the previous call, the prefix is unchanged. - Graceful degradation, said out loud. When the question cannot be embedded, the facts fall back to the most recently confirmed ones and no memories are included; if only the memory search is unavailable, the block holds facts only. Either way the call succeeds,
degradedistrue,degraded_partsnames the parts, and a block with any content closes with aNOTE:for the model, insidevolatile. - End-user scope.
user_idis the end-user scope; omitting it builds context across all end users of the key's project. - Counts as a read. Every call that passes the key and scope checks counts one query against your monthly quota. A call refused with
422for a malformed parameter does not: its query is given back. It still counts against the per-minute, per-hour and per-day rate limits.
Related
- Get context, guide, the narrative walkthrough with context.
- Get facts, the raw typed facts for an end user.
- Search memories, the underlying memory search.
- Add a memory, write the memories this block is built from.