Core operations
Get context
Retrieve a single, prompt-ready context block assembled from the most relevant memories and current facts, one call, right before your LLM generates.
get_context is the read call an agent makes immediately before sending a prompt to
an LLM. It searches across all stored memories and current facts for the given query,
assembles them into a single formatted string that fits inside your token budget, and
returns the source ids so you can surface citations. The context string is ready to drop
directly into a system prompt or user message, no post-processing required.
get_context is the primary recall path, and the one place Korely's moat shows up in a
single call. It does not just rank raw text: it assembles the user's currently-valid typed
facts, (subject, predicate, object) triples with bi-temporal validity, so superseded
facts are excluded and only what holds right now is included, alongside the most relevant memories. The
returned sources array mixes fact ids (fct_) and memory ids (mem_),
showing exactly which assembled the block.
Unlike search, which runs semantic vector retrieval and
returns a list of raw memory hits, get_context ranks, deduplicates, and formats the most
relevant material, facts first, into a cohesive block. Use it when you want memory with zero glue code.
Use search when you need the raw results to apply your own ranking or filtering logic.
flowchart LR
A([Agent]) -->|"GET /v1/context?query=..."| B[Korely API]
B --> C[(Memories)]
B --> D[(Facts)]
C --> E[Rank + dedup]
D --> E
E --> F[Format to token budget]
F -->|"{ context, tokens, sources }"| A
A -->|"system prompt + context"| G([LLM]) Request
Endpoint: GET /v1/context. SDK: korely.get_context(query, *, user_id=None, agent_id=None, token_budget=800).
| Param | Type | Notes |
|---|---|---|
query | string | Required, 1-2,000 characters. The text to retrieve relevant context for. Typically the user's latest message or the topic your agent is about to reason over. |
user_id | string | Optional, max 255 characters. Scopes retrieval to one end user's facts and memories. Pass your app's user identifier (e.g. customer-4812). If omitted, facts are not filtered by end user and memories are searched across all end users of the key's project. |
agent_id | string | Optional. Narrows the facts and the memories to one agent namespace, as on search and list. |
token_budget | integer | Optional, integer from 50 to 8,000. Maximum approximate token count for the returned context string. Default: 800. Korely packs the highest-scoring content that fits within this budget. Lower values keep context tight for small models; raise it for more coverage. |
Example
from korely_memory import Korely
korely = Korely(api_key="kor_live_...")
# Call right before you build the LLM promptresult = korely.get_context( query="What are this user's dietary preferences?", user_id="customer-4812", token_budget=600,)
system_prompt = f"""You are a helpful nutrition assistant.
Relevant context about this user:{result.context}
Answer using the context above when relevant."""
# result.tokens → 230# result.sources → ["fct_b91e", "fct_c30a", "mem_8f2c1a"]Response
{ "context": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. Ground your answer in this context. [...] do not abstain when it is present._\n\n## Known facts\n- The user allergic_to tree nuts (since 2026-06-11)\n- The user avoids gluten (since 2026-06-11)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-06-11] User is vegetarian, avoids gluten and is allergic to tree nuts.", "stable": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. Ground your answer in this context. [...] do not abstain when it is present._", "volatile": "## Known facts\n- The user allergic_to tree nuts (since 2026-06-11)\n- The user avoids gluten (since 2026-06-11)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-06-11] User is vegetarian, avoids gluten and is allergic to tree nuts.", "stable_hash": "a345b30abc62923c0ac2bd5f87a6682ef3843a8c787ef3a3a96aec91015f7bbe", "tokens": 230, "sources": [ "fct_b91e", "fct_c30a", "mem_8f2c1a" ], "degraded": false, "degraded_parts": []}| Field | Type | Description |
|---|---|---|
context | string | A formatted, prompt-ready Markdown block: a note for the model that reads it (about 180 tokens, shortened in the example above; present whenever the block has content and token_budget is at least 400), then ## Known facts with the active facts, then ## Relevant memories, each under the day the memory was said ([YYYY-MM-DD]). It is stable and volatile joined by a blank line. Ready to embed directly into a system or user message. |
stable | string | The note: the part that does not depend on the question, the same text from one call to the next. Put it in the system prompt, where your provider's prompt cache reuses it. |
volatile | string | What the question brings: ## Known facts, ## Relevant memories, and a closing NOTE: when part of the retrieval failed. Put it in the user turn. |
stable_hash | string | SHA-256 (hex) of stable. Unchanged from the previous call means the cached system prompt is still valid. |
tokens | integer | Approximate token count of the returned context string. Always at or below the requested token_budget. |
sources | string[] | Ids of the memories and facts included in the context block. Use these to power citations or to retrieve full objects via search or GET /v1/memories/{id}. |
degraded | boolean | true when part of the block could not be retrieved the normal way. A partial context that looks complete is worse than an error: check it before trusting an empty answer. |
degraded_parts | string[] | Which part: facts (ranked by recency instead of relevance) and memories (not retrieved). If the question cannot be embedded, both. |
Prompt caching
Model providers cache only an identical prompt prefix. stable
holds no date, no echo of the query and no count that depends on it, so
from one question to the next it is byte for byte the same. Put it at the
start of your system prompt and volatile in the user turn, and
the provider can reuse the cached prefix on every turn, within its own
rules (a minimum prompt length, a cache lifetime, on some providers an
explicit cache marker). The same stable_hash as on the
previous call means the same prefix.
[ {"role": "system", "content": "You are a helpful nutrition assistant.\n\n<stable>"}, {"role": "user", "content": "<volatile>\n\nWhat can I cook tonight?"}]Errors
| Status | Code string | When it happens |
|---|---|---|
401 | invalid_key | The Authorization header is missing, malformed, or the key has been revoked. Re-authenticate and retry. |
403 | forbidden | The key lacks the memories:read scope. |
422 | invalid_request | Request validation failed: query missing or over 2,000 characters, token_budget outside 50-8,000 or not an integer, or user_id over 255 characters. Like every Korely error, the body is the flat envelope {"code": "invalid_request", "message": "query.query: Field required"}. |
429 | quota_exceeded | Your monthly query quota (including the +10% grace) is exhausted. Upgrade your plan or wait for the reset. The monthly-quota 429 does not carry a Retry-After header; only the per-minute, per-hour or per-day rate-limit 429 does, as integer seconds. |
429 | rate_limit_exceeded | Too many requests in the current minute, hour or day. Wait the seconds in Retry-After. |
Notes
- Idempotency.
GET /v1/contextis a pure read, calling it multiple times with the same parameters produces the same result (assuming no new memories have been written). It is safe to retry on transient network errors without risk of duplicate writes. - Scoping. Pass
user_idin customer-facing agents to restrict retrieval to one end user. A call withoutuser_iddraws from all end users of the key's project, which is correct for an internal ops agent and wrong for a per-customer chat.agent_idnarrows the facts and the memories to one agent, on top ofuser_id. - Token budget. The budget is an approximation based on a 4-character-per-token estimate. The actual token count your LLM sees may differ by a few percent depending on the model's tokenizer. Leave a 10-20% buffer below your model's hard context limit.
- Rate limiting.
get_contextcounts against your monthly query quota, the same pool assearch. Hobby plans include 25k queries per month; Developer 250k; Team 1M; Scale 10M. It also counts toward the per-minute, per-hour and per-day rate limits of your plan (429 rate_limit_exceededwith aRetry-Afterheader), and hitting the monthly cap returns429 quota_exceeded. - Empty context. If no memories or facts are stored for the given scope,
context,stableandvolatileare empty strings,tokensis0, andsourcesis an empty array. This is a valid 200 response, not an error. Agents should handle the empty-context case gracefully by generating a neutral reply rather than surfacing an error to the user.
Related
- Search memories, returns raw ranked hits when you need to apply your own ranking or filtering before building a prompt.
- Add a memory, write new content and run the full extraction pipeline so future context calls include it.
- Delete a memory, soft-delete a memory so it is excluded from future context blocks.
- API reference, complete endpoint contract with all parameters and response shapes.