Korely

Core operations

Get context

Retrieve a single, prompt-ready context block assembled from the most relevant memories and current facts, one call, right before your LLM generates.

get_context is the read call an agent makes immediately before sending a prompt to an LLM. It searches across all stored memories and current facts for the given query, assembles them into a single formatted string that fits inside your token budget, and returns the source ids so you can surface citations. The context string is ready to drop directly into a system prompt or user message, no post-processing required.

get_context is the primary recall path, and the one place Korely's moat shows up in a single call. It does not just rank raw text: it assembles the user's currently-valid typed facts, (subject, predicate, object) triples with bi-temporal validity, so superseded facts are excluded and only what holds right now is included, alongside the most relevant memories. The returned sources array mixes fact ids (fct_) and memory ids (mem_), showing exactly which assembled the block.

Unlike search, which runs semantic vector retrieval and returns a list of raw memory hits, get_context ranks, deduplicates, and formats the most relevant material, facts first, into a cohesive block. Use it when you want memory with zero glue code. Use search when you need the raw results to apply your own ranking or filtering logic.

flowchart LR
    A([Agent]) -->|"GET /v1/context?query=..."| B[Korely API]
    B --> C[(Memories)]
    B --> D[(Facts)]
    C --> E[Rank + dedup]
    D --> E
    E --> F[Format to token budget]
    F -->|"{ context, tokens, sources }"| A
    A -->|"system prompt + context"| G([LLM])
get_context assembles ranked memories and current facts into one block before generation.

Request

Endpoint: GET /v1/context. SDK: korely.get_context(query, *, user_id=None, agent_id=None, token_budget=800).

ParamTypeNotes
query string Required, 1-2,000 characters. The text to retrieve relevant context for. Typically the user's latest message or the topic your agent is about to reason over.
user_id string Optional, max 255 characters. Scopes retrieval to one end user's facts and memories. Pass your app's user identifier (e.g. customer-4812). If omitted, facts are not filtered by end user and memories are searched across all end users of the key's project.
agent_id string Optional. Narrows the facts and the memories to one agent namespace, as on search and list.
token_budget integer Optional, integer from 50 to 8,000. Maximum approximate token count for the returned context string. Default: 800. Korely packs the highest-scoring content that fits within this budget. Lower values keep context tight for small models; raise it for more coverage.

Example

from korely_memory import Korely
korely = Korely(api_key="kor_live_...")
# Call right before you build the LLM prompt
result = korely.get_context(
query="What are this user's dietary preferences?",
user_id="customer-4812",
token_budget=600,
)
system_prompt = f"""You are a helpful nutrition assistant.
Relevant context about this user:
{result.context}
Answer using the context above when relevant."""
# result.tokens → 230
# result.sources → ["fct_b91e", "fct_c30a", "mem_8f2c1a"]

Response

{
"context": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. Ground your answer in this context. [...] do not abstain when it is present._\n\n## Known facts\n- The user allergic_to tree nuts (since 2026-06-11)\n- The user avoids gluten (since 2026-06-11)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-06-11] User is vegetarian, avoids gluten and is allergic to tree nuts.",
"stable": "_The facts below are a compact profile of the user; the memories are the verbatim source of truth. Ground your answer in this context. [...] do not abstain when it is present._",
"volatile": "## Known facts\n- The user allergic_to tree nuts (since 2026-06-11)\n- The user avoids gluten (since 2026-06-11)\n\n## Relevant memories\n_A time in square brackets after a relative expression, as in \"next month [June 2023]\", was already resolved from the day that memory was written: use it as written and never recompute it from the current date. The date at the start of a memory, as in [2023-05-25], is the day it was said._\n- [2026-06-11] User is vegetarian, avoids gluten and is allergic to tree nuts.",
"stable_hash": "a345b30abc62923c0ac2bd5f87a6682ef3843a8c787ef3a3a96aec91015f7bbe",
"tokens": 230,
"sources": [
"fct_b91e",
"fct_c30a",
"mem_8f2c1a"
],
"degraded": false,
"degraded_parts": []
}
FieldTypeDescription
context string A formatted, prompt-ready Markdown block: a note for the model that reads it (about 180 tokens, shortened in the example above; present whenever the block has content and token_budget is at least 400), then ## Known facts with the active facts, then ## Relevant memories, each under the day the memory was said ([YYYY-MM-DD]). It is stable and volatile joined by a blank line. Ready to embed directly into a system or user message.
stable string The note: the part that does not depend on the question, the same text from one call to the next. Put it in the system prompt, where your provider's prompt cache reuses it.
volatile string What the question brings: ## Known facts, ## Relevant memories, and a closing NOTE: when part of the retrieval failed. Put it in the user turn.
stable_hash string SHA-256 (hex) of stable. Unchanged from the previous call means the cached system prompt is still valid.
tokens integer Approximate token count of the returned context string. Always at or below the requested token_budget.
sources string[] Ids of the memories and facts included in the context block. Use these to power citations or to retrieve full objects via search or GET /v1/memories/{id}.
degraded boolean true when part of the block could not be retrieved the normal way. A partial context that looks complete is worse than an error: check it before trusting an empty answer.
degraded_parts string[] Which part: facts (ranked by recency instead of relevance) and memories (not retrieved). If the question cannot be embedded, both.

Prompt caching

Model providers cache only an identical prompt prefix. stable holds no date, no echo of the query and no count that depends on it, so from one question to the next it is byte for byte the same. Put it at the start of your system prompt and volatile in the user turn, and the provider can reuse the cached prefix on every turn, within its own rules (a minimum prompt length, a cache lifetime, on some providers an explicit cache marker). The same stable_hash as on the previous call means the same prefix.

[
{"role": "system", "content": "You are a helpful nutrition assistant.\n\n<stable>"},
{"role": "user", "content": "<volatile>\n\nWhat can I cook tonight?"}
]

Errors

StatusCode stringWhen it happens
401 invalid_key The Authorization header is missing, malformed, or the key has been revoked. Re-authenticate and retry.
403 forbidden The key lacks the memories:read scope.
422 invalid_request Request validation failed: query missing or over 2,000 characters, token_budget outside 50-8,000 or not an integer, or user_id over 255 characters. Like every Korely error, the body is the flat envelope {"code": "invalid_request", "message": "query.query: Field required"}.
429 quota_exceeded Your monthly query quota (including the +10% grace) is exhausted. Upgrade your plan or wait for the reset. The monthly-quota 429 does not carry a Retry-After header; only the per-minute, per-hour or per-day rate-limit 429 does, as integer seconds.
429 rate_limit_exceeded Too many requests in the current minute, hour or day. Wait the seconds in Retry-After.

Notes

  • Idempotency. GET /v1/context is a pure read, calling it multiple times with the same parameters produces the same result (assuming no new memories have been written). It is safe to retry on transient network errors without risk of duplicate writes.
  • Scoping. Pass user_id in customer-facing agents to restrict retrieval to one end user. A call without user_id draws from all end users of the key's project, which is correct for an internal ops agent and wrong for a per-customer chat. agent_id narrows the facts and the memories to one agent, on top of user_id.
  • Token budget. The budget is an approximation based on a 4-character-per-token estimate. The actual token count your LLM sees may differ by a few percent depending on the model's tokenizer. Leave a 10-20% buffer below your model's hard context limit.
  • Rate limiting. get_context counts against your monthly query quota, the same pool as search. Hobby plans include 25k queries per month; Developer 250k; Team 1M; Scale 10M. It also counts toward the per-minute, per-hour and per-day rate limits of your plan (429 rate_limit_exceeded with a Retry-After header), and hitting the monthly cap returns 429 quota_exceeded.
  • Empty context. If no memories or facts are stored for the given scope, context, stable and volatile are empty strings, tokens is 0, and sources is an empty array. This is a valid 200 response, not an error. Agents should handle the empty-context case gracefully by generating a neutral reply rather than surfacing an error to the user.

Related

  • Search memories, returns raw ranked hits when you need to apply your own ranking or filtering before building a prompt.
  • Add a memory, write new content and run the full extraction pipeline so future context calls include it.
  • Delete a memory, soft-delete a memory so it is excluded from future context blocks.
  • API reference, complete endpoint contract with all parameters and response shapes.