SDK
korely-memory is a typed Python and Node.js client over the
REST API. Every method maps 1:1
onto an endpoint, so anything you can do with curl you can do with the SDK,
and the JSON shapes in the API reference are the attribute shapes here.
Same backend, same memory, as the
CLI.
# 1. Connectfrom korely_memory import Korelykorely = Korely(api_key="kor_live_...")
# 2. Remember (typed facts + bi-temporal built-in)korely.add("User lives in Lisbon", user_id="dana")
# 3. The contradiction resolves itself: the new fact# invalidates the old one, never deletes itkorely.add("User just moved to Berlin", user_id="dana")
# 4. Recall once extraction has run (a few seconds after each add):# the active facts, assembled into a prompt-ready blockctx = korely.get_context(query="where does the user live", user_id="dana")print(ctx.context)# A note for the model that reads the block, then:# ## Known facts# - The user lives_in Berlin (since 2026-06-11)## ## Relevant memories# ...# The superseded Lisbon fact is not among the known facts.Install
pip install korely-memoryPython 3.9 or later. The Node SDK needs Node 18+. No heavy dependencies: the SDK is a thin HTTP client. Embeddings, entity extraction, and fact extraction with contradiction checking all run on Korely's side, not in your process, so your install stays small and your process stays light.
Initialize
from korely_memory import Korely
korely = Korely(api_key="kor_live_...")
# Or read the key from the environment (KORELY_API_KEY)korely = Korely()
Keys look like kor_live_... and belong to a
project of your account: a key
reads and writes only its own project's memories and facts, and
agent_id and user_id scope the data inside it.
The client's region argument picks the API address, and
eu (https://api.korely.ai) is the only one. Your
data is stored in the EU, on our own infrastructure; the model that writes
your facts runs in the
region of the project.
Core methods
Every method wraps exactly one REST endpoint:
| Method | REST endpoint | Path |
|---|---|---|
korely.add(...) | POST /v1/memories | Write |
korely.search(...) | POST /v1/memories/search | Read |
korely.get_all(...) | GET /v1/memories | Read |
korely.get(id) | GET /v1/memories/:id | Read |
korely.update(id, ...) | PATCH /v1/memories/:id | Write |
korely.delete(id) | DELETE /v1/memories/:id | Write |
korely.delete_all(user_id=...) | DELETE /v1/users/:user_id/memories | Write |
korely.add_fact_triple(...) | POST /v1/facts | Write |
korely.correct_fact(id, ...) | PATCH /v1/facts/:id | Write |
korely.forget_fact(id) | POST /v1/facts/:id/forget | Write |
korely.get_facts(...) | GET /v1/facts | Read |
korely.get_profile(user_id=...) | GET /v1/profile | Read |
korely.get_context(...) | GET /v1/context | Read |
korely.history(id) | GET /v1/memories/:id/history | Read |
korely.users(...) | GET /v1/users | Read |
korely.list_agents(...) | GET /v1/agents | Read |
korely.delete_agent(agent_id) | DELETE /v1/agents/:agent_id | Write |
korely.events(...) | GET /v1/events | Read |
korely.batch(...) | POST /v1/batch | Write |
korely.batch_status(id) | GET /v1/batch/:id | Read |
korely.audit(...) | GET /v1/audit | Read |
korely.iter_audit(...) | GET /v1/audit, every page | Read |
korely.ping() | GET /v1/ping | Key check, no quota |
korely.delete_account(confirm=True) | DELETE /v1/account?confirm=true | Write |
Korely.init_agent(agent_caller?) | POST /v1/agents/init, no key | Signup |
The same methods in Node, in camelCase (getAll,
iterAudit, deleteAccount, the static
Korely.initAgent). init_agent is a class method
because it is how you get a key. In Python, AsyncKorely has
every method as a coroutine (async for over
iter_audit); every Node method already returns a promise.
add also accepts a list of chat messages (role / content
dicts), not just a string, they are joined into one block before storing,
so you can hand it a conversation as-is.
Reads are retrieval, not generation. No generative model ever composes output on the read path. There is no reranker model and no answer synthesis; your agent's own model does the reasoning. The write path is where the intelligence runs: embeddings, entity extraction, typed-fact extraction with contradiction checking and bi-temporal validity, about a tenth of a cent per memory, all included. That is why read quotas are an order of magnitude more generous than write quotas.
add
Maps to POST /v1/memories. One call stores the memory and
queues the full write pipeline. The return value is the stored memory, with
status processing; the extracted facts, and the
older facts they supersede, follow a few seconds later.
Fact extraction is asynchronous. add returns
as soon as the memory is stored, so on the hosted service
memory.facts is empty on the immediate response and the facts
land within a few seconds as extraction
and contradiction checking finish server-side. To read the facts back
deterministically, poll get_facts (GET /v1/facts)
a moment later, or rely on get_context, which always assembles
the current active facts.
memory = korely.add( "Northwind Hosting costs 50 euro per month since the June upgrade.", agent_id="infra-bot", user_id="customer-4812", metadata={"source": "slack"},)
print(memory.id) # mem_8f2c1aprint(memory.status) # processing: facts are extracted after the call returnsprint(memory.facts) # []
# A few seconds laterfacts = korely.get_facts(entity="Northwind Hosting", user_id="customer-4812")print(facts[0].subject, facts[0].predicate, facts[0].object)# Northwind Hosting costs 50 euro per monthsearch
Maps to POST /v1/memories/search. Semantic vector search:
cosine similarity over the memory embeddings. The only model call on the
read path is the query embedding, a fraction of a hundredth of a cent.
Keyword-style queries (1 to 5 words) work best. For recall you almost always
want get_context (below) instead, it assembles the active
typed facts into a prompt-ready block, which is the primary recall path.
search returns raw memory snippets and is the secondary path.
results = korely.search( "northwind pricing", user_id="customer-4812", limit=5,)
for hit in results: print(hit.id, hit.score, hit.snippet) # mem_8f2c1a 0.91 Northwind Hosting costs 50 euro per month...
Optional filters mirror the REST params: agent_id,
user_id, run_id, metadata (keys ANDed,
each value compared as text with the stored one, a number also as a number) and limit (server default 15, max 50).
Filters are additive (AND).
get_all
Maps to GET /v1/memories. List memories in a scope, newest
first, no search query needed. Pass user_id,
agent_id or run_id to narrow the scope, and page with
limit (default 50, max 200) and offset. Listing is plain
SQL on the read path, no model calls.
page = korely.get_all( user_id="customer-4812", limit=50, offset=0,)
print(page.total) # 218for memory in page: print(memory.id, memory.created_at) # mem_8f2c1a 2026-06-07T09:14:00.412345+00:00get
Maps to GET /v1/memories/:id. Full content, metadata, and the
facts extracted from this memory.
memory = korely.get("mem_8f2c1a")
print(memory.content) # Northwind Hosting costs 50 euro per month...print(memory.metadata) # {"source": "slack"}print(memory.created_at) # 2026-06-07T09:14:00Zupdate
Maps to PATCH /v1/memories/:id. Updating content re-runs
extraction, so facts stay in sync with what the memory says. As with
add, the hosted service answers processing with
an empty facts list, and the new facts follow a few seconds
later. An edit counts as writes of the month like an add (one per
6,000 characters of the new text). Pass
expected_updated_at for optimistic concurrency: if another
writer got there first, the call raises StaleWriteError
(409 stale_write) instead of clobbering. Pass back the exact
updated_at the API returned (created_at before
the first edit): the server compares it to the microsecond, so a
timestamp typed by hand or rounded to the second is refused as stale.
current = korely.get("mem_8f2c1a")memory = korely.update( "mem_8f2c1a", content="Northwind Hosting costs 55 euro per month after the storage add-on.", # the exact value the API returned, never one typed by hand expected_updated_at=current.updated_at or current.created_at,)
print(memory.status) # processingprint(memory.facts) # [], the new facts arrive a few seconds later
# Read them afterwards (newest first)facts = korely.get_facts(entity="Northwind Hosting", user_id="customer-4812")print(facts[0].object) # 55 euro per monthdelete
Maps to DELETE /v1/memories/:id. Forget one memory.
Audited invalidation, not a hard row delete: the memory and its facts
drop out of every default read, and an audit stub records when and by
which key it was forgotten.
receipt = korely.delete("mem_8f2c1a")
print(receipt.status) # forgottenprint(receipt.facts_invalidated) # 1print(receipt.audit_id) # aud_3d0fdelete_all
Maps to DELETE /v1/users/:user_id/memories. Bulk erasure
for one end user: every memory and fact scoped to that
user_id is physically deleted in a single call, superseded
facts included, with one audit record. Deleting every memory for a user becomes one method call.
receipt = korely.delete_all(user_id="customer-4812")
print(receipt.memories_deleted) # 218print(receipt.facts_deleted) # 64print(receipt.audit_id) # aud_91xbget_facts
Maps to GET /v1/facts. Typed (subject, predicate, object)
triples with bi-temporal validity, the heart of the moat. Reads from the
fact store are deterministic SQL, no model calls.
Returns a flat list of facts (use get_profile for the
grouped-by-family view). Pass as_of for a point-in-time query:
what was true on that date. Works on every tier, hobby included.
# Current state: only active factsfacts = korely.get_facts(entity="Northwind Hosting")print(facts[0].object) # 50 euro per monthprint(facts[0].invalid_at) # None, active
# Point-in-time: what did we believe on June 1?facts = korely.get_facts(entity="Northwind Hosting", as_of="2026-06-01")print(facts[0].object) # 40 euro per monthprint(facts[0].invalid_at) # 2026-06-07T09:14:00Z, superseded since
# Full history chain, superseded facts includedfacts = korely.get_facts( entity="Northwind Hosting", include_invalidated=True,)
Filters mirror the REST contract: user_id,
agent_id, subject,
entity (matches either side of the triple),
predicate, predicate_family,
include_invalidated, as_of, limit,
offset. It returns a list with .total, the count
before paging.
See temporal facts for
how invalidation works.
get_context
Maps to GET /v1/context. The one-call method: it assembles a
prompt-ready context block (profile plus relevant facts plus relevant
memories) within a token budget. Assembly is deterministic retrieval and
formatting, not generation. Drop the returned string into your system
prompt. This is the method most agent frameworks use.
ctx = korely.get_context( query="plan infra budget", user_id="customer-4812", token_budget=800,)
print(ctx.tokens) # 642print(ctx.sources) # ["fct_b91e", "mem_8f2c1a"]
messages = [ {"role": "system", "content": f"You are a helpful assistant.\n\n{ctx.context}"}, {"role": "user", "content": user_message},]batch
Maps to POST /v1/batch. Bulk import for migrations: up to 500
memory objects per call (same shape as add), processed
asynchronously. Items count against the memory quota.
job = korely.batch([ {"content": "Prefers async standups over meetings.", "user_id": "customer-0001"}, {"content": "Renewal date moved to October 1st.", "user_id": "customer-0002"},])
print(job.status) # processing
job = korely.batch_status(job.id)print(job.status) # completedprint(job.imported) # 2add_fact_triple
Maps to POST /v1/facts. Write a typed
(subject, predicate, object) fact directly, skipping extraction, when your
agent already has the structured form. The contradiction check still runs,
and the fact is bi-temporal, pass valid_from for a historical
fact.
fact = korely.add_fact_triple( "Marco", "works_at", "Acme GmbH", user_id="customer-4812", subject_type="person", valid_from="2026-06-01",)
print(fact.invalidated) # ids of any facts this one supersededget_profile
Maps to GET /v1/profile. The assembled profile of one end user:
the active facts known about them, the end user's own facts first, grouped
by family. Pass as_of for the profile as it stood on a past date.
profile = korely.get_profile(user_id="customer-4812")print(profile.total) # 7print(list(profile.by_family)) # ["places", "work", "preferences"]
# The profile as it stood on March 1stpast = korely.get_profile(user_id="customer-4812", as_of="2026-03-01")history
Maps to GET /v1/memories/:id/history. The lifecycle of one
memory, keyed on the mem_ id: the events
created, updated, fact_extracted,
and fact_invalidated, each timestamped. (To walk a single
fact's supersede chain, use get_facts with
include_invalidated=True instead.)
h = korely.history("mem_8f2c1a")for event in h.events: print(event.event, event.at) # created ... / fact_extracted ...users
Maps to GET /v1/users. The end users you've stored data for,
each with active memory and fact counts. Returns a page: iterable like a
list, with .total for pagination.
page = korely.users()print(page.total) # 1for u in page: print(u.user_id, u.memories, u.facts)correct_fact
Maps to PATCH /v1/facts/:id. Supersede a fact with a
corrected one: pass at least one of subject,
predicate, object. Not an edit in place: the old
row keeps its dates and gains invalidated_by, so
as_of before the correction still returns what you believed
then. Returns the new fact.
fact = korely.correct_fact("fct_a1", object="Beacon Labs")print(fact.object, fact.invalid_at) # Beacon Labs Noneforget_fact
Maps to POST /v1/facts/:id/forget. Close a fact: it stops
being current and drops out of default reads, and stays in history
(include_invalidated=True still shows it). Pass
at (ISO date) to record when it stopped being true; the
default is now. Forgetting twice is safe: the second call answers
already_forgotten.
receipt = korely.forget_fact("fct_9c1e", at="2026-06-30")print(receipt["status"]) # forgottenlist_agents
Maps to GET /v1/agents. The agent namespaces you have written
under (the distinct agent_id values), each with memory and
fact counts and last activity, plus your plan's agent cap and
how many slots are used. Call it when a write fails with
agent_cap_exceeded, to reuse an existing id.
page = korely.list_agents()print(page.used, "of", page.cap) # 2 of 2for a in page: print(a.agent_id, a.memories, a.facts)delete_agent
Maps to DELETE /v1/agents/:agent_id. Delete an agent
namespace for good: every memory and fact written under that
agent_id in the key's project is purged, and its slot in your
plan's agent cap is freed unless another project of the account still uses
the same name. Returns the counts, an audit id and slot_freed.
receipt = korely.delete_agent("staging-bot")print(receipt.memories_deleted, receipt.facts_deleted, receipt.audit_id)events
Maps to GET /v1/events. Which writes have finished
processing. add returns as soon as the memory is stored and
fact extraction runs behind it, so a read taken right away can find no
facts yet. Each event carries memory_id and a
status of processing, ready or
error; processing counts what is still in
flight, so a batch import can wait on one number.
result = korely.events(user_id="customer-4812")print(result["processing"]) # 0for e in result["events"]: print(e["memory_id"], e["status"]) # mem_8f2c1a readyMigrating from another memory API? The migration guide maps the request shapes side by side.
Method signatures at a glance
In Python the first argument of add, search,
get_context, the id of get, update,
delete, history, delete_agent,
forget_fact, correct_fact, batch_status,
and the triple of add_fact_triple are positional; every other
parameter is keyword-only (the * in the signatures below). In
Node the options are one object: after the positional arguments, or as the
only argument for getAll, users,
listAgents, getFacts, getProfile,
events and deleteAll (getContext takes
the query as a string or inside the object). Node method names are camelCase
versions of the Python names (get_all becomes
getAll, delete_all becomes deleteAll,
and so on).
Memory
| Method (Python) | Parameters | Returns |
|---|---|---|
add(content, *, user_id?, agent_id?, run_id?, metadata?, timestamp?) | content: str or list of role/content dicts.
Scope with any combination of user_id, agent_id, run_id.
metadata: free-form dict stored alongside the memory.
timestamp: ISO date the events happened (facts take it as valid_from).
| Memory with id, status, facts (empty while status is processing) |
search(query, *, user_id?, agent_id?, run_id?, metadata?, limit?) | query: str, keyword or natural-language.
Filters are AND-combined.
limit: default 15, max 50.
| list of SearchHit with id, score, snippet, user_id, agent_id, metadata |
get_all(*, user_id?, agent_id?, run_id?, limit?, offset?) | No query. Lists by recency. limit default 50, max 200. | MemoryPage with memories, total |
get(id) | id: str, e.g. mem_8f2c1a | Memory full object including facts |
update(id, *, content, expected_updated_at?) | Re-runs extraction on new content. expected_updated_at enables optimistic concurrency. Counts one write per 6,000 characters of the new text. | Memory updated, status processing until the new facts land |
delete(id) | id: str | DeleteReceipt with status, facts_invalidated, audit_id |
delete_all(*, user_id) | user_id: str, erases every memory for one end user | BulkReceipt with user_id, memories_deleted, facts_deleted, erasure, audit_id (memories_forgotten and facts_invalidated are deprecated aliases) |
history(id) | id: str | MemoryHistory with list of events |
batch(memories) | memories: list of add-shaped objects, max 500. Each takes content and optionally user_id, agent_id, run_id, metadata, timestamp; an unknown key refuses the batch with a 422. | BatchJob with id, status, received |
batch_status(id) | id: str, job id from batch() | BatchJob with status, received, imported, failed, errors |
Facts
| Method (Python) | Parameters | Returns |
|---|---|---|
get_facts(*, entity?, subject?, predicate?, predicate_family?, user_id?, agent_id?, as_of?, include_invalidated?, limit?, offset?) |
All filters optional. entity matches either side of the triple.
predicate is normalized server-side (the raw verb is returned as
predicate_raw); predicate_family is one of eleven families, chosen once per project
for each new relation; filter by entity when you are not sure which one.
as_of: ISO-8601 date string for point-in-time queries.
include_invalidated: bool, default false.
Works on every tier, hobby included.
| list of Fact with .total; each fact has the fields of Get facts: id, subject, subject_type, predicate, predicate_raw, object, object_is_literal, predicate_family, confidence, user_id, agent_id, valid_from, invalid_at, invalidated_by, source_memory_id, created_at, subject_canonical, object_canonical, tense, last_confirmed_at, observation_count, source_memory_ids |
correct_fact(fact_id, *, subject?, predicate?, object?) | At least one field. Supersedes the old fact, which stays readable as history. | Fact (the new, current one) |
forget_fact(fact_id, *, at?) | at: ISO date the fact stopped being true, default now. | ForgetReceipt (a dict, so r["status"] works too) with id, status (forgotten or already_forgotten), invalid_at, audit_id |
add_fact_triple(subject, predicate, object, *, user_id?, agent_id?, run_id?, subject_type?, object_is_literal?, confidence?, valid_from?, tense?) |
Writes a typed triple directly, skipping NLP extraction.
Contradiction check still runs.
valid_from: ISO-8601 for historical back-dating.
| Fact with invalidated list |
Context and profile
| Method (Python) | Parameters | Returns |
|---|---|---|
get_context(query, *, user_id?, agent_id?, token_budget?) | query: the current user turn or topic.
token_budget: default 800, from 50 to 8000. Assembly is deterministic retrieval, no generation.
| Context with context (str ready for system prompt), tokens, sources, and stable, volatile, stable_hash, degraded, degraded_parts (Python from 0.1.18; Node types them too). An older server that does not send them leaves them empty (None, or [] for degraded_parts). |
get_profile(*, user_id, agent_id?, as_of?) | user_id: required.
as_of: ISO-8601 for a historical snapshot.
| Profile with user_id, as_of, facts, by_family dict, total, truncated |
Admin
| Method (Python) | Parameters | Returns |
|---|---|---|
users(*, agent_id?, limit?, offset?) | Lists end users you have stored data for. | UsersPage with users, total |
list_agents(*, limit?, offset?) | Lists your agent namespaces. | AgentsPage with agents, total, cap, used |
delete_agent(agent_id) | Purges one agent namespace of the key's project and frees its cap slot. | AgentDeleteReceipt with agent_id, memories_deleted, facts_deleted, audit_id, slot_freed |
events(*, user_id?, status?, limit?) | status: processing, ready or error. | EventsResponse (a dict) with events, processing |
audit(*, user_id?, action?, since?, until?, limit?, offset?) | The trail of the key's project, newest first. since / until: ISO text, datetime or date. limit 1 to 1000, default 100. Counts against no quota. | AuditPage with events, total |
iter_audit(*, user_id?, action?, since?, until?, page_size?, offset?) | Every page of audit(), for an export, with until pinned when it starts. | iterator of AuditEvent |
ping() | Checks the key: no scope, no rate limit, no quota. | PingResponse with ok, tier, region, scopes |
delete_account(*, confirm) | Cloud, for an account made by init_agent() or korely init --agent: deletes it with every key, memory and fact. confirm=True is required. An account with a login raises ConflictError (account_has_login). | AccountDeleteReceipt with deleted, removed |
Korely.init_agent(agent_caller?, *, base_url?) | Class method, no key: signs up for a free hobby account. Cloud only. | AgentInitResult with api_key (shown once), tier, region, scopes, quotas |
End-to-end example
A support bot that remembers each customer across sessions. On every turn it
pulls a prompt-ready context block, calls the LLM, then stores what the user
said. Three SDK calls per turn: get_context,
llm.chat, add.
"""Support bot with persistent per-customer memory.Uses korely-memory for context recall and fact storage."""import osimport google.generativeai as genaifrom korely_memory import Korely, APIError
korely = Korely(api_key=os.environ["KORELY_API_KEY"])genai.configure(api_key=os.environ["GOOGLE_API_KEY"])model = genai.GenerativeModel("gemini-2.0-flash")
AGENT = "support-bot"
def chat(user_id: str, user_message: str) -> str: # 1. Assemble memory context for this user + query ctx = korely.get_context( query=user_message, agent_id=AGENT, user_id=user_id, token_budget=600, )
system = ( "You are a concise support assistant. " "Use the memory context below to personalise your reply.\n\n" + ctx.context )
# 2. Call the LLM reply = model.generate_content( [{"role": "user", "parts": [system + "\n\nUser: " + user_message]}] ).text
# 3. Persist what the user said (extracts facts, detects contradictions) try: korely.add( user_message, agent_id=AGENT, user_id=user_id, metadata={"turn": "user"}, ) except APIError as err: if err.code != "quota_exceeded": raise pass # over the write quota, log and continue, the reply was already generated
return reply
# --- run a two-turn session ---uid = "customer-4812"
print(chat(uid, "Hi, I am on the Developer plan and I prefer async updates."))# context is empty on first turn, the bot greets + confirms
print(chat(uid, "What plan am I on again?"))# second turn: context block contains the plan + preference facts from turn 1# → bot answers correctly without the user repeating themselves
What happens inside add on turn 1: Korely extracts two typed
facts, (customer-4812, subscribed_to, Developer plan) and
(customer-4812, likes, async updates) (the raw verb "prefers"
is normalized to likes, and kept verbatim in
predicate_raw), and stores them alongside the full memory
text. Extraction runs server-side and is asynchronous, so
memory.facts may be empty on the immediate response and fill in
a moment later. On turn 2, get_context retrieves both: the
context block your system prompt receives contains the plan and preference
before the LLM sees the query. No prompt engineering required beyond
dropping in ctx.context.
One agent, unlimited customers. The same
agent_id="support-bot" can serve thousands of
user_id values. Quotas count writes and queries, never the
number of end users you remember.
Scoping
Three identifiers, three levels of scope. They are the same parameters everywhere: SDK, REST, and CLI.
| Param | What it identifies | Example |
|---|---|---|
agent_id | Your application or agent. One namespace per product surface. | "support-bot" |
user_id | Your end user. Free-form string, you choose the identifier. Scopes memory to one person your agent serves. | "customer-4812" |
run_id | One session or agent run. Sub-scope inside a user. | "session-2026-06-11" |
End users are unlimited on every tier. One support-bot
agent can remember thousands of distinct customers; quotas count memories
and queries, never people. A typical product setup is one
agent_id per surface and one user_id per
customer:
# Write: scoped to this customerkorely.add( "Asked to be contacted on Slack, not email.", agent_id="support-bot", user_id="customer-4812", run_id="session-2026-06-11",)
# Read: only this customer's memory comes backresults = korely.search("contact preference", user_id="customer-4812") Always pass user_id on reads in multi-tenant
products. Filters are additive (AND). A search without
user_id spans every end user in the namespace, which is
what you want for an internal ops agent and not what you want inside a
customer-facing chat.
Error handling
Every error response carries the same envelope,
{"code": "...", "message": "..."}, and the SDK
raises it as an APIError with .status,
.code, .message and .retry_after /
.retryAfter, the server's Retry-After in seconds
on any status that sends it (Python also keeps the server's answer in
.body). The class tells the status:
AuthenticationError: 401.NamespaceForbiddenError: 403.NotFoundError: 404.ConflictError: 409, with its subclassStaleWriteErrorforstale_writeonly. Other codes:fact_not_currentfromcorrect_fact, whosecurrent_fact_id(Python) orcurrentFactId(Node) is the fact to correct instead, orNone/undefinedwhen nothing took its place;account_has_loginfromdelete_account.QuotaExceededError: 429, with its subclassTooManyBatchesErrorfortoo_many_batches. Onquota_exceededit carrieslimit,usedandresets_at(Node:resetsAt).
All of them are APIErrors, and APIError is a
KorelyError, the base of everything the SDK raises: a
KorelyError that is not an APIError never got an
answer from the server (no key, a timeout, a connection error). Branch on
err.code to handle each case.
| Status | code | When |
|---|---|---|
401 | invalid_key | Missing, malformed, or revoked API key. Message: "Invalid or missing API key: ...", then what is wrong. |
403 | forbidden, agent_cap_exceeded | The key lacks the scope the call needs, or a new agent_id would exceed your plan's agent cap. |
404 | not_found | Memory id does not exist, or was forgotten. Message: "Memory not found". |
409 | stale_write | update with an expected_updated_at that is not the record's current version (StaleWriteError). |
409 | fact_not_current | correct_fact on a fact that is history; current_fact_id (Node: currentFactId) is the current one (ConflictError). |
409 | account_has_login | delete_account with the key of an account somebody signs in to (ConflictError). |
422 | invalid_request | Validation failure, e.g. "content: Field required". |
429 | quota_exceeded | Monthly write quota (for writes) or query quota (for reads) used up, 10% grace included. No Retry-After. The error carries limit, used and resets_at (Node: resetsAt), from PyPI 0.1.20 and npm 0.1.12. |
429 | rate_limit_exceeded | Too many requests in the current minute, hour or day. Carries Retry-After in integer seconds. |
429 | too_many_batches | batch() while three batches are still importing (TooManyBatchesError). No Retry-After: send it again when one finishes. |
500 | internal_error | The request failed on our side. |
503 | search_unavailable, model_unavailable, writes_paused | A dependency did not answer, or the service's daily model budget is spent. Nothing was written; retry. |
from korely_memory import Korely, APIError
korely = Korely(api_key="kor_live_...")
try: memory = korely.get("mem_8f2c1a")except APIError as err: if err.code == "invalid_key": # 401: check the key, or create a new one in the dashboard raise elif err.code == "not_found": # 404: the memory does not exist, or an end user forgot it memory = None elif err.code == "quota_exceeded": # 429: monthly query quota reached (get() is a read), upgrade or wait for the reset memory = None else: raise
There is no overage billing, ever. At 80% of quota you get a
quota.warning webhook; past 100% there is a +10% soft cap so
a busy day does not break your agent; past that, writes return
429 with code: "quota_exceeded" (a
QuotaExceededError, which is an APIError) until
the month rolls over at resets_at, you top up the month on a
paid plan, or you upgrade.
Your bill is always exactly the tier price. See
the API reference for quotas per
tier.
Related
Explore further by topic:
| Topic | Where to go |
|---|---|
| Full REST contract | API reference, every endpoint, parameter, and response shape. The SDK is a thin wrapper; the reference is the source of truth. |
| CLI surface | CLI reference, same memory, same key, from the terminal. korely context --user-id alice "what do I know about billing?" |
| MCP surface | MCP reference, connect Claude, Cursor, or any MCP-compatible agent to your namespace with your kor_live_ key. |
| Bi-temporal facts | Temporal facts, how contradiction detection works, what as_of queries return, and how the fact timeline is stored. |
| End-to-end chatbot | Cookbook: chatbot that remembers, a fuller version of the example above, with conversation history and streaming. |
| Bulk import | Cookbook: bulk import, migrate an existing user history using batch(), with progress tracking and error handling. |
| Surfaces overview | When to use SDK vs CLI vs REST, a decision table for choosing the right surface per use case. |
| Migration from Mem0 | Migration guide, request shapes mapped side by side, and what the Korely fact graph adds. |
| Pricing and quotas | Pricing, hobby (free), developer (€19/mo), team (€79/mo), scale (€249/mo). Writes and queries count; end users are always unlimited. |