Korely

SDK

korely-memory is a typed Python and Node.js client over the REST API. Every method maps 1:1 onto an endpoint, so anything you can do with curl you can do with the SDK, and the JSON shapes in the API reference are the attribute shapes here. Same backend, same memory, as the CLI.

# 1. Connect
from korely_memory import Korely
korely = Korely(api_key="kor_live_...")
# 2. Remember (typed facts + bi-temporal built-in)
korely.add("User lives in Lisbon", user_id="dana")
# 3. The contradiction resolves itself: the new fact
# invalidates the old one, never deletes it
korely.add("User just moved to Berlin", user_id="dana")
# 4. Recall once extraction has run (a few seconds after each add):
# the active facts, assembled into a prompt-ready block
ctx = korely.get_context(query="where does the user live", user_id="dana")
print(ctx.context)
# A note for the model that reads the block, then:
# ## Known facts
# - The user lives_in Berlin (since 2026-06-11)
#
# ## Relevant memories
# ...
# The superseded Lisbon fact is not among the known facts.

Install

Terminal window
pip install korely-memory

Python 3.9 or later. The Node SDK needs Node 18+. No heavy dependencies: the SDK is a thin HTTP client. Embeddings, entity extraction, and fact extraction with contradiction checking all run on Korely's side, not in your process, so your install stays small and your process stays light.

Initialize

from korely_memory import Korely
korely = Korely(api_key="kor_live_...")
# Or read the key from the environment (KORELY_API_KEY)
korely = Korely()

Keys look like kor_live_... and belong to a project of your account: a key reads and writes only its own project's memories and facts, and agent_id and user_id scope the data inside it. The client's region argument picks the API address, and eu (https://api.korely.ai) is the only one. Your data is stored in the EU, on our own infrastructure; the model that writes your facts runs in the region of the project.

Core methods

Every method wraps exactly one REST endpoint:

MethodREST endpointPath
korely.add(...)POST /v1/memoriesWrite
korely.search(...)POST /v1/memories/searchRead
korely.get_all(...)GET /v1/memoriesRead
korely.get(id)GET /v1/memories/:idRead
korely.update(id, ...)PATCH /v1/memories/:idWrite
korely.delete(id)DELETE /v1/memories/:idWrite
korely.delete_all(user_id=...)DELETE /v1/users/:user_id/memoriesWrite
korely.add_fact_triple(...)POST /v1/factsWrite
korely.correct_fact(id, ...)PATCH /v1/facts/:idWrite
korely.forget_fact(id)POST /v1/facts/:id/forgetWrite
korely.get_facts(...)GET /v1/factsRead
korely.get_profile(user_id=...)GET /v1/profileRead
korely.get_context(...)GET /v1/contextRead
korely.history(id)GET /v1/memories/:id/historyRead
korely.users(...)GET /v1/usersRead
korely.list_agents(...)GET /v1/agentsRead
korely.delete_agent(agent_id)DELETE /v1/agents/:agent_idWrite
korely.events(...)GET /v1/eventsRead
korely.batch(...)POST /v1/batchWrite
korely.batch_status(id)GET /v1/batch/:idRead
korely.audit(...)GET /v1/auditRead
korely.iter_audit(...)GET /v1/audit, every pageRead
korely.ping()GET /v1/pingKey check, no quota
korely.delete_account(confirm=True)DELETE /v1/account?confirm=trueWrite
Korely.init_agent(agent_caller?)POST /v1/agents/init, no keySignup

The same methods in Node, in camelCase (getAll, iterAudit, deleteAccount, the static Korely.initAgent). init_agent is a class method because it is how you get a key. In Python, AsyncKorely has every method as a coroutine (async for over iter_audit); every Node method already returns a promise.

add also accepts a list of chat messages (role / content dicts), not just a string, they are joined into one block before storing, so you can hand it a conversation as-is.

Reads are retrieval, not generation. No generative model ever composes output on the read path. There is no reranker model and no answer synthesis; your agent's own model does the reasoning. The write path is where the intelligence runs: embeddings, entity extraction, typed-fact extraction with contradiction checking and bi-temporal validity, about a tenth of a cent per memory, all included. That is why read quotas are an order of magnitude more generous than write quotas.

add

Maps to POST /v1/memories. One call stores the memory and queues the full write pipeline. The return value is the stored memory, with status processing; the extracted facts, and the older facts they supersede, follow a few seconds later.

Fact extraction is asynchronous. add returns as soon as the memory is stored, so on the hosted service memory.facts is empty on the immediate response and the facts land within a few seconds as extraction and contradiction checking finish server-side. To read the facts back deterministically, poll get_facts (GET /v1/facts) a moment later, or rely on get_context, which always assembles the current active facts.

memory = korely.add(
"Northwind Hosting costs 50 euro per month since the June upgrade.",
agent_id="infra-bot",
user_id="customer-4812",
metadata={"source": "slack"},
)
print(memory.id) # mem_8f2c1a
print(memory.status) # processing: facts are extracted after the call returns
print(memory.facts) # []
# A few seconds later
facts = korely.get_facts(entity="Northwind Hosting", user_id="customer-4812")
print(facts[0].subject, facts[0].predicate, facts[0].object)
# Northwind Hosting costs 50 euro per month

search

Maps to POST /v1/memories/search. Semantic vector search: cosine similarity over the memory embeddings. The only model call on the read path is the query embedding, a fraction of a hundredth of a cent. Keyword-style queries (1 to 5 words) work best. For recall you almost always want get_context (below) instead, it assembles the active typed facts into a prompt-ready block, which is the primary recall path. search returns raw memory snippets and is the secondary path.

results = korely.search(
"northwind pricing",
user_id="customer-4812",
limit=5,
)
for hit in results:
print(hit.id, hit.score, hit.snippet)
# mem_8f2c1a 0.91 Northwind Hosting costs 50 euro per month...

Optional filters mirror the REST params: agent_id, user_id, run_id, metadata (keys ANDed, each value compared as text with the stored one, a number also as a number) and limit (server default 15, max 50). Filters are additive (AND).

get_all

Maps to GET /v1/memories. List memories in a scope, newest first, no search query needed. Pass user_id, agent_id or run_id to narrow the scope, and page with limit (default 50, max 200) and offset. Listing is plain SQL on the read path, no model calls.

page = korely.get_all(
user_id="customer-4812",
limit=50,
offset=0,
)
print(page.total) # 218
for memory in page:
print(memory.id, memory.created_at)
# mem_8f2c1a 2026-06-07T09:14:00.412345+00:00

get

Maps to GET /v1/memories/:id. Full content, metadata, and the facts extracted from this memory.

memory = korely.get("mem_8f2c1a")
print(memory.content) # Northwind Hosting costs 50 euro per month...
print(memory.metadata) # {"source": "slack"}
print(memory.created_at) # 2026-06-07T09:14:00Z

update

Maps to PATCH /v1/memories/:id. Updating content re-runs extraction, so facts stay in sync with what the memory says. As with add, the hosted service answers processing with an empty facts list, and the new facts follow a few seconds later. An edit counts as writes of the month like an add (one per 6,000 characters of the new text). Pass expected_updated_at for optimistic concurrency: if another writer got there first, the call raises StaleWriteError (409 stale_write) instead of clobbering. Pass back the exact updated_at the API returned (created_at before the first edit): the server compares it to the microsecond, so a timestamp typed by hand or rounded to the second is refused as stale.

current = korely.get("mem_8f2c1a")
memory = korely.update(
"mem_8f2c1a",
content="Northwind Hosting costs 55 euro per month after the storage add-on.",
# the exact value the API returned, never one typed by hand
expected_updated_at=current.updated_at or current.created_at,
)
print(memory.status) # processing
print(memory.facts) # [], the new facts arrive a few seconds later
# Read them afterwards (newest first)
facts = korely.get_facts(entity="Northwind Hosting", user_id="customer-4812")
print(facts[0].object) # 55 euro per month

delete

Maps to DELETE /v1/memories/:id. Forget one memory. Audited invalidation, not a hard row delete: the memory and its facts drop out of every default read, and an audit stub records when and by which key it was forgotten.

receipt = korely.delete("mem_8f2c1a")
print(receipt.status) # forgotten
print(receipt.facts_invalidated) # 1
print(receipt.audit_id) # aud_3d0f

delete_all

Maps to DELETE /v1/users/:user_id/memories. Bulk erasure for one end user: every memory and fact scoped to that user_id is physically deleted in a single call, superseded facts included, with one audit record. Deleting every memory for a user becomes one method call.

receipt = korely.delete_all(user_id="customer-4812")
print(receipt.memories_deleted) # 218
print(receipt.facts_deleted) # 64
print(receipt.audit_id) # aud_91xb

get_facts

Maps to GET /v1/facts. Typed (subject, predicate, object) triples with bi-temporal validity, the heart of the moat. Reads from the fact store are deterministic SQL, no model calls. Returns a flat list of facts (use get_profile for the grouped-by-family view). Pass as_of for a point-in-time query: what was true on that date. Works on every tier, hobby included.

# Current state: only active facts
facts = korely.get_facts(entity="Northwind Hosting")
print(facts[0].object) # 50 euro per month
print(facts[0].invalid_at) # None, active
# Point-in-time: what did we believe on June 1?
facts = korely.get_facts(entity="Northwind Hosting", as_of="2026-06-01")
print(facts[0].object) # 40 euro per month
print(facts[0].invalid_at) # 2026-06-07T09:14:00Z, superseded since
# Full history chain, superseded facts included
facts = korely.get_facts(
entity="Northwind Hosting",
include_invalidated=True,
)

Filters mirror the REST contract: user_id, agent_id, subject, entity (matches either side of the triple), predicate, predicate_family, include_invalidated, as_of, limit, offset. It returns a list with .total, the count before paging. See temporal facts for how invalidation works.

get_context

Maps to GET /v1/context. The one-call method: it assembles a prompt-ready context block (profile plus relevant facts plus relevant memories) within a token budget. Assembly is deterministic retrieval and formatting, not generation. Drop the returned string into your system prompt. This is the method most agent frameworks use.

ctx = korely.get_context(
query="plan infra budget",
user_id="customer-4812",
token_budget=800,
)
print(ctx.tokens) # 642
print(ctx.sources) # ["fct_b91e", "mem_8f2c1a"]
messages = [
{"role": "system", "content": f"You are a helpful assistant.\n\n{ctx.context}"},
{"role": "user", "content": user_message},
]

batch

Maps to POST /v1/batch. Bulk import for migrations: up to 500 memory objects per call (same shape as add), processed asynchronously. Items count against the memory quota.

job = korely.batch([
{"content": "Prefers async standups over meetings.", "user_id": "customer-0001"},
{"content": "Renewal date moved to October 1st.", "user_id": "customer-0002"},
])
print(job.status) # processing
job = korely.batch_status(job.id)
print(job.status) # completed
print(job.imported) # 2

add_fact_triple

Maps to POST /v1/facts. Write a typed (subject, predicate, object) fact directly, skipping extraction, when your agent already has the structured form. The contradiction check still runs, and the fact is bi-temporal, pass valid_from for a historical fact.

fact = korely.add_fact_triple(
"Marco", "works_at", "Acme GmbH",
user_id="customer-4812",
subject_type="person",
valid_from="2026-06-01",
)
print(fact.invalidated) # ids of any facts this one superseded

get_profile

Maps to GET /v1/profile. The assembled profile of one end user: the active facts known about them, the end user's own facts first, grouped by family. Pass as_of for the profile as it stood on a past date.

profile = korely.get_profile(user_id="customer-4812")
print(profile.total) # 7
print(list(profile.by_family)) # ["places", "work", "preferences"]
# The profile as it stood on March 1st
past = korely.get_profile(user_id="customer-4812", as_of="2026-03-01")

history

Maps to GET /v1/memories/:id/history. The lifecycle of one memory, keyed on the mem_ id: the events created, updated, fact_extracted, and fact_invalidated, each timestamped. (To walk a single fact's supersede chain, use get_facts with include_invalidated=True instead.)

h = korely.history("mem_8f2c1a")
for event in h.events:
print(event.event, event.at) # created ... / fact_extracted ...

users

Maps to GET /v1/users. The end users you've stored data for, each with active memory and fact counts. Returns a page: iterable like a list, with .total for pagination.

page = korely.users()
print(page.total) # 1
for u in page:
print(u.user_id, u.memories, u.facts)

correct_fact

Maps to PATCH /v1/facts/:id. Supersede a fact with a corrected one: pass at least one of subject, predicate, object. Not an edit in place: the old row keeps its dates and gains invalidated_by, so as_of before the correction still returns what you believed then. Returns the new fact.

fact = korely.correct_fact("fct_a1", object="Beacon Labs")
print(fact.object, fact.invalid_at) # Beacon Labs None

forget_fact

Maps to POST /v1/facts/:id/forget. Close a fact: it stops being current and drops out of default reads, and stays in history (include_invalidated=True still shows it). Pass at (ISO date) to record when it stopped being true; the default is now. Forgetting twice is safe: the second call answers already_forgotten.

receipt = korely.forget_fact("fct_9c1e", at="2026-06-30")
print(receipt["status"]) # forgotten

list_agents

Maps to GET /v1/agents. The agent namespaces you have written under (the distinct agent_id values), each with memory and fact counts and last activity, plus your plan's agent cap and how many slots are used. Call it when a write fails with agent_cap_exceeded, to reuse an existing id.

page = korely.list_agents()
print(page.used, "of", page.cap) # 2 of 2
for a in page:
print(a.agent_id, a.memories, a.facts)

delete_agent

Maps to DELETE /v1/agents/:agent_id. Delete an agent namespace for good: every memory and fact written under that agent_id in the key's project is purged, and its slot in your plan's agent cap is freed unless another project of the account still uses the same name. Returns the counts, an audit id and slot_freed.

receipt = korely.delete_agent("staging-bot")
print(receipt.memories_deleted, receipt.facts_deleted, receipt.audit_id)

events

Maps to GET /v1/events. Which writes have finished processing. add returns as soon as the memory is stored and fact extraction runs behind it, so a read taken right away can find no facts yet. Each event carries memory_id and a status of processing, ready or error; processing counts what is still in flight, so a batch import can wait on one number.

result = korely.events(user_id="customer-4812")
print(result["processing"]) # 0
for e in result["events"]:
print(e["memory_id"], e["status"]) # mem_8f2c1a ready

Migrating from another memory API? The migration guide maps the request shapes side by side.

Method signatures at a glance

In Python the first argument of add, search, get_context, the id of get, update, delete, history, delete_agent, forget_fact, correct_fact, batch_status, and the triple of add_fact_triple are positional; every other parameter is keyword-only (the * in the signatures below). In Node the options are one object: after the positional arguments, or as the only argument for getAll, users, listAgents, getFacts, getProfile, events and deleteAll (getContext takes the query as a string or inside the object). Node method names are camelCase versions of the Python names (get_all becomes getAll, delete_all becomes deleteAll, and so on).

Memory

Method (Python)ParametersReturns
add(content, *, user_id?, agent_id?, run_id?, metadata?, timestamp?) content: str or list of role/content dicts. Scope with any combination of user_id, agent_id, run_id. metadata: free-form dict stored alongside the memory. timestamp: ISO date the events happened (facts take it as valid_from). Memory with id, status, facts (empty while status is processing)
search(query, *, user_id?, agent_id?, run_id?, metadata?, limit?) query: str, keyword or natural-language. Filters are AND-combined. limit: default 15, max 50. list of SearchHit with id, score, snippet, user_id, agent_id, metadata
get_all(*, user_id?, agent_id?, run_id?, limit?, offset?) No query. Lists by recency. limit default 50, max 200. MemoryPage with memories, total
get(id) id: str, e.g. mem_8f2c1a Memory full object including facts
update(id, *, content, expected_updated_at?) Re-runs extraction on new content. expected_updated_at enables optimistic concurrency. Counts one write per 6,000 characters of the new text. Memory updated, status processing until the new facts land
delete(id) id: str DeleteReceipt with status, facts_invalidated, audit_id
delete_all(*, user_id) user_id: str, erases every memory for one end user BulkReceipt with user_id, memories_deleted, facts_deleted, erasure, audit_id (memories_forgotten and facts_invalidated are deprecated aliases)
history(id) id: str MemoryHistory with list of events
batch(memories) memories: list of add-shaped objects, max 500. Each takes content and optionally user_id, agent_id, run_id, metadata, timestamp; an unknown key refuses the batch with a 422. BatchJob with id, status, received
batch_status(id) id: str, job id from batch() BatchJob with status, received, imported, failed, errors

Facts

Method (Python)ParametersReturns
get_facts(*, entity?, subject?, predicate?, predicate_family?, user_id?, agent_id?, as_of?, include_invalidated?, limit?, offset?) All filters optional. entity matches either side of the triple. predicate is normalized server-side (the raw verb is returned as predicate_raw); predicate_family is one of eleven families, chosen once per project for each new relation; filter by entity when you are not sure which one. as_of: ISO-8601 date string for point-in-time queries. include_invalidated: bool, default false. Works on every tier, hobby included. list of Fact with .total; each fact has the fields of Get facts: id, subject, subject_type, predicate, predicate_raw, object, object_is_literal, predicate_family, confidence, user_id, agent_id, valid_from, invalid_at, invalidated_by, source_memory_id, created_at, subject_canonical, object_canonical, tense, last_confirmed_at, observation_count, source_memory_ids
correct_fact(fact_id, *, subject?, predicate?, object?) At least one field. Supersedes the old fact, which stays readable as history. Fact (the new, current one)
forget_fact(fact_id, *, at?) at: ISO date the fact stopped being true, default now. ForgetReceipt (a dict, so r["status"] works too) with id, status (forgotten or already_forgotten), invalid_at, audit_id
add_fact_triple(subject, predicate, object, *, user_id?, agent_id?, run_id?, subject_type?, object_is_literal?, confidence?, valid_from?, tense?) Writes a typed triple directly, skipping NLP extraction. Contradiction check still runs. valid_from: ISO-8601 for historical back-dating. Fact with invalidated list

Context and profile

Method (Python)ParametersReturns
get_context(query, *, user_id?, agent_id?, token_budget?) query: the current user turn or topic. token_budget: default 800, from 50 to 8000. Assembly is deterministic retrieval, no generation. Context with context (str ready for system prompt), tokens, sources, and stable, volatile, stable_hash, degraded, degraded_parts (Python from 0.1.18; Node types them too). An older server that does not send them leaves them empty (None, or [] for degraded_parts).
get_profile(*, user_id, agent_id?, as_of?) user_id: required. as_of: ISO-8601 for a historical snapshot. Profile with user_id, as_of, facts, by_family dict, total, truncated

Admin

Method (Python)ParametersReturns
users(*, agent_id?, limit?, offset?) Lists end users you have stored data for. UsersPage with users, total
list_agents(*, limit?, offset?) Lists your agent namespaces. AgentsPage with agents, total, cap, used
delete_agent(agent_id) Purges one agent namespace of the key's project and frees its cap slot. AgentDeleteReceipt with agent_id, memories_deleted, facts_deleted, audit_id, slot_freed
events(*, user_id?, status?, limit?) status: processing, ready or error. EventsResponse (a dict) with events, processing
audit(*, user_id?, action?, since?, until?, limit?, offset?) The trail of the key's project, newest first. since / until: ISO text, datetime or date. limit 1 to 1000, default 100. Counts against no quota. AuditPage with events, total
iter_audit(*, user_id?, action?, since?, until?, page_size?, offset?) Every page of audit(), for an export, with until pinned when it starts. iterator of AuditEvent
ping() Checks the key: no scope, no rate limit, no quota. PingResponse with ok, tier, region, scopes
delete_account(*, confirm) Cloud, for an account made by init_agent() or korely init --agent: deletes it with every key, memory and fact. confirm=True is required. An account with a login raises ConflictError (account_has_login). AccountDeleteReceipt with deleted, removed
Korely.init_agent(agent_caller?, *, base_url?) Class method, no key: signs up for a free hobby account. Cloud only. AgentInitResult with api_key (shown once), tier, region, scopes, quotas

End-to-end example

A support bot that remembers each customer across sessions. On every turn it pulls a prompt-ready context block, calls the LLM, then stores what the user said. Three SDK calls per turn: get_context, llm.chat, add.

"""
Support bot with persistent per-customer memory.
Uses korely-memory for context recall and fact storage.
"""
import os
import google.generativeai as genai
from korely_memory import Korely, APIError
korely = Korely(api_key=os.environ["KORELY_API_KEY"])
genai.configure(api_key=os.environ["GOOGLE_API_KEY"])
model = genai.GenerativeModel("gemini-2.0-flash")
AGENT = "support-bot"
def chat(user_id: str, user_message: str) -> str:
# 1. Assemble memory context for this user + query
ctx = korely.get_context(
query=user_message,
agent_id=AGENT,
user_id=user_id,
token_budget=600,
)
system = (
"You are a concise support assistant. "
"Use the memory context below to personalise your reply.\n\n"
+ ctx.context
)
# 2. Call the LLM
reply = model.generate_content(
[{"role": "user", "parts": [system + "\n\nUser: " + user_message]}]
).text
# 3. Persist what the user said (extracts facts, detects contradictions)
try:
korely.add(
user_message,
agent_id=AGENT,
user_id=user_id,
metadata={"turn": "user"},
)
except APIError as err:
if err.code != "quota_exceeded":
raise
pass # over the write quota, log and continue, the reply was already generated
return reply
# --- run a two-turn session ---
uid = "customer-4812"
print(chat(uid, "Hi, I am on the Developer plan and I prefer async updates."))
# context is empty on first turn, the bot greets + confirms
print(chat(uid, "What plan am I on again?"))
# second turn: context block contains the plan + preference facts from turn 1
# → bot answers correctly without the user repeating themselves

What happens inside add on turn 1: Korely extracts two typed facts, (customer-4812, subscribed_to, Developer plan) and (customer-4812, likes, async updates) (the raw verb "prefers" is normalized to likes, and kept verbatim in predicate_raw), and stores them alongside the full memory text. Extraction runs server-side and is asynchronous, so memory.facts may be empty on the immediate response and fill in a moment later. On turn 2, get_context retrieves both: the context block your system prompt receives contains the plan and preference before the LLM sees the query. No prompt engineering required beyond dropping in ctx.context.

One agent, unlimited customers. The same agent_id="support-bot" can serve thousands of user_id values. Quotas count writes and queries, never the number of end users you remember.

Scoping

Three identifiers, three levels of scope. They are the same parameters everywhere: SDK, REST, and CLI.

ParamWhat it identifiesExample
agent_id Your application or agent. One namespace per product surface. "support-bot"
user_id Your end user. Free-form string, you choose the identifier. Scopes memory to one person your agent serves. "customer-4812"
run_id One session or agent run. Sub-scope inside a user. "session-2026-06-11"

End users are unlimited on every tier. One support-bot agent can remember thousands of distinct customers; quotas count memories and queries, never people. A typical product setup is one agent_id per surface and one user_id per customer:

# Write: scoped to this customer
korely.add(
"Asked to be contacted on Slack, not email.",
agent_id="support-bot",
user_id="customer-4812",
run_id="session-2026-06-11",
)
# Read: only this customer's memory comes back
results = korely.search("contact preference", user_id="customer-4812")

Always pass user_id on reads in multi-tenant products. Filters are additive (AND). A search without user_id spans every end user in the namespace, which is what you want for an internal ops agent and not what you want inside a customer-facing chat.

Error handling

Every error response carries the same envelope, {"code": "...", "message": "..."}, and the SDK raises it as an APIError with .status, .code, .message and .retry_after / .retryAfter, the server's Retry-After in seconds on any status that sends it (Python also keeps the server's answer in .body). The class tells the status:

  • AuthenticationError: 401.
  • NamespaceForbiddenError: 403.
  • NotFoundError: 404.
  • ConflictError: 409, with its subclass StaleWriteError for stale_write only. Other codes: fact_not_current from correct_fact, whose current_fact_id (Python) or currentFactId (Node) is the fact to correct instead, or None/undefined when nothing took its place; account_has_login from delete_account.
  • QuotaExceededError: 429, with its subclass TooManyBatchesError for too_many_batches. On quota_exceeded it carries limit, used and resets_at (Node: resetsAt).

All of them are APIErrors, and APIError is a KorelyError, the base of everything the SDK raises: a KorelyError that is not an APIError never got an answer from the server (no key, a timeout, a connection error). Branch on err.code to handle each case.

StatuscodeWhen
401invalid_keyMissing, malformed, or revoked API key. Message: "Invalid or missing API key: ...", then what is wrong.
403forbidden, agent_cap_exceededThe key lacks the scope the call needs, or a new agent_id would exceed your plan's agent cap.
404not_foundMemory id does not exist, or was forgotten. Message: "Memory not found".
409stale_writeupdate with an expected_updated_at that is not the record's current version (StaleWriteError).
409fact_not_currentcorrect_fact on a fact that is history; current_fact_id (Node: currentFactId) is the current one (ConflictError).
409account_has_logindelete_account with the key of an account somebody signs in to (ConflictError).
422invalid_requestValidation failure, e.g. "content: Field required".
429quota_exceededMonthly write quota (for writes) or query quota (for reads) used up, 10% grace included. No Retry-After. The error carries limit, used and resets_at (Node: resetsAt), from PyPI 0.1.20 and npm 0.1.12.
429rate_limit_exceededToo many requests in the current minute, hour or day. Carries Retry-After in integer seconds.
429too_many_batchesbatch() while three batches are still importing (TooManyBatchesError). No Retry-After: send it again when one finishes.
500internal_errorThe request failed on our side.
503search_unavailable, model_unavailable, writes_pausedA dependency did not answer, or the service's daily model budget is spent. Nothing was written; retry.
from korely_memory import Korely, APIError
korely = Korely(api_key="kor_live_...")
try:
memory = korely.get("mem_8f2c1a")
except APIError as err:
if err.code == "invalid_key":
# 401: check the key, or create a new one in the dashboard
raise
elif err.code == "not_found":
# 404: the memory does not exist, or an end user forgot it
memory = None
elif err.code == "quota_exceeded":
# 429: monthly query quota reached (get() is a read), upgrade or wait for the reset
memory = None
else:
raise

There is no overage billing, ever. At 80% of quota you get a quota.warning webhook; past 100% there is a +10% soft cap so a busy day does not break your agent; past that, writes return 429 with code: "quota_exceeded" (a QuotaExceededError, which is an APIError) until the month rolls over at resets_at, you top up the month on a paid plan, or you upgrade. Your bill is always exactly the tier price. See the API reference for quotas per tier.

Related

Explore further by topic:

TopicWhere to go
Full REST contract API reference, every endpoint, parameter, and response shape. The SDK is a thin wrapper; the reference is the source of truth.
CLI surface CLI reference, same memory, same key, from the terminal. korely context --user-id alice "what do I know about billing?"
MCP surface MCP reference, connect Claude, Cursor, or any MCP-compatible agent to your namespace with your kor_live_ key.
Bi-temporal facts Temporal facts, how contradiction detection works, what as_of queries return, and how the fact timeline is stored.
End-to-end chatbot Cookbook: chatbot that remembers, a fuller version of the example above, with conversation history and streaming.
Bulk import Cookbook: bulk import, migrate an existing user history using batch(), with progress tracking and error handling.
Surfaces overview When to use SDK vs CLI vs REST, a decision table for choosing the right surface per use case.
Migration from Mem0 Migration guide, request shapes mapped side by side, and what the Korely fact graph adds.
Pricing and quotas Pricing, hobby (free), developer (€19/mo), team (€79/mo), scale (€249/mo). Writes and queries count; end users are always unlimited.

Python and Node.js are available now, both korely-memory, on PyPI and npm.