Support
FAQs
The questions developers ask before they wire Korely in.
What is Korely Agents?
A managed cloud memory layer for AI agents. Your agent writes what it learns
with add() and pulls a prompt-ready block back with
get_context(), across sessions, scoped per end user. Reachable
from the SDK, the CLI, or the REST API.
How is it different from a vector database?
A vector database stores embeddings and returns nearest neighbours. Korely
stores the memory and mines it into typed
(subject, predicate, object) facts, resolves contradictions over
time, and links entities in a graph. A read returns the current
truth, not just the closest text.
What is the difference between a memory and a fact?
A memory is the text you store ("Maria moved to the Pro
plan"). Facts are the typed triples extracted from it,
like (Maria, plan, Pro), each with a valid_from and an
invalid_at, so the timeline of what was true stays queryable.
How do I keep one customer's data separate from another's?
Scope every call with user_id (your end user). agent_id
scopes one of your apps and run_id a single session. The scope is
enforced server-side on every query, cross-tenant reads are impossible by
construction, not by prompt discipline.
Where is my data stored?
In the EU: Postgres and pgvector, on our own infrastructure, on every tier. It is not replicated to the US. See Data & governance.
How do reads and writes count against my plan?
Writes (add, batch items, edits, fact writes and corrections) count against your monthly memories quota, one write per 6,000 characters of a memory's text. Reads (search, context, facts, get, users) count against your queries quota. Read quotas are an order of magnitude more generous because reads run no generative model: they embed the query and look things up.
Is there a free tier?
Yes, Hobby is free (250 memories and 25,000 queries a month, 2 agents). Developer (19 euro a month) and Team (79 euro a month) raise the limits. End users are unlimited on every tier. Pricing shows the prices where you are.
Can a person see and edit what my agent remembered?
Yes, through your product. get_profile shows everything
known about one end user, correct_fact and
forget_fact fix or close a single fact, and
delete_all erases the user with one call. You have the same
Correct and Forget actions in the dashboard. See
Human in the loop.
Do you train models on my data?
No. Your stored memories are never used to train models.
What happens when I hit my quota (429)?
Your plan's figure has a 10% grace on top. Once a write would take you past
both, it returns 429 with the standard error envelope,
code is quota_exceeded, and nothing is written.
Reads past the queries quota get the same answer. The body:
{
"code": "quota_exceeded",
"message": "Monthly write limit reached: 250 writes on the Hobby plan this month, plus a 10% grace. It starts again on 2026-11-01. To go on now, move to a paid plan at https://agent.korely.ai/usage (an account made with korely init connects its key to a login there first: https://korely.ai/docs/agent-signup/#connect)",
"limit": 250,
"used": 275,
"resets_at": "2026-11-01T00:00:00Z"
}
Branch on code, never on the message. limit is
the plan's monthly figure (top-ups included), used what the
month has used, resets_at when the count starts again (00:00
UTC on the 1st). The quota 429 carries no Retry-After header: read resets_at. Read calls
(search, context, facts, get, users) are never blocked by a write quota:
they have their own, larger queries quota. A batch import refused for the
quota carries the same limit, used and
resets_at, and a message that says how many
writes are left.
A separate rate-limit 429 (too many requests in a minute, hour or day) does carry
a Retry-After header in integer seconds. All Korely errors
share one envelope, {"code": "...", "message": "..."},
with these three extra fields on quota_exceeded. On a paid plan you
can add 1,000 writes to the month once, from the
Usage page of the dashboard; past that, upgrade your plan to raise the write
quota. Current limits per plan: Hobby 250 writes /
25k queries, Developer 2k /
250k, Team 8k /
1M, Scale 25k /
10M. See Pricing.
What is the agent_id cap and how does the 403 work?
Each plan allows a fixed number of distinct agent_id values
(Hobby: 2, Developer: 10, Team: 100, Scale: 500). The first time you pass a
new agent_id that would push you past that ceiling,
the request returns 403 agent_cap_exceeded, the memory is
not written. Existing agents keep working. To add a new
agent, either upgrade or delete an unused agent with
DELETE /v1/agents/{id}.
There is no cap on the number of end users (user_id) or
sessions (run_id) on any tier.
Can I import a large history in bulk?
Yes, use the batch endpoint. Submit a batch of memories in a single request;
the server processes them asynchronously and returns a job id
(e.g. job_4e1aa0) with its received count. Poll
GET /v1/batch/{id} until status is
completed or failed; the result reports how many were
imported versus failed. Every fact the import
extracts is run through the same contradiction check, so a bulk load lands
as resolved, typed facts, not just dumped rows.
POST /v1/batch
{
"memories": [
{ "content": "Alice prefers dark mode", "user_id": "customer-4812" },
{ "content": "Bob is on the Team plan", "user_id": "customer-7203" }
]
}
Response: { "id": "job_4e1aa0", "status": "processing", "received": 2 }
GET /v1/batch/job_4e1aa0
Response: {
"id": "job_4e1aa0",
"status": "completed",
"received": 2,
"imported": 2,
"failed": 0,
"errors": []
}
Batch writes count against the same monthly writes quota as individual
add() calls.
How fast are reads in practice?
No generative model runs on a read. What each read does:
GET /v1/factsandGET /v1/memories/{id}: a database query, no model call at all.POST /v1/memories/search: the query becomes a vector on our server (EmbeddingGemma), then a pgvector search. No third party sees the query.GET /v1/context(the primary recall path): the active typed facts plus the same query embedding and vector search, assembled into one block.
POST /v1/memories (add) is slower, it triggers entity extraction
and embedding, but that happens asynchronously after the 201 is returned,
so your agent is never blocked.
Is there a self-hosted option?
Yes. The same engine runs on your own infrastructure: Docker Compose, your
Postgres, your model provider account, and a dashboard served by the
install itself. It is the Self-hosted plan on the
pricing page, and it starts with a call. The
four public tiers stay hosted in our EU cloud: on every tier, including
Hobby, your memories, facts and their vectors are stored in the EU. To
extract facts, the text of each memory you write goes to the language
model of your project's region: Google's Gemini API, in the United
States, in the default global region, or gpt-oss-120b on
Scaleway, in Paris, in the eu region. Searches and context
queries run no language model. See Data & governance
and the sub-processors on the Security page.