Korely

Support

FAQs

The questions developers ask before they wire Korely in.

What is Korely Agents?

A managed cloud memory layer for AI agents. Your agent writes what it learns with add() and pulls a prompt-ready block back with get_context(), across sessions, scoped per end user. Reachable from the SDK, the CLI, or the REST API.

How is it different from a vector database?

A vector database stores embeddings and returns nearest neighbours. Korely stores the memory and mines it into typed (subject, predicate, object) facts, resolves contradictions over time, and links entities in a graph. A read returns the current truth, not just the closest text.

What is the difference between a memory and a fact?

A memory is the text you store ("Maria moved to the Pro plan"). Facts are the typed triples extracted from it, like (Maria, plan, Pro), each with a valid_from and an invalid_at, so the timeline of what was true stays queryable.

How do I keep one customer's data separate from another's?

Scope every call with user_id (your end user). agent_id scopes one of your apps and run_id a single session. The scope is enforced server-side on every query, cross-tenant reads are impossible by construction, not by prompt discipline.

Where is my data stored?

In the EU: Postgres and pgvector, on our own infrastructure, on every tier. It is not replicated to the US. See Data & governance.

How do reads and writes count against my plan?

Writes (add, batch items, edits, fact writes and corrections) count against your monthly memories quota, one write per 6,000 characters of a memory's text. Reads (search, context, facts, get, users) count against your queries quota. Read quotas are an order of magnitude more generous because reads run no generative model: they embed the query and look things up.

Is there a free tier?

Yes, Hobby is free (250 memories and 25,000 queries a month, 2 agents). Developer (19 euro a month) and Team (79 euro a month) raise the limits. End users are unlimited on every tier. Pricing shows the prices where you are.

Can a person see and edit what my agent remembered?

Yes, through your product. get_profile shows everything known about one end user, correct_fact and forget_fact fix or close a single fact, and delete_all erases the user with one call. You have the same Correct and Forget actions in the dashboard. See Human in the loop.

Do you train models on my data?

No. Your stored memories are never used to train models.

What happens when I hit my quota (429)?

Your plan's figure has a 10% grace on top. Once a write would take you past both, it returns 429 with the standard error envelope, code is quota_exceeded, and nothing is written. Reads past the queries quota get the same answer. The body:

{
  "code": "quota_exceeded",
  "message": "Monthly write limit reached: 250 writes on the Hobby plan this month, plus a 10% grace. It starts again on 2026-11-01. To go on now, move to a paid plan at https://agent.korely.ai/usage (an account made with korely init connects its key to a login there first: https://korely.ai/docs/agent-signup/#connect)",
  "limit": 250,
  "used": 275,
  "resets_at": "2026-11-01T00:00:00Z"
}

Branch on code, never on the message. limit is the plan's monthly figure (top-ups included), used what the month has used, resets_at when the count starts again (00:00 UTC on the 1st). The quota 429 carries no Retry-After header: read resets_at. Read calls (search, context, facts, get, users) are never blocked by a write quota: they have their own, larger queries quota. A batch import refused for the quota carries the same limit, used and resets_at, and a message that says how many writes are left.

A separate rate-limit 429 (too many requests in a minute, hour or day) does carry a Retry-After header in integer seconds. All Korely errors share one envelope, {"code": "...", "message": "..."}, with these three extra fields on quota_exceeded. On a paid plan you can add 1,000 writes to the month once, from the Usage page of the dashboard; past that, upgrade your plan to raise the write quota. Current limits per plan: Hobby 250 writes / 25k queries, Developer 2k / 250k, Team 8k / 1M, Scale 25k / 10M. See Pricing.

What is the agent_id cap and how does the 403 work?

Each plan allows a fixed number of distinct agent_id values (Hobby: 2, Developer: 10, Team: 100, Scale: 500). The first time you pass a new agent_id that would push you past that ceiling, the request returns 403 agent_cap_exceeded, the memory is not written. Existing agents keep working. To add a new agent, either upgrade or delete an unused agent with DELETE /v1/agents/{id}. There is no cap on the number of end users (user_id) or sessions (run_id) on any tier.

Can I import a large history in bulk?

Yes, use the batch endpoint. Submit a batch of memories in a single request; the server processes them asynchronously and returns a job id (e.g. job_4e1aa0) with its received count. Poll GET /v1/batch/{id} until status is completed or failed; the result reports how many were imported versus failed. Every fact the import extracts is run through the same contradiction check, so a bulk load lands as resolved, typed facts, not just dumped rows.

POST /v1/batch
{
  "memories": [
    { "content": "Alice prefers dark mode", "user_id": "customer-4812" },
    { "content": "Bob is on the Team plan", "user_id": "customer-7203" }
  ]
}

Response: { "id": "job_4e1aa0", "status": "processing", "received": 2 }

GET /v1/batch/job_4e1aa0
Response: {
  "id": "job_4e1aa0",
  "status": "completed",
  "received": 2,
  "imported": 2,
  "failed": 0,
  "errors": []
}

Batch writes count against the same monthly writes quota as individual add() calls.

How fast are reads in practice?

No generative model runs on a read. What each read does:

  • GET /v1/facts and GET /v1/memories/{id}: a database query, no model call at all.
  • POST /v1/memories/search: the query becomes a vector on our server (EmbeddingGemma), then a pgvector search. No third party sees the query.
  • GET /v1/context (the primary recall path): the active typed facts plus the same query embedding and vector search, assembled into one block.

POST /v1/memories (add) is slower, it triggers entity extraction and embedding, but that happens asynchronously after the 201 is returned, so your agent is never blocked.

Is there a self-hosted option?

Yes. The same engine runs on your own infrastructure: Docker Compose, your Postgres, your model provider account, and a dashboard served by the install itself. It is the Self-hosted plan on the pricing page, and it starts with a call. The four public tiers stay hosted in our EU cloud: on every tier, including Hobby, your memories, facts and their vectors are stored in the EU. To extract facts, the text of each memory you write goes to the language model of your project's region: Google's Gemini API, in the United States, in the default global region, or gpt-oss-120b on Scaleway, in Paris, in the eu region. Searches and context queries run no language model. See Data & governance and the sub-processors on the Security page.