Core operations
Search memories
Semantic vector retrieval that returns the raw memories closest to a query, in one call. For the recall path that assembles typed facts into a ready-to-inject block, reach for get_context first.
Reach for get_context first. The differentiator
is fact-assembled recall: get_context resolves your end user's
active typed facts, bi-temporal (subject, predicate, object)
triples with contradiction resolution, and folds them, plus the most
relevant memories, into a single prompt block. Raw search below
is the secondary path: it ranks individual memories by semantic similarity
when you want the verbatim source rather than an assembled answer.
Search returns the memories most similar to a query so an agent can surface
source context before composing a reply. Given a natural-language or keyword
query, Korely embeds the query and ranks stored memories by cosine similarity
over their embeddings, then returns the top hits. Each hit carries a relevance
score and a short snippet so your agent can decide whether to fetch the full
memory or pass the snippet directly into the prompt. Search is
semantic vector retrieval only, it does not run a keyword
index or a graph walk. The typed-fact graph is reached through
get_context and
get_facts, not through this endpoint.
An agent typically calls search when it needs the verbatim memory behind a topic, scoped to the current end user. To assemble a compact, fact-aware memory block for the system prompt in one call, use get_context. For a full time-ordered list of every memory you've stored for a user, see get_all.
flowchart LR
Q([query]) --> EMB[Embed<br/>query]
EMB --> VEC[Cosine similarity<br/>over memory embeddings]
VEC --> R([ranked hits]) Request
Endpoint: POST /v1/memories/search. SDK: korely.search(query, *, user_id=None, agent_id=None, run_id=None, metadata=None, limit=None) (the server default is 15).
| Parameter | Type | Required | Description |
|---|---|---|---|
query | string | Required | The search query, 1-2,000 characters. Keyword-style (1-5 words) works best; longer conversational strings are also accepted. The query is embedded and ranked by cosine similarity against stored memories. A whitespace-only query returns an empty list. |
user_id | string | Optional | Strongly recommended in multi-tenant products. Max 255 characters. Scopes results to one end user. Without it, the search spans all end users of the key's project. |
agent_id | string | Optional | Max 255 characters. Further narrows to a specific agent surface. Omit to search across all agents of the key's project. |
run_id | string | Optional | Max 255 characters. Scopes results to one run or session. |
metadata | object | Optional | Filter on the metadata you stored at write time. Keys are ANDed; each value is compared as text with the stored one, e.g. {"tier": "pro"}: true matches a stored true, 5 a stored 5 (a number also matches an equal number, 5.0), "5" a stored 5 or "5", null a stored null. An object, an array, NaN or an infinity is refused with 422. |
limit | integer | Optional | Number of hits to return. Default 15, from 1 to 50. Hits are ordered by descending relevance score. |
min_score | number | Optional | From 0 to 1. Drops the hits whose score is below it. Absent: no floor. See A floor on the score. |
Any other field is refused with 422 invalid_request, Mem0's top_k and filters included.
Always pass user_id in customer-facing agents.
A search without user_id spans every end user stored in the
namespace. That is correct for an internal ops tool and wrong for a
per-customer chat. Filters are additive (AND): adding agent_id
on top of user_id narrows further, it does not broaden.
Example
from korely_memory import Korely
korely = Korely(api_key="kor_live_...", region="eu")
results = korely.search( "northwind pricing", user_id="customer-4812", limit=5,)
for hit in results: print(hit.id, hit.score, hit.snippet) # mem_8f2c1a 0.91 Northwind Hosting costs 50 euro per month since the June upgrade. # mem_3c77d9 0.74 The team agreed to renegotiate the hosting contract in Q3.Response
{ "results": [ { "id": "mem_8f2c1a", "score": 0.91, "snippet": "Northwind Hosting costs 50 euro per month since the June upgrade.", "user_id": "customer-4812", "agent_id": "infra-bot", "metadata": { "source": "slack" } }, { "id": "mem_3c77d9", "score": 0.74, "snippet": "The team agreed to renegotiate the hosting contract in Q3.", "user_id": "customer-4812", "agent_id": "infra-bot", "metadata": {} } ]}Response fields
| Field | Type | Description |
|---|---|---|
results | array | Ranked list of memory hits, best score first. |
results[].id | string | Memory ID. Pass to korely.get(id) to retrieve the full content and extracted facts. |
results[].score | float | 1 minus the cosine distance between the query and the memory, rounded to 4 decimals: at most 1.0, near 0 for unrelated text. Higher means closer in meaning. Mem0 returns a fused, normalized score; Korely returns the cosine. |
results[].snippet | string | Short excerpt (up to 280 characters) of the stored memory, suitable for direct inclusion in a prompt or for display in a UI. For the complete text, fetch the memory by ID. |
results[].user_id | string or null | The end user the memory belongs to. |
results[].agent_id | string | The agent surface that wrote the memory, or null if none was set. |
results[].metadata | object | The metadata object passed when the memory was written. Empty object if none was set. |
The response contains only the results array. There is no
total field and no memories wrapper. An empty
result is {"results": []}.
The read path is zero-generation. Search embeds your query (a fraction of a cent) and ranks stored memories by cosine similarity. No model ever composes or rewrites the snippets you receive. Your agent's own model does the reasoning over what comes back. Reads are retrieval, not generation.
A floor on the score
Without min_score, a search always returns its nearest
memories, even for a question whose answer was never stored. A sentence
that is close to nothing still has a nearest neighbour.
Pass min_score to drop the hits below a score:
{ "query": "what is the customer's dog called?", "user_id": "customer-4812", "min_score": 0.2}
If no memory reaches 0.2, the answer is {"results": []}.
We measured 0.2 on LongMemEval, 60 questions and 5,097 memories, with
the embeddings Korely uses. Every question kept its best memory, and more than half
of the searches whose answer was not stored came back empty. At 0.25,
2 of the 60 questions lost their best memory. No value separates the two cleanly,
so the floor is yours to set: start at 0.2 and read your own scores.
/v1/context takes no floor, because the model that reads the block
skips the lines that do not apply.
Fetching full content
Snippets are truncated for speed. When a hit's score is high enough to include in the prompt, fetch the complete memory to avoid cutting off important context:
results = korely.search("northwind pricing", user_id="customer-4812", limit=3)
# Pull full content for the top hitif results and results[0].score > 0.8: memory = korely.get(results[0].id) print(memory.content) # Northwind Hosting costs 50 euro per month since the June upgrade. # Contract renewed through December 2026. See invoice INV-2026-0611.Timeline and fact history
Search returns only current, active memories. There is no
include_history or time_filter parameter on this
endpoint. For the full lifecycle of a single memory (edits, supersessions,
deletion), use history. To query
the typed fact store with temporal filters, including superseded facts, use
get_facts(include_invalidated=True).
Errors
Search is a read-only operation, it counts against your monthly query quota, never the write quota. The table below lists every status code this endpoint can return.
| Status | Code string | Cause |
|---|---|---|
200 | None | Success. Results may be an empty array if no memories match; this is not an error. |
401 | invalid_key | The Authorization header is missing, malformed, or the key has been revoked. |
403 | forbidden | The key lacks the memories:read scope. |
422 | invalid_request | Request validation failed. The most common cause is a missing or empty query field, limit outside the 1-50 range, min_score outside 0-1, a metadata value that is an object, an array, NaN or an infinity, or an unknown field. The body is the flat {"code": "invalid_request", "message": "query: Field required"} envelope. |
429 | quota_exceeded | Monthly query quota is exhausted (past the grace allowance). The quota 429 carries no Retry-After, it resets on the 1st of the month (UTC). Upgrade your plan or wait for the monthly reset. |
429 | rate_limit_exceeded | Too many requests in the current minute, hour or day. Carries a Retry-After header, in seconds. |
500 | internal_error | The request failed on our side. Retry. |
503 | search_unavailable | The retrieval engine did not answer. Retry shortly; other read paths are unaffected. |
Every error response uses the same envelope: {"code": "<slug>", "message": "<text>"}.
A 401 is {"code": "invalid_key", "message": "Invalid or missing API key: no live key matches it (a revoked key stops at once); your keys are at https://agent.korely.ai/keys"}.
There is no error or detail field, and quota
information is never returned in the body.
Notes
- Read-only, non-destructive. Search never modifies stored memories or facts. Calling it any number of times with the same parameters is safe and produces the same ranked result set given the same corpus.
- Scoping is additive. Filters narrow results:
user_idalone returns all memories for that user across every agent; addingagent_idon top narrows to that agent's memories for the user. You cannot broaden results by combining filters. - No cross-workspace access. The API key determines the namespace. A search with
user_id="customer-4812"only touches end users stored under that key's workspace, it cannot reach another customer's data even if they share the sameuser_idstring. - Semantic vector ranking. Hits are ranked by cosine similarity between the query embedding and stored memory embeddings. There is no keyword index and no graph walk on this endpoint; for typed-fact recall use get_context or
get_facts. - Rate-limit behaviour. Each search call counts as one query against your plan's monthly query quota. Hobby (25k/month), Developer (250k/month), Team (1M/month), Scale (10M/month). When the monthly quota is used up, 10% grace included, the API returns
429 quota_exceededwith noRetry-After; its body sayslimit,usedandresets_at(00:00 UTC on the 1st). A separate per-minute, per-hour or per-day rate limit can also return429, that one carries aRetry-Afterheader in seconds. - Empty results are not errors. A
200response with"results": []means no memories matched the query for the given scope. This is normal for a new end user or a very specific query.
Related
- Add a memory, store content and run the full write pipeline before searching for it.
- Get context, the primary recall path: assembles a ready-to-inject prompt block from active typed facts plus the most relevant memories.
- Delete a memory, forget a memory by id so it no longer appears in search results.
- API reference, full endpoint contract including the
GET /v1/memories/{id}/historyandGET /v1/factsendpoints.