Textual Memory

Long-context, persona, script, and conversation-memory benchmarks.

Measure what your agents remember.

Compare what truly matters.

A public benchmark space for comparing textual, multimodal, and coding-agent memory systems under a consistent evaluation flow.

Each track keeps its own result table and detailed metric breakdown.

Long-context, persona, script, and conversation-memory benchmarks.

Memory retrieval and generation over image-rich or multimodal tasks.

Agent memory support for coding tasks and repository-context recall.

Open-source methods and commercial products both use participant-hosted Add/Search APIs. AML does not deploy repository-only submissions.

Operate publicly reachable Add/Search endpoints and submit the fixed API version for review.

Use the issued key to verify the synchronous Add/Search flow.

After smoke passes, submit the full scored evaluation.

Use the product pages to inspect rankings, run evaluations, and prepare an integration.

Public ranking with filters, dataset columns, and score bars.

Create eval jobs, watch progress, and inspect private results.

Eligibility, submission routes, required materials, timelines, rewards, and publication rules.

User guide, evaluation workflow, API contract, security, and result publication.

Add/search API contract, request fields, polling, and response schemas.

Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.

Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.

Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.

Create API-gated eval jobs against the Leaderboard Suite, monitor task progress, inspect private results, and submit eligible full-suite runs for administrator review.

Choose a bound version. Use Run label to distinguish repeated evaluations of the same version.

Confirm each item before starting a public-board candidate run. The button remains locked until every item is checked.

Live job status and progress.

Review completed full-suite runs from non-admin leaderboard keys. Approved results enter the public board; rejected results remain private.

Private scores stay scoped to the current leaderboard key. Admin runs remain separate from external review candidates.

参赛免费,不设组队要求。参赛方承担自身 API、数据库、带宽和计算成本,平台承担统一 Answer、Eval 与评测编排成本。

2026 年 9 月 20 日 00:00

2026 年 10 月 31 日 23:59(UTC+8)

2026 年 11 月 4 日 23:59(UTC+8)

2026 年 11 月中旬

文本、代码与多模态记忆

评测类型决定系统接受什么任务;参赛组别决定结果展示在哪个榜单。两者是互相独立的两个维度。

第二期所有参赛组别都必须由参赛方自行部署 Add / Search API。Eval Key 是在线接口申请审核通过后获得的评测凭证。

提交固定版本、Add / Search 地址、鉴权方式和运行说明;公开仓库作为开放性与署名材料,不能替代在线接口。

提交固定产品版本、Add / Search 地址、鉴权和容量说明。无需公开内部实现,审核通过后获得 Eval Key,结果进入商业产品榜。

开源方法榜和商业产品榜都只通过参赛方自托管的 Add / Search API 接入。第二期不接受只提交 GitHub 仓库、Docker 镜像或启动说明并由 AML 代为部署。开源材料仍用于开放性、署名和复现审核。

请在提交申请前固定参评版本。正式 Full 评测受理后,不得因结果不理想更换版本或撤回。

系统名称与版本、联系人、机构或团队、拟参评类型、方法或产品说明、允许公开展示的信息,以及完整的提交说明。

已部署的 Add / Search API、鉴权与容量声明,以及公开仓库、固定 commit、原始工作引用和方法改动说明。

固定产品版本、Add / Search API、鉴权方式、评测专用密钥、容量限制、超时与限流说明。

两个 Key 的签发方和用途不同,请勿将它们放入公开仓库、URL、截图、邮件正文或群聊。

第二届 Agent Memory Challenge 面向文本、代码与多模态三个赛道设置合计人民币 150,000 元奖金池。奖金仅面向符合赛事规则并通过审核的开源方法榜参赛团队,商业产品榜不参与奖金评选。

奖金人民币 20,000 元。文本、代码与多模态三个赛道分别评选。

二等奖 2 名,每名人民币 8,000 元;三等奖 3 名,每名人民币 3,000 元。

奖金人民币 5,000 元。每个赛道奖金合计 50,000 元,三个赛道合计 150,000 元。

提交申请前请先阅读 API 接入指南并准备完整材料。报名与评测问题可发送至 [email protected]。

Agent Memory Challenge is the second public evaluation cycle of Agent Memory Leaderboard. It is open to researchers, open-source maintainers, and commercial product teams worldwide. Participants provide Add and Search; the platform runs Answer, Eval, result review, and leaderboard publication.

Participation is free, with no team-size requirement. Participants cover the cost of their own APIs, databases, bandwidth, and compute; the platform covers unified Answer, Eval, and evaluation orchestration.

September 20, 2026 · 00:00 (UTC+8)

October 31, 2026 · 23:59 (UTC+8)

November 4, 2026 · 23:59 (UTC+8)

Mid-November 2026

Textual, Coding, and Multimodal Memory

The evaluation type determines the tasks your system receives; the participant division determines where the result is listed. These are two independent dimensions.

In Cycle 2, every participant division must host stable Add / Search APIs. An Eval Key is issued after the submitted online endpoints and materials pass review.

Submit fixed Add / Search endpoints, authentication and operational details. A public repository supports openness and attribution review but cannot replace deployed endpoints.

Submit a fixed product version, Add / Search endpoints, authentication, and capacity details. Internal implementation may remain closed; an Eval Key is issued after approval and results enter the Commercial Products board.

Open-source and commercial entries integrate through participant-hosted Add/Search APIs. Cycle 2 does not accept repository-only or Docker-only submissions for AML deployment. Open-source materials remain part of openness and attribution review.

Freeze a clear evaluation version before applying. Once a formal Full evaluation is accepted, the version may not be replaced or withdrawn because of an unfavorable result.

System name and version, contact details, organization or team, intended evaluation type, method or product description, information approved for public display, and complete submission notes.

Deployed Add/Search endpoints, authentication and capacity declarations, plus a public repository, fixed commit, attribution of prior work, and method changes.

Fixed product version, Add / Search APIs, authentication, a dedicated evaluation credential, capacity, timeout, and rate-limit details.

These credentials are issued by different parties and serve different purposes. Never place either credential in a public repository, URL, screenshot, email body, or group chat.

The second Agent Memory Challenge provides a total prize pool of RMB 150,000 across the Textual, Coding, and Multimodal tracks. Prizes are available only to eligible Open-source Methods teams that pass review; Commercial Products entries are not eligible.

RMB 20,000. The Textual, Coding, and Multimodal tracks are judged separately.

Second Prize: 2 winners, RMB 8,000 each. Third Prize: 3 winners, RMB 3,000 each.

RMB 5,000. Each track awards RMB 50,000, for a total prize pool of RMB 150,000.

Read the API Guide and prepare all required materials before applying. For participation and evaluation questions, contact [email protected].

OVERVIEW

Agent Memory Leaderboard 使用统一的端到端流程评估不同记忆系统。我们不限定你使用的数据库、索引、向量模型或内部架构,只要求 Add / Search 接口符合现行规范,且评测过程和检索结果能够复核。

我们按样本和来源会话调用 Add 接口。

你的系统负责持久化、组织、索引和更新记忆。

我们针对每道题调用 Search 接口并接收排序结果。

我们使用固定流程生成答案、完成评分并汇总结果。

你只负责 Add / Search 环节。回答模型、提示词、评分器、数据集组合、Top K 和汇总规则由我们统一固定,使榜单分数尽可能反映记忆系统本身的能力。

WHO THIS IS FOR

你可以在 Textual、Multimodal 或 Coding Track 中查看 Overall、分项指标、系统版本和发布日期。不同 Track 使用的任务和指标不同,分数不宜跨 Track 直接比较。

实现 Add / Search → 提交准入申请 → 获取 API Key → 运行公开 smoke → 提交 full 正式评测 → 查看私有结果 → 进入公榜审核。

PARTICIPATION WORKFLOW

先提交 Evaluation Access Request,完成参测系统、版本、在线 Add / Search 接口和鉴权方式审核。第二期所有榜单均要求参赛方自行部署接口,AML 不接受仅提交仓库后代为部署。

选择申请新 Key,或输入已有 Key 为其增加版本;填写联系人、系统版本、Add / Search 接口、鉴权、容量和运行限制。开源方法另附仓库与来源披露。

审核通过后,新申请会签发 Leaderboard API Key;新增版本申请会直接绑定到原 Key,无需管理员再次手工录入接口信息。

平台直接校验参赛方提交的在线接口,并按照现行同步规范执行 Add → Search → Answer → Evaluate Smoke 流程,确认接口可用后再进入正式评测。

选择 full 前,必须逐项勾选提交清单:smoke 已通过、API 契约正确、运行说明完整,以及原创性披露和诚信承诺均已完成。

在 Evaluation 页面验证 API Key,选择已绑定的版本,填写唯一的 Run Label 后提交 full 评测。正式任务使用统一的数据集套件和 Top K。

任务依次执行写入、检索、回答、评分和结果汇总。页面会显示当前阶段、完成进度和错误摘要;我们同时保留复核所需的过程记录。

评测结果首先仅对绑定的 API Key 可见。成功完成的 full 任务会进入公榜资格校验和审核队列;审核通过后生成公榜提交记录。

允许复现已有论文或封装已有仓库,但必须披露原始作者、技术报告以及方法改动;未披露的重复代码可能被视为抄袭。发现重复代码时,平台会优先保护本人已披露并先提交的评测结果。重复提交相似或低质量代码、提示词注入、数据或结果操纵、恶意刷榜等行为,均可能取消参赛资格。

EVALUATION MODES

首次接入或接口行为变更后,先使用独立的兼容性 smoke 检查接口规范;接口稳定且版本准备完成后,再提交唯一的正式评测模式 full。

ADD / SEARCH CONTRACT

参赛选手配置 Add 和 Search 的接口地址,请求和响应格式按照现行规范固定,不随 URL 路径变化。生产环境建议使用 HTTPS。URL 中不得包含用户名、密码等凭据,也不得指向私有、回环或链路本地地址。

记忆写入完成后,Add 接口返回 HTTP 200

返回数量不得超过 top_k,超量会判为契约错误

Add 和 Search 必须使用完全相同的 user_id

Add / Search 支持 Token、Bearer 和 X-Api-Key;none 仅用于公开 smoke。Health 接口通过无需鉴权的 GET 请求调用,返回任意 2xx 状态码即表示服务正常。

{

"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",

"messages": [{

"role": "user",

"timestamp": 1704067200000,

"content": "memory text"

}],

"user_id": "eval:run_abc123:locomo:conv-0",

"session_id": "eval:run_abc123:sample:0"

}request_id必填。本次写入请求的唯一标识;成功响应必须原样返回该值。

messages必填。消息按原顺序排列,每条消息包含 role 和非空 content。文本与代码使用字符串;多模态可使用字符串或有序 text/image_url 内容数组。timestamp 可选,单位为 Unix 毫秒。

user_id必填。Search 接口唯一使用的检索范围标识;写入和检索时必须保持一致。

session_id必填。用于标识来源会话,可以用于组织记忆,但不作为 Search 的筛选条件。

未使用字段现行规范不发送 metadata、app_id、agent_id 或 async_mode;同步语义由 Add 完成写入后返回 HTTP 200 保证。

HTTP 200

{

"success": true,

"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",

"user_id": "eval:run_abc123:locomo:conv-0",

"session_id": "eval:run_abc123:sample:0"

}success必填,且必须是布尔值 true。

request_id / user_id / session_id全部必填,并且必须与请求中的值完全一致。

同步完成服务内部可以采用异步处理,但 Add 接口必须等待处理完成后再返回。

HTTP 202仅当提交版本已审核绑定包含 {task_id} 的 Add 状态查询地址时支持;响应须返回 task_id。未绑定时返回 202 属于契约错误。响应不需要 memory_ids。

{

"query": "Which answer best matches the memory?",

"options": ["A. First answer", "B. Second answer"],

"user_id": "eval:run_abc123:locomo:conv-0",

"top_k": 100

}query必填。按照原文检索相关记忆;多模态可使用与 Add 相同的 text/image_url 内容数组。不得替换为最终答案或使用评测金标。

options可选。选择题(含 Streaming)在 Search 顶层传入选项文本数组;开放题不发送,不包含金标答案。

user_id必填。只能在该 user_id 对应的记忆范围内检索。

top_k必填。返回的记忆数量不得超过该值;正式外部评测固定为 100。

请求格式现行规范不发送 filters、rerank 或 keyword_search。

{

"data": [{

"id": "mem_1",

"content": "remembered fact text",

"score": 0.87,

"created_at": "2026-07-01T12:00:00Z"

}]

}data必填,类型为数组。不要增加 items 包装层,也不要直接返回顶层数组。

id必填,非空字符串,用于稳定标识该条记忆。

content必填,非空字符串;多模态可返回 text/image_url 内容数组。平台保留原始候选并传入统一 Answer 流程。

score可选,数值类型。数值越大应表示相关性越高。

created_at可选,用于记录记忆的来源时间或持久化时间。

其他字段我们只读取上述字段;metadata 等未声明字段会被忽略。

ERROR HANDLING

参测接口应使用标准 HTTP 状态码,并返回便于排查问题且不含密钥的错误信息。平台业务错误通常采用 {"detail":{"reason":"..."}} 格式;字段校验失败时,HTTP 422 会返回结构化错误明细。

Add 遇到网络错误及 408、409、425、429、500、502、503、504、524 时会有界重试;409 不单独代表成功。Search 遇到网络错误及 408、425、429、500、502、503、504 时会有界重试;400、401、403、404、409、422 等永久性 4xx 不重试。重试和断点续跑保持同一 Add 的 request_id 与请求体。

即使 HTTP 状态码为 200,只要 Add 未返回 success=true、三个 ID 未正确返回,或者 Search 未返回 data 数组、某条记录缺少 id / content,当前阶段都会立即失败。

DATA, SECURITY & PRIVACY

你的接口只会收到当前任务所需的记忆片段、user/session 标识和检索问题。我们不会提供金标答案、评分依据或完整数据集下载。

user_id 是 Search 接口唯一使用的检索范围标识,存储和检索时必须完全一致。session_id 只用于组织来源会话。禁止跨 user_id 返回记忆。

Memory System Key 通过受控的申请流程提交并加密保存,任务信息中只保留不可读的引用。密钥不会出现在邮件、公榜或公开 API 响应中。

我们只连接通过网络校验的公开 HTTP(S) 接口,并拒绝 URL 中包含凭据,或者指向私有、回环、链路本地地址的目标。

我们会保留复核所需的请求结果、耗时、错误、候选记忆和接口格式校验记录,用于确认评测是否完整,以及结果是否符合公榜条件。

私有任务和结果仅对绑定的 Leaderboard API Key 可见。公榜只展示审核通过的系统、版本、分数和必要的评测信息。

评测数据及其派生副本只能用于完成当前任务,不得用于模型训练、微调、产品分析、数据集重建或对外传播。请仅向必要人员开放访问权限,避免记录不必要的请求正文,并在任务完成后 30 天内删除相关数据;如需延长保留时间,必须事先获得我们的书面同意。

EVALUATION & RESULTS

每个记忆分块都必须在返回 HTTP 200 前完成持久化,并且能够立即检索。

我们会检查 data 数组、必填字段和 Top K,并按照接口返回的顺序接收候选记忆。

通过校验的候选记忆会进入统一的回答模型和提示词流程。

我们按照题型使用固定的评分规则,并汇总各数据集和 Overall 结果。

任务完成后,你可以使用绑定的 API Key 查看运行状态、分项结果、错误摘要和任务信息。

full 结果通过资格校验和审核后,会生成公榜提交记录并展示在公开榜单中。

结果必须来自成功完成的 full 模式固定套件,且所有评测任务均执行成功。统一回答模型、评测规范、pipeline code hash、dataset bundle hash 和题量记录必须完整,并与现行发布基线一致。结果不得重复提交,同时还需通过平台审核。

BENCHMARK SUITE

按能力维度浏览基准数据集。每张卡片均链接到源仓库,并说明任务格式、指标与评测备注。

OVERVIEW

Agent Memory Leaderboard compares memory systems under a consistent end-to-end evaluation. The platform does not prescribe a database, index, embedding model, or internal architecture. It requires a conformant external API and verifiable evidence for retrieval and audit.

The platform calls Add by sample and source session.

Your system persists, organizes, indexes, and updates memory.

The platform calls Search per question and receives ranked results.

A fixed platform workflow generates answers, scores them, and aggregates results.

Participant-controlled behavior is limited to Add / Search. The platform locks the answer model, prompts, evaluators, dataset suite, Top K, and aggregation rules so that score differences primarily reflect the memory system.

WHO THIS IS FOR

Review Overall scores, metric breakdowns, system versions, and publication dates within the Textual, Multimodal, or Coding track. Each track has a distinct task and metric contract; scores are not comparable across tracks.

Implement Add / Search → submit an access request → receive an API Key → run public smoke → submit the full evaluation → review private results → enter public review.

PARTICIPATION WORKFLOW

Open-source Methods and Commercial Products submissions both use participant-hosted Add/Search APIs and follow the same review, Key issuance, and public Smoke flow. AML does not deploy repository-only submissions.

All entries provide deployed Add/Search endpoints, authentication, capacity, and operational details. Open-source Methods entries additionally disclose the public repository, original authors, technical report, and method changes.

Approval either issues a new Leaderboard API Key or binds the submitted version to the verified existing key without administrator re-entry.

The platform checks the participant-operated endpoints and runs the synchronous Add → Search → Answer → Evaluate Smoke flow. AML does not build or deploy participant systems.

Before selecting full, confirm the interactive checklist: smoke passed, API contract followed, run instructions complete, original work disclosed, and no integrity violations.

Verify the API Key on the Evaluation page, select a bound version, assign a unique Run Label, and submit the full evaluation. Formal jobs use the fixed suite and platform Top K.

The job proceeds through ingestion, retrieval, answering, scoring, and aggregation. The interface reports stage, progress, and error summaries while the platform retains evidence required for review.

Results are initially visible only to the bound API Key. A successful full job enters the public-eligibility gate and review queue; approval creates a public submission.

Reproductions of papers or existing repositories are allowed only with full attribution to the original authors, technical report, and all method changes. Undisclosed reuse may be treated as plagiarism. When duplicate code is discovered, the platform prioritizes the participant's first disclosed submission. Repeated near-duplicate or low-quality code, prompt injection, benchmark or result manipulation, malicious behavior, and other leaderboard abuse may cancel eligibility.

EVALUATION MODES

Use the separate compatibility smoke after initial integration or an API behavior change. Submit full, the only formal participant mode, after the version is ready for publication.

ADD / SEARCH CONTRACT

Participants configure the Add and Search URLs; request and response schemas are fixed and do not vary by path. Production deployments should use HTTPS. URLs must not embed credentials or resolve to private, loopback, or link-local addresses.

Add succeeds only with HTTP 200 after persistence

The response must not exceed top_k; excess items are a contract error

Add and Search must use the identical isolation boundary

Add / Search support Token, Bearer, and X-Api-Key; none is limited to public smoke. Health is an unauthenticated GET where any 2xx means healthy.

{

"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",

"messages": [{

"role": "user",

"timestamp": 1704067200000,

"content": "memory text"

}],

"user_id": "eval:run_abc123:locomo:conv-0",

"session_id": "eval:run_abc123:sample:0"

}request_idRequired. Unique identifier for this chunk request; the success response must echo it exactly.

messagesRequired ordered array. Each item includes role and non-empty content. Textual and Coding use strings; Multimodal may use strings or ordered text/image_url content parts. timestamp is optional Unix milliseconds.

user_idRequired. The sole retrieval-isolation field; Search must use the identical value.

session_idRequired. Identifies the source session and may be used for grouping, but is not a Search filter.

Fields not sentThe current contract omits metadata, app_id, agent_id, and async_mode. Synchronous behavior is guaranteed by returning HTTP 200 only after Add completes.

HTTP 200

{

"success": true,

"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",

"user_id": "eval:run_abc123:locomo:conv-0",

"session_id": "eval:run_abc123:sample:0"

}successRequired and must be the boolean value true.

request_id / user_id / session_idAll are required and must match the request byte for byte.

Synchronous completionInternal work may be asynchronous, but the endpoint must wait for completion before responding.

HTTP 202Supported only when the submitted version has an approved Add-status URL containing {task_id}; the response must return task_id. Without that binding, 202 is a contract error. memory_ids are not required.

{

"query": "Which answer best matches the memory?",

"options": ["A. First answer", "B. Second answer"],

"user_id": "eval:run_abc123:locomo:conv-0",

"top_k": 100

}queryRequired original question. Multimodal may use the same text/image_url content parts as Add. Do not replace it with a final answer or use benchmark gold data.

optionsOptional. Choice text is sent as a top-level string array for multiple-choice questions, including Streaming; omitted for open questions and never contains the gold answer.

user_idRequired. Retrieve only from memory stored under this exact value.

top_kRequired. The response must not exceed this number; formal external evaluations use 100.

Fixed schemaThe current contract does not send filters, rerank, or keyword_search.

{

"data": [{

"id": "mem_1",

"content": "remembered fact text",

"score": 0.87,

"created_at": "2026-07-01T12:00:00Z"

}]

}dataRequired array. Do not add an items wrapper or return a top-level array.

idRequired non-empty string that stably identifies the memory.

contentRequired non-empty string; Multimodal may return text/image_url content parts. The original candidate is retained and passed to Answer.

scoreOptional number. Higher values should indicate greater relevance.

created_atOptional source or persistence timestamp.

Other fieldsThe platform reads only the declared fields above; undeclared fields such as metadata are ignored.

ERROR HANDLING

Participant endpoints should use standard HTTP status codes and return actionable errors without secrets. Platform business errors normally use {"detail":{"reason":"..."}}; HTTP 422 provides structured field-validation details.

Add retries 408, 409, 425, 429, 500, 502, 503, 504, and 524. Search retries 408, 425, 429, 500, 502, 503, and 504. Network timeouts and transport failures are also eligible.

Even with HTTP 200, the stage fails if Add omits success=true or mis-echoes an ID, or if Search omits the data array or an item lacks id / content.

DATA, SECURITY & PRIVACY

Participant endpoints receive only the memory chunks, user/session identifiers, and retrieval questions needed for the current job. Gold answers, scoring criteria, and bulk dataset downloads are not provided.

user_id is the sole Search-isolation field and must match exactly during storage and retrieval. session_id is only for source-session organization. Cross-user_id retrieval is prohibited.

The Memory System Key is submitted through a controlled request flow and stored encrypted; job metadata contains only an opaque reference. Secrets do not appear in email, public rankings, or public API responses.

The platform connects only to network-validated public HTTP(S) endpoints and rejects URLs with embedded credentials or targets resolving to private, loopback, or link-local addresses.

The platform retains request outcomes, latency, errors, returned memories, and contract-validation evidence required to verify evaluation integrity and public eligibility.

Private jobs and results are visible only to the bound Leaderboard API Key. The public board shows approved systems, versions, scores, and required evaluation metadata.

Evaluation data and derived copies may be used only to complete the current job. Do not use them for training, fine-tuning, product analytics, dataset reconstruction, or redistribution. Restrict access to authorized personnel, avoid unnecessary payload logging, and delete the data within 30 days after job completion unless the platform approves another retention period in writing.

EVALUATION & RESULTS

Each chunk must be persisted and searchable before its HTTP 200 response.

The platform validates the data array, required fields, and Top K, then accepts candidates in participant order.

Valid candidate memories enter the locked platform Answer model and prompt workflow.

The platform applies fixed scoring contracts by question type and aggregates dataset and Overall results.

After completion, the bound API Key can inspect run status, breakdowns, error summaries, and job metadata.

An eligible full result creates a public submission only after qualification checks and review.

The result must come from a successful full run on the fixed suite; every evaluation task must succeed; the Answer model, evaluation contract, pipeline code hash, dataset bundle hashes, and question counts must be complete and match the current release baseline; and the result must be unique and pass platform review.

BENCHMARK SUITE

Browse benchmark datasets by dimension. Each card links to its source repository with task format, metrics, and evaluation notes.

150 software-engineering tasks run under relevant and noisy memory conditions, for 300 scored attempts.

Not currently used as the scored AML Coding benchmark.

Cycle 2 uses participant-hosted Add and Search APIs for the Textual, Coding, and Multimodal tracks. Implement the formats below, apply for and bind an evaluation key, pass Smoke, and then submit a Full evaluation.

Use this ordered array wherever a Multimodal field carries mixed text and images. Textual and Coding use plain strings.

[

{"type": "text", "text": "caption or text evidence"},

{

"type": "image_url",

"image_url": {"url": "data:image/jpeg;base64,..."}

}

]Textual and Coding use string content. For Multimodal messages[].content, use the ordered array shown above.

{

"request_id": "eval:<run_id>:locomo_refined:conv-0:chunk-0",

"messages": [{

"role": "user",

"timestamp": 1704067200000,

"content": "raw memory text"

}],

"user_id": "eval:<run_id>:locomo:conv-0",

"session_id": "eval:<run_id>:sample:0"

}Textual, Coding, and Multimodal use the same response schema.

{

"success": true,

"request_id": "eval:<run_id>:locomo_refined:conv-0:chunk-0",

"user_id": "eval:<run_id>:locomo:conv-0",

"session_id": "eval:<run_id>:sample:0"

}request_id received in the Add request.

user_idRequiredExact user_id received in the Add request.

session_idRequiredExact session_id received in the Add request.

Textual and Coding use a string query. For a Multimodal query, use the ordered array shown above.

{

"query": "Which answer best matches the memory?",

"options": ["A. First answer", "B. Second answer"],

"user_id": "eval:<run_id>:locomo:conv-0",

"top_k": 100

}Textual and Coding return string content. For Multimodal data[].content, use the ordered array shown above.

{

"data": [

{

"id": "mem_123",

"content": "remembered fact text",

"score": 0.87,

"created_at": "2026-07-01T12:00:00Z"

}

]

}Your API receives benchmark content during each run.

Your API receives evaluation memories and questions.

Use this data only for the evaluation. Do not train on it, analyze it, or share it.

Keep it private, avoid storing logs, and delete it within 30 days after the run.

A compact map of the product surface so users know where to go after their endpoints are ready.

Overview and benchmark preview.

Public rankings by track.

Verify key, start eval jobs, inspect private results and attribution.

User workflow, API contract, security, and result publication.

Propose a documented benchmark with its task, provenance, package, and license.

Help extend the Agent Memory Leaderboard by proposing a well-documented benchmark. Include the dataset context, evaluation goal, and provenance so reviewers can assess fit, safety, and reproducibility.