atlas-agent-control-plane
AtlasAgent - an auditable AI agent control plane: evidence-backed memory, governed tool runtime, checkpoint DAG recovery, and a 55-chapter engineering tutorial. FastAPI / Next.js PWA / Textual TUI
Links
README
From the repo.
Why I built this
When I deployed OpenClaw for the first time, the experience was unlike anything before. But a question followed me for weeks: these general agents are already this strong — is there any point in building another one?
The answer I settled on: a general agent is a good knife; a vertical agent is a machine assembled for one workflow. They don't compete. So this project is my attempt to build the machine — not a chat shell wired to an LLM, but a system that plans a task, executes it step by step, calls tools, observes results, and answers with evidence it can point to.
Along the way I kept running into the same wall: the demos were impressive, but the moment you ask "where did this fact come from?" or "what did the agent do while I was asleep?" — nothing holds up. AtlasAgent is what that wall looks like when you take it seriously. Every fact carries its source. Every tool call passes a gate before it runs. Every long task leaves behind checkpoints you can resume from. And when a task claims it's done, an acceptance command has to exit 0 before the system believes it.
One more thing: the whole system is paired with a 62-chapter tutorial that rebuilds it from an empty directory — because reading source code tells you what, but rarely why. The trade-offs, the dead ends, the reasons things are the way they are: that's what the tutorial keeps. 写在前面 has the full story, in my own words.
Six boundaries, taken as first principles
| Memory with provenance | Facts pass a write gate before they reach the model — source, scope, expiry. Then Ebbinghaus decay, consolidation, conflict resolution, and a typed memory graph. A memory should be able to forget. |
| RAG that cites its work | Query rewriting, multi-query RRF fusion, parent-document retrieval, LLM reranking with a lexical fallback. Every citation carries a confidence score; every query leaves a retrieval trace. |
| A tool runtime with a gate | Risk tier → approval → idempotency → timeout → redaction → artifact. The handler runs only after the gate says yes. Add result caching, dependency-aware batching, retries with fallback, per-step budgets. |
| A task must earn "done" | Before a task completes, three gates run in order: the acceptance command (exit 0 or it isn't done), a scope audit (only files the plan declared), and an LLM coverage review of the tests. Three gates, one shared retry budget — no infinite arguing with the model. |
| Checkpoints over restarts | State hashes and environment fingerprints. A task that dies at step 38 resumes from step 17. |
| Events are the only truth | Task state is a rebuildable view; checkpoints are verified recovery points. Everything the agent ever did is a queryable record. |
Run it
Just Node 22+ — no Docker, no PostgreSQL, no Redis:
git clone https://github.com/malevrigns/atlas-agent-control-plane.git
cd atlas-agent-control-plane
cd backend/api-ts && pnpm install && pnpm dev
# or from the repo root: node scripts/quickstart-ts.mjs
That's SQLite plus the TypeScript Hono control plane. Open http://localhost:8000/api/status.
For the full stack — Web workbench, Postgres, Redis, Nginx, sandbox container — you need Git, Docker and Docker Compose v2:
cp .env.example .env
BUILD=true ./scripts/start.sh # Git Bash / WSL
# Windows PowerShell: $env:BUILD="true"; ./scripts/start.ps1
Open http://localhost:8088 — the Web workbench. The start script prints the API key you'll log in with. No LLM key? Everything still boots; a local hash embedding powers offline mode, and any OpenAI-compatible endpoint (DeepSeek, Qwen, DashScope, Ollama…) plugs in via one line in backend/api/config/llm.yaml.
Calling the control plane directly
ATLAS_KEY="$(sed -n 's/^ATLAS_API_KEY=//p' .env)"
# a structured task, with its acceptance command declared up front
curl -X POST http://localhost:8088/api/control-plane/tasks \
-H "X-Atlas-API-Key: ${ATLAS_KEY}" -H "Content-Type: application/json" \
-d '{"title": "Upgrade deps", "goal": "A verifiable upgrade",
"acceptance_criteria": ["pytest exits 0"], "project_id": "atlas"}'
# a tool call — risk, approval and idempotency are checked before the handler runs
curl -X POST http://localhost:8088/api/agent-core/tools/draft_plan/invoke \
-H "X-Atlas-API-Key: ${ATLAS_KEY}" -H "Content-Type: application/json" \
-d '{"arguments": {"task": "Check delivery quality"}, "project_id": "atlas",
"idempotency_key": "demo-001"}'
# every invocation left a record you can query
curl -H "X-Atlas-API-Key: ${ATLAS_KEY}" \
"http://localhost:8088/api/control-plane/tool-invocations?project_id=atlas"
More: Memory & Tool Control Plane · RAG & Skills · Architecture
Configuration, security, and local development
Key settings (see .env.example for the rest): LLM_API_KEY (optional), ATLAS_API_KEY (auto-generated), RAG_VECTOR_BACKEND=pgvector|qdrant, RAG_EMBEDDING_PROVIDER=auto|local_hash, NGINX_PORT (default 8088). The built-in key is a single-tenant boundary; internet-facing deployments need TLS, OIDC/RBAC and rate limiting in front. MCP and A2A transports are off or allowlist-only by default.
Development, after docker compose up -d postgres redis:
| Module | Command |
|---|---|
| API | cd backend/api-ts && pnpm install && pnpm dev — docs/TYPESCRIPT_RUNTIME.md |
| Web | cd frontend/web && pnpm install && pnpm dev |
| TUI | cd frontend/tui && uv sync && ATLAS_API_URL=http://localhost:8000 uv run atlas-tui |
| Sandbox | cd backend/sandbox && docker build -t atlas-sandbox . && docker run -d -p 127.0.0.1:8100:8100 -e SANDBOX_AUTH_ENABLED=true atlas-sandbox |
Tests: cd backend/api-ts && pnpm test.
The repository
atlas-agent-control-plane/
├── frontend/
│ ├── web/ Next.js client (PWA)
│ └── tui/ Textual terminal client
├── backend/
│ ├── api-ts/ TypeScript Hono control plane
│ ├── sandbox-ts/ TypeScript sandbox (files / shell / VNC)
│ ├── api/ legacy Python (not the runtime)
│ └── sandbox/ legacy Python (not the runtime)
├── docs/ Deep-dive documentation
├── tutorial/ 62-chapter engineering tutorial (62,000+ lines)
├── scripts/ start.sh / start.ps1 / stop.sh
└── docker-compose.yml
Issues and PRs are welcome — CONTRIBUTING.md has the setup, commit conventions, and PR flow.
为什么做这个
第一次部署 OpenClaw 的时候,体验是前所未有的。但随之而来的问题困扰了我很久:这些通用 Agent 已经这么强了,还有必要从零做一个新的吗?
后来我想明白了:通用 Agent 像一把好刀,垂类 Agent 更像一台按业务流程装配好的机器。两者并不冲突。所以这个项目就是我去造那台机器的尝试——不是一个接了 LLM 接口的聊天壳,而是一个能围绕任务做规划、逐步执行、调用工具、观察结果、最后用可指认的证据回答的系统。
做的过程中反复撞到同一堵墙:演示都很惊艳,但只要你问一句「这条事实从哪来的」「我睡着的时候 Agent 到底干了什么」,就没有一样东西站得住。AtlasAgent 就是把这堵墙当真之后的样子:每条事实带来源,每个工具调用先过门禁,每个长任务留下可以恢复的 Checkpoint,任务说自己完成了,验收命令必须 exit 0,系统才信它。
还有一件事:整个系统配了一套 62 章的教程,从空目录开始把它重新造一遍。因为读源码能知道「是什么」,很难知道「为什么」——那些取舍、死路、和「为什么它是现在这个样子」,教程里都留着。写在前面 里有完整的来龙去脉。
六个边界,当成第一性约束
| 带来源的记忆 | 事实先过写入门禁才到模型——来源、作用域、有效期,一个都不能少。再往上叠艾宾浩斯衰减、自动巩固、冲突消解、类型化记忆图谱。记忆应该会遗忘。 |
| 会引用出处的 RAG | 查询改写、多查询 RRF 融合、父文档检索、LLM 重排(词法降级兜底)。每条引用带置信度,每次检索落审计。 |
| 有门禁的工具运行时 | 风险分级 → 审批 → 幂等 → 超时 → 脱敏 → 制品化。门禁说可以,handler 才执行。另有结果缓存、依赖分批、重试降级、步骤预算。 |
| 「完成」要挣来的 | 任务完成前,三道关依次跑:验收命令(exit 0 才算完)、范围审计(只改计划声明的文件)、LLM 覆盖度评审。三关共享重试额度——不和模型无限拉扯。 |
| Checkpoint 优先于重启 | 状态哈希 + 环境指纹。第 38 步挂了,从第 17 步恢复。 |
| 事件是唯一事实源 | 任务状态只是可重建的视图,Checkpoint 是验证过的恢复点。Agent 干过的每件事都是一条可查询的记录。 |
跑起来
只要 Node 22+,不用 Docker、不用 PostgreSQL、不用 Redis:
git clone https://github.com/malevrigns/atlas-agent-control-plane.git
cd atlas-agent-control-plane
cd backend/api-ts && pnpm install && pnpm dev
# 或仓库根目录:node scripts/quickstart-ts.mjs
SQLite + TypeScript Hono 控制平面。打开 http://localhost:8000/api/status 就能调。
要完整形态(Web 工作台、Postgres、Redis、Nginx、沙箱容器),才需要 Git、Docker、Docker Compose v2:
cp .env.example .env
BUILD=true ./scripts/start.sh # Git Bash / WSL
# Windows PowerShell: $env:BUILD="true"; ./scripts/start.ps1
打开 http://localhost:8088 就是 Web 工作台,启动脚本会打印登录用的 API Key。没配模型密钥也能跑——本地哈希 embedding 撑起离线模式;要接模型的话,DeepSeek、Qwen、DashScope、Ollama,任何 OpenAI 兼容接口在 backend/api/config/llm.yaml 里改一行就行。
直接调用控制平面
ATLAS_KEY="$(sed -n 's/^ATLAS_API_KEY=//p' .env)"
# 一个结构化任务,验收命令一开始就声明好
curl -X POST http://localhost:8088/api/control-plane/tasks \
-H "X-Atlas-API-Key: ${ATLAS_KEY}" -H "Content-Type: application/json" \
-d '{"title": "升级依赖", "goal": "可验证升级",
"acceptance_criteria": ["pytest 退出码 0"], "project_id": "atlas"}'
# 一次工具调用——handler 执行前先过风险、审批、幂等检查
curl -X POST http://localhost:8088/api/agent-core/tools/draft_plan/invoke \
-H "X-Atlas-API-Key: ${ATLAS_KEY}" -H "Content-Type: application/json" \
-d '{"arguments": {"task": "检查交付质量"}, "project_id": "atlas",
"idempotency_key": "demo-001"}'
# 每次调用都留了记录,随时可以查
curl -H "X-Atlas-API-Key: ${ATLAS_KEY}" \
"http://localhost:8088/api/control-plane/tool-invocations?project_id=atlas"
配置、安全与本地开发
常用配置(其余见 .env.example):LLM_API_KEY(可选)、ATLAS_API_KEY(自动生成)、RAG_VECTOR_BACKEND=pgvector|qdrant、RAG_EMBEDDING_PROVIDER=auto|local_hash、NGINX_PORT(默认 8088)。内置 API Key 是单租户边界;公网多用户部署需要 TLS、OIDC/RBAC 和限流。MCP 与 A2A 传输默认关闭或仅白名单。
本地开发,先 docker compose up -d postgres redis:
| 模块 | 命令 |
|---|---|
| API | cd backend/api-ts && pnpm install && pnpm dev — docs/TYPESCRIPT_RUNTIME.md |
| Web | cd frontend/web && pnpm install && pnpm dev |
| TUI | cd frontend/tui && uv sync && ATLAS_API_URL=http://localhost:8000 uv run atlas-tui |
| Sandbox | cd backend/sandbox && docker build -t atlas-sandbox . && docker run -d -p 127.0.0.1:8100:8100 -e SANDBOX_AUTH_ENABLED=true atlas-sandbox |
测试:cd backend/api-ts && pnpm test。
仓库结构
atlas-agent-control-plane/
├── frontend/
│ ├── web/ Next.js 客户端(PWA)
│ └── tui/ Textual 终端客户端
├── backend/
│ ├── api-ts/ TypeScript Hono 控制平面
│ ├── sandbox-ts/ TypeScript 沙箱(文件 / Shell / VNC)
│ ├── api/ 旧 Python(不再作为运行时)
│ └── sandbox/ 旧 Python(不再作为运行时)
├── docs/ 专题深潜文档
├── tutorial/ 62 章中文工程教程(62000+ 行)
├── scripts/ start.sh / start.ps1 / stop.sh
└── docker-compose.yml
欢迎 Issue 与 PR——流程见 CONTRIBUTING.md。
Collected info
- ★ 105 stars
- ⎇ 1 forks
- Language: Python
- Source updated: 9/15/2026