Sovereign-OS
Constitution-first AI orchestration: one Charter (YAML) defines mission, budget & rules. CEO plans → CFO approves → Ledger tracks every cent & token → Auditor scores. 16 workers, Stripe, MCP. Think. Audit. Execute.
Links
README
From the repo.
Governance for an autonomous AI workforce — runs on your machine, on your keys.
Submit a goal: Sovereign-OS plans it, checks the budget, executes with built-in workers, and delivers a cryptographically verified result.
Open source • self-hosted • bring your own keys • no fund custody
Quick Start • Architecture • Features • Deployment • Configuration • Custom Workers • Docs
Overview
Sovereign-OS is not another chatbot wrapper or agent framework. It is an operating system for autonomous AI work: a governance layer that enforces budget, quality, and permissions before any token is spent or any task is executed.
The core contract is simple. One YAML file — the Charter — declares the agent's mission, spending limits, KPIs, and allowed capabilities. Everything else — planning, approval, execution, auditing, payment — flows from that file.
Charter → CEO (plan) → CFO (approve budget) → Workers (execute) → Auditor (verify) → Ledger (record)
The Ledger is append-only. The Auditor is cryptographically bound. Neither can be overridden at runtime.
Why Sovereign-OS?
| Typical agent frameworks | Sovereign-OS | |
|---|---|---|
| Cost control | API key = burn until empty. | Every cent and token is tracked. The budget is checked before every task. Daily caps + circuit breaker enforced. |
| Quality | Hope the output is correct. | Every task verified against Charter KPIs. Fail audit → TrustScore drops. Proof hash on every report. |
| Permissions | All-or-nothing. | Agents start sandboxed. They earn capabilities (SPEND_USD, CALL_API, WRITE_FILES) via TrustScore. |
| Monetization | None. | Built-in job queue, Stripe integration, ingest from any source, webhook on delivery. |
Quick Start
Everything runs on your machine, on your keys — nothing is sent to a central service.
Docker — the whole workspace:
git clone https://github.com/Justin0504/Sovereign-OS.git && cd Sovereign-OS
export ANTHROPIC_API_KEY=sk-ant-… # or OPENAI_API_KEY
docker compose up -d redis web # web console on :8000
# Open http://localhost:8000
CLI — run a single mission in 30 seconds:
pip install -e .
sovereign run --charter charter.example.yaml "Summarize the current AI agent landscape."
Output: task plan → budget check → execution → audit report → ledger entry.
Web dashboard — submit missions and watch the governed run:
pip install -e ".[llm]"
cp .env.example .env
# Edit .env: set ANTHROPIC_API_KEY or OPENAI_API_KEY (optionally STRIPE_API_KEY)
python -m sovereign_os.web.app
# Open http://localhost:8000
From the dashboard you can submit missions, inspect the job queue, retry jobs, and monitor token usage and the audit trail in real time. Deploying to a server or a custom domain? See docs/DEPLOY.md.
Architecture
Five layers. Data flows down; accountability flows back up.
Charter (YAML)
└─ Governance Engine
├─ CEO (Strategist) — decomposes goals into ordered tasks; prunes to budget & re-plans on failure
├─ CFO (Treasury) — approves/rejects per-task spend; daily cap, runway, margin + session circuit breaker
└─ Worker Registry — routes each task to the right worker by skill
└─ Workers — execute; emit TaskResult with token usage
└─ Auditor — verifies output against Charter KPIs (category-tuned rubric); signed AuditReport
└─ Ledger (UnifiedLedger) — append-only USD + token accounting
Mermaid diagram
flowchart LR
Charter --> CEO[CEO · Plan]
CEO --> CFO[CFO · Approve]
CFO --> Registry[Workers]
Registry --> Ledger[Ledger]
Registry --> Auditor[Auditor]
Auditor --> CFO
Key design decisions:
|
|
Marketplace oversight & delivery categories
Beyond running its own missions, Sovereign-OS acts as a governance layer over the agent-task economy — in both directions, applying the same two gates (budget + delivery quality) wherever work flows.
INBOUND human/agent posts ─▶ ingest (TaskBounty · StacksTasker · BotBounty · ClawTasks · APB/x402)
─▶ [CFO budget] ─▶ governed workers ─▶ [Auditor quality] ─▶ submit
OUTBOUND Sovereign-OS posts ─▶ [CFO budget gate] ─▶ fund escrow (RentAHuman)
─▶ worker delivers ─▶ [Auditor quality gate] ─▶ release | dispute
- Inbound sources (
ingest_bridge/) pull open tasks from real marketplaces (live-validated APIs); Sovereign-OS governs its own agents doing the work. Includes an APB (Agent Payment Bounty) source that discovers x402-native bounties from publishers'/.well-known/bounties.json— the standard the x402 Foundation rail is built on (see docs/AUTONOMY.md). - Outbound oversight (
oversight/) posts tasks for external workers: the CFO approves the spend and reserves escrow before funding, and the Auditor must pass the deliverable before payment is released (else it's disputed/refunded). All money-moving calls are dry-run until explicitly set live; runpython -m sovereign_os.oversight.rentahuman_preflightfor a go-live check.
A task-category backbone (agents/categories.py) ties delivery together:
each category (coding, data, design, email, research, writing, automation) maps
to a top-tier worker, a risk-scaled budget ceiling (CategoryBudgetPolicy),
a permission tier earned per category (per-category TrustScore), and the
connectors it needs (connectors/, MCP / built-in / HTTP, with readiness).
Inspect it with sovereign categories and sovereign connectors.
See docs/OVERSIGHT.md, docs/CATEGORIES.md, docs/CONNECTORS.md, docs/CLAWTASKS.md, docs/COST.md, and docs/X402.md.
Capability map
INBOUND platforms GOVERNANCE AGENTIC WORKERS (by category)
TaskBounty · StacksTasker ──ingest──▶ ┌───────────────┐ ──route──▶ coding → read·write·run_tests·submit_PR
BotBounty · ClawTasks │ CFO budget gate│ research → web_fetch
│ (per-category │ data → web_fetch
OUTBOUND (agent hires) │ ceilings) │ design → read_figma · generate_image
RentAHuman escrow ◀──post/fund────── │ Auditor quality│ writing → draft → self-critique → revise
release | dispute ◀──gate───────────│ (rubric) │ ─────────────────────────────────────
└───────────────┘ tools dispatched via connectors/, code
Ledger (append-only, thread-safe) · per-model/category cost · TrustScore earned per category · Docker sandbox
Every task is categorized → routed to a top-tier worker → budget-gated by a per-category ceiling → executed with real tools (web/repo/Figma/image, in a Docker sandbox for untrusted code) → quality-gated before delivery or payout. All money/exec is dry-run by default.
Try it (no keys, no funds):
python examples/coding_bounty_demo.py # bug-fix: route → budget → read·write·test·PR → audit
python examples/oversight_demo.py # outbound: budget gate + quality gate (release vs dispute)
python examples/category_demo.py # category → worker · budget · permission · connectors
python examples/oversight_e2e.py # full readiness check (inbound + outbound + preflights)
sovereign categories # the delivery-category table
sovereign connectors # connector readiness + required MCP servers
Features
Governance
- Charter-driven: mission, competencies, KPIs, fiscal boundaries — all in one YAML.
- CEO (Strategist) decomposes natural-language goals into executable task plans with dependencies. Reactive:
prune_to_budgetsheds lowest-value tasks to fit a budget without breaking the dependency DAG;corrective_taskbuilds a high-priority retry carrying the audit's failure reason and fix. - Profitability-first task selection — before spending any compute, an ingest screen (
governance/economics.py) estimates a task's fully-loaded cost (LLM by category × complexity + settlement fee + gas) and skips anything that can't clear the margin floor. Stops the classic autonomous-agent failure mode where fees/gas turn shipped work into a net loss. Opt in withSOVEREIGN_PROFIT_SCREEN=true; tune withSOVEREIGN_{SETTLEMENT_FEE_RATIO,GAS_COST_CENTS,MIN_MARGIN_RATIO}. - Reactive self-repair — a task that fails audit is automatically re-run with the Auditor's reason + fix folded into its brief (
SOVEREIGN_MAX_REPAIR_ATTEMPTS=N), recovering quality with no human in the loop. Only a fully-passing mission reaches delivery/settlement — a failed audit never gets submitted to a platform or charged. - CFO (Treasury) enforces
max_task_cost_usd,daily_budget_usd,runway_days, andmin_job_margin_ratiobefore any task runs. A runtime SpendCircuitBreaker adds fast-fail during a session — it halts the loop when cumulative spend hits a session ceiling, audits fail past a streak limit, or ROI collapses (the guard that stops runaway agent loops the pre-flight gates miss). Configure viaSOVEREIGN_SESSION_CEILING_CENTS,SOVEREIGN_MAX_CONSECUTIVE_FAILURES,SOVEREIGN_ROI_FLOOR(all default off); watch it live on the dashboard Guardrails tab (session spend vs. ceiling, failure streak, ROI, one-click reset). - Permissions — TrustScore-gated capabilities (
READ_FILES,WRITE_FILES,EXECUTE_SHELL,SPEND_USD,CALL_EXTERNAL_API), earned per category, with graduated autonomous-spend ceilings. High-risk grants use just-in-time leases: scoped to one task, bounded by TTL and use-count, auto-revoked on completion (grant_lease/use_lease/revoke_task_leases). Active leases and per-agent trust are shown live on the dashboard Guardrails tab. - Two dashboards, same guardrails — the web dashboard (Guardrails tab + per-task quality scorecard) and a Textual terminal Command Center (
python -m sovereign_os.ui.app) both surface the CFO circuit breaker (session spend vs. ceiling, failure streak, ROI; pressbto reset), active JIT leases, and the category-rubric breakdown streamed live under each audit verdict. - Prometheus metrics — all three guardrails export on
/metricsfor Grafana: breaker session spend/ceiling/ROI/trips, active JIT leases, per-agent trust, and per-category audit-quality histograms (overall + per rubric criterion). See docs/METRICS.md for the metric list and PromQL starter panels.
Execution
- 16 built-in workers:
summarize,research,reply,write_article,write_email,write_post,meeting_minutes,translate,rewrite_polish,collect_info,extract_structured,spec_writer,solve_problem,assistant_chat,code_assistant,code_review. - Verification-driven coding — the coding worker runs a harness-enforced loop (
run_with_verified_tools): it won't accept "done" until the test suite actually passes. A premature answer bounces back with the failing test output and the model must fix and re-verify — code that can't reach green is markedtests_verified=false/success=false, so broken work never ships to a paid bounty. Gates only when execution is enabled (SOVEREIGN_CODE_EXEC_ENABLED); otherwise it's a no-op skip. - Multi-model: Strategist and workers can use different backends (e.g. GPT-4o for planning, Claude for execution).
- Pluggable agent backends — for complex delivery a worker can delegate a whole task to a purpose-built coding agent (Claude Code, OpenAI Codex, Gemini CLI, Aider, or any headless CLI via a command template) instead of a single chat call. One stable
AgentBackendseam; new agents plug in by config (SOVEREIGN_AGENT_BACKEND/SOVEREIGN_BACKEND_<CATEGORY>), dry-run untilSOVEREIGN_AGENT_BACKEND_ENABLED=true. - MCP connectors in every worker — register MCP servers via
SOVEREIGN_MCP_SERVERSand their tools appear inline in every worker's tool-use loop, next to built-ins likeweb_fetch. Any MCP-provided connector (DBs, search, SaaS, another agent over MCP) becomes callable mid-task with no code change — the universal connector layer both Claude Agent SDK and Codex speak. See docs/BACKENDS.md. - Dynamic worker loading: drop a Python file into
sovereign_os/agents/user_workers/— no registration boilerplate.
Auditing
- Every task output is verified against Charter KPIs by the ReviewEngine.
- Category-tuned analytic rubric: high-value deliverables are scored criterion-by-criterion on a rubric matched to the work type — coding gets correctness/completeness/robustness/relevance, writing gets clarity/voice, etc. — always ending in a
safetycriterion. Criterion order is deterministically shuffled per task to blunt LLM-judge positional bias. - Value-aware bar: higher-paid jobs must clear a stricter passing score; cheap tasks can skip the (relatively expensive) LLM judge.
- AuditReport carries
score,passed,reason,suggested_fix,sub_scores, andproof_hash(SHA-256 of inputs + output). Each job's delivery view renders a quality scorecard — the per-criterion rubric breakdown (bars per criterion, category label, overall score) so you can see why a deliverable passed or failed, not just the verdict. - Append-only audit trail (JSONL). Integrity verifiable offline.
Monetization & job queue
- SQLite-backed job queue (Redis optional for multi-instance).
- Stripe integration: set
STRIPE_API_KEYand each completed job triggers a real charge; transactions appear in your Stripe Dashboard. - Auto-approval mode or manual review per job.
- Ingest from any HTTP endpoint, Reddit (PRAW), Shopify, WooCommerce, or custom scrapers via the ingest bridge.
- Webhook delivery:
POSTjob result to any URL on completion.
Left: job result delivered in the dashboard. Right: charges recorded directly in Stripe Dashboard.
Observability
- OpenTelemetry tracing.
- Prometheus metrics at
GET /metrics: job counters, queue depth, task duration histograms. GET /healthreturns config warnings (missing API keys, payment mode).- Structured JSON logs with correlation IDs.
Security
- API key authentication (constant-time comparison).
- Optional IP allowlist and per-IP rate limiting.
- Job input validation (Pydantic v2).
- No secrets in code; everything via environment variables.
TrustScore-gated permission system — agents earn capabilities through passing audits.
Project Layout
sovereign_os/
├── models/ # Charter schema (Pydantic v2)
├── ledger/ # UnifiedLedger — append-only USD + token accounting
├── governance/ # CEO (Strategist), CFO (Treasury), SpendCircuitBreaker, GovernanceEngine
├── agents/ # Workers, WorkerRegistry, SovereignAuth (+ JIT leases), categories, user_workers/
├── auditor/ # ReviewEngine, category rubric, AuditReport (proof_hash), KPIValidator
├── jobs/ # JobStore (SQLite), RedisJobStore
├── ingest/ # HTTP poller → job queue
├── ingest_bridge/ # Reddit, scrapers, Shopify/WooCommerce → jobs
├── oversight/ # Outbound task posting + escrow (RentAHuman, StacksTasker)
├── delivery/ # Deliver-back adapters (TaskBounty, StacksTasker, ClawTasks, Reddit, APB/x402)
├── connectors/ # web_fetch, code_workspace, figma, image_gen, submit_pr, sandbox
├── payments/ # StripePaymentService, X402PaymentService, DummyPaymentService
├── ui/ # Textual terminal Command Center (TaskTree · DecisionStream · Finance & Guardrails)
├── web/ # FastAPI app, dashboard, /api/jobs, /health, Stripe webhook
└── telemetry/ # OpenTelemetry, Prometheus
charters/ # Example Charter YAML files
examples/ # Example job payloads
tests/ # pytest
docs/ # Full documentation
Deployment
Local (bare metal)
git clone https://github.com/Justin0504/Sovereign-OS.git
cd Sovereign-OS
pip install -e ".[llm]"
cp .env.example .env
# Edit .env with your API keys
python -m sovereign_os.web.app
Listens on http://0.0.0.0:8000 by default. Set SOVEREIGN_HOST and SOVEREIGN_PORT to change.
Docker Compose (recommended for production)
cp .env.example .env
# Edit .env
docker compose up -d
# Web UI: http://localhost:8000
# Metrics: http://localhost:8000/metrics
# Health: http://localhost:8000/health
docker-compose.yml includes the web app, Redis (for multi-instance job queue), and named volumes for the ledger and job database.
# docker-compose.yml excerpt
services:
web:
build: .
ports: ["8000:8000"]
env_file: .env
volumes:
- sovereign_data:/app/data
redis:
image: redis:7-alpine
volumes:
- redis_data:/data
See docs/DEPLOY.md for volume strategy, health checks, and graceful shutdown.
Ingest bridge (optional)
Pull real orders from Reddit, scrapers, or Shopify:
pip install -e ".[bridge]"
python -m sovereign_os.ingest_bridge # serves on :9000
# In .env: SOVEREIGN_INGEST_URL=http://localhost:9000/jobs?take=true
See docs/INGEST_BRIDGE.md for Reddit credentials, scraper targets, and retail connectors.
Configuration
All configuration is via environment variables. Copy .env.example to .env.
| Variable | Required | Description |
|---|---|---|
OPENAI_API_KEY | One of these | OpenAI API key for LLM workers |
ANTHROPIC_API_KEY | One of these | Anthropic API key for LLM workers |
STRIPE_API_KEY | Optional | Stripe secret key (sk_test_… or sk_live_…). Without this, charges are simulated. |
SOVEREIGN_CHARTER_PATH | Optional | Path to Charter YAML. Default: charter.default.yaml. |
SOVEREIGN_JOB_DB | Optional | SQLite path. Default: data/jobs.db. |
SOVEREIGN_AUTO_APPROVE_JOBS | Optional | true to skip manual approval. Default: false. |
SOVEREIGN_JOB_WORKER_ENABLED | Optional | true to start the background job processor. Default: false. |
SOVEREIGN_INGEST_URL | Optional | HTTP endpoint to poll for incoming jobs. |
SOVEREIGN_INGEST_ENABLED | Optional | true to enable the ingest poller. Default: false. |
SOVEREIGN_API_KEY | Optional | Bearer token for /api/* endpoints. |
SOVEREIGN_ALLOWED_IPS | Optional | Comma-separated IP allowlist. |
SOVEREIGN_WEBHOOK_URL | Optional | URL to POST job results to on completion. |
REDIS_URL | Optional | Redis connection string for multi-instance job queue. |
Full reference: docs/CONFIG.md.
Custom Workers
Workers are the execution layer. Each worker maps to a skill name and receives a task description; it returns a TaskResult with output and token usage.
# sovereign_os/agents/user_workers/my_worker.py
from sovereign_os.agents.base import BaseWorker, TaskResult
class MyWorker(BaseWorker):
skill = "my_skill"
async def run(self, task_description: str, **kwargs) -> TaskResult:
result = await self._chat(f"Do this: {task_description}")
return TaskResult(output=result, success=True)
Drop the file into sovereign_os/agents/user_workers/ — it is loaded automatically on startup. Reference it in your Charter:
competencies:
- name: my_skill
description: "Does the custom thing."
See docs/WORKER.md for the full API, LLM helper methods, and advanced patterns.
Testing
pip install -e ".[dev]"
pytest tests/ -v
Tests (257, all passing) cover the governance engine, CFO circuit breaker, JIT capability leases, category-tuned audit rubric, reactive planning, ledger accounting, marketplace oversight, payments (Stripe + x402), job store, compliance hooks, web API, and webhook delivery. Run with LLM keys unset so tests use stubs (no network).
Docs
| Document | Description |
|---|---|
| QUICKSTART.md | Stripe + LLM key, first paid job, curl examples, troubleshooting. |
| CONFIG.md | All SOVEREIGN_* environment variables. |
| CHARTER.md | How to write mission, competencies, KPIs, and fiscal bounds. |
| WORKER.md | How to write and register custom workers. |
| DEPLOY.md | Docker Compose, volumes, health checks, graceful shutdown. |
| INGEST_BRIDGE.md | Pull orders from Reddit, scrapers, Shopify/WooCommerce. |
| MONETIZATION.md | Job queue, Stripe charging, approval flow, compliance. |
| POSITIONING.md | Order sources (open platforms, partners, API), service model. |
| CEO_CFO_PROFITABILITY.md | Unit economics: margin floor, rejecting unprofitable jobs. |
| ARCHITECTURE_FAQ.md | Design decisions and trade-offs. |
| AUDIT_PROOF.md | Verifiable audit trail, proof_hash, integrity check. |
| METRICS.md | Prometheus metrics (breaker, JIT leases, audit quality) + Grafana/PromQL starter. |
| BACKENDS.md | Pluggable agent backends (Claude Code / Codex / CLI agents), MCP compatibility, adapting to platform updates. |
| AUTONOMY.md | Human-out-of-loop runbook: profit screening, self-repair, platform notes, safety switches. |
| MULTI_INSTANCE.md | Redis queue, horizontal scaling, concurrency. |
| BACKUP.md | Back up job DB and ledger for disaster recovery. |
| FUTURE_DEVELOPMENT.md | Roadmap: stability, UX, workers, scale, community. |
| CONTRIBUTING.md | How to contribute, run tests, submit PRs. |
| SECURITY.md | Vulnerability disclosure policy. |
Contributing
Pull requests are welcome. Run the test suite before submitting:
pip install -e ".[dev]"
pytest tests/ -v
See CONTRIBUTING.md for the full process.
License
MIT — see LICENSE.
Sovereign-OS — Think. Approve. Execute. Audit. Every time.
Collected info
- ★ 92 stars
- ⎇ 13 forks
- Language: Python
- Source updated: 9/17/2026