maggy
What started as an opinionated Claude Code setup kit is now an autonomous AI engineering command center
Links
README
From the repo.
Claude Bootstrap + Maggy
Turn Claude Code into a self-reviewing, test-enforced engineering system that remembers context across sessions — then route work across 13 models from a single dashboard.
Claude Bootstrap is an installable config pack (skills, hooks, rules, templates) for Claude Code. Maggy is the optional local server that adds multi-model routing, a web dashboard, intent-driven protocols, and plugin orchestration. Both live in this repo. Start with Bootstrap; add Maggy when you need the harness.
1100+ tests. 73 skills. 15 MCP tools. Used daily across production codebases.
Who This Is For
- Solo engineers using Claude Code who want TDD enforcement, quality gates, and memory that survives context compaction — without changing their workflow
- Teams routing work across Claude, DeepSeek, Kimi, Gemini, and Codex from a single dashboard with cost-aware model selection
- Platform engineers building AI-assisted developer tooling who need a reference implementation with intent tracking, protocol execution, and plugin architecture
Choose Your Path
| Claude Bootstrap | Maggy Harness | |
|---|---|---|
| What it is | Skills, hooks, rules installed into ~/.claude/ | Local FastAPI server + web dashboard |
| Install time | ~30 seconds | ~5 minutes (Python 3.11+, API keys) |
| Requires | Claude Code (also works with Codex, Kimi, Gemini CLI) | Everything in Bootstrap + Python + optional Docker |
| You get | TDD enforcement, 73 skills, quality gates, ADR reviews, iCPG, Mnemos memory | All of Bootstrap + 13-tier routing, skill protocols, Telos testing, Cortex MCP, plugins, dashboard |
Bootstrap — 30-second install
git clone https://github.com/alinaqi/maggy.git
cd maggy && ./install.sh
Your next Claude Code session picks it up automatically.
Full Harness — zero-config
pipx install maggy-harness # or: pip install maggy-harness
maggy bootstrap # installs skills, hooks, ~/bin model wrappers, plugins
maggy serve # auto-configures from your local repos,
# then opens the dashboard at localhost:8080
(or from source: cd maggy && ./install.sh && maggy serve)
No API keys required to start — Maggy runs in local mode and, on first launch,
discovers your local git repos and opens the dashboard pointed at them. Add
GITHUB_TOKEN / ANTHROPIC_API_KEY later only if you want GitHub sync or
API-model features. See GETTING_STARTED.md for details.
What It Looks Like in Practice
Routing a task:
You: "review the auth middleware for timing attacks"
→ Blast score: 8/10 (security + architecture)
→ Routed to: Claude (Tier 11)
→ ADR gate: found docs/adr/0003-jwt-strategy.md → injected as context
→ Review runs with full architectural context
Skill Protocol execution:
You: "push to git"
→ Intent matched: git-push protocol
→ ✅ lint (2.1s)
→ ✅ typecheck (4.3s)
→ ✅ tests (11.2s)
→ ✅ stage
→ ✅ commit [AI-generated: "fix: resolve token refresh race condition"]
→ ✅ push
Fatigue-aware memory:
Session fatigue: 0.61 (PRE-SLEEP)
→ Mnemos: auto-checkpoint written
→ Micro-consolidation: 3 ResultNodes compressed
→ iCPG context injected: 2 ReasonNodes, 1 constraint
→ Context freed: ~18k tokens
The Problem This Solves
You're using Claude Code. It's impressive — but:
- It picks the most expensive model for everything, including trivial tasks
- Context fills up, state is lost, you re-explain yourself every session
- There's no enforcement: code quality, test coverage, and ADR compliance only happen if you remember to ask
- Running multiple agents on the same repo causes file conflicts
- You have no visibility into what Claude is actually doing inside your codebase
What Bootstrap Gives You
| Layer | What it does |
|---|---|
| 73 skills | Python, TypeScript, React, React Native, Flutter, Supabase, Firebase, Stripe, Playwright, visual validation, security & audit, ADRs, cross-agent delegation |
| TDD enforcement | Stop hooks — tests must pass before Claude considers a task done |
| Visual validation | Default for web projects — demo-video records a captioned Playwright walkthrough (proof mp4 that doubles as a passing E2E test); visual-validation screenshots catch regressions. A user-facing web flow isn't "done" without it |
| Quality gates | Max 20 lines/function, 3 params, 2 nesting levels. Enforced per file |
| iCPG | Intent-Augmented Code Property Graph. Stores why code exists. 6-dimension drift detection. Prevents duplicate implementations |
| Mnemos | Task-scoped memory with 4-dimension fatigue model. Survives context compaction with typed checkpoints |
| ADR enforcement | Non-trivial changes require an Architectural Decision Record. Missing one? Reverse-engineered from git history |
| Agent teams | 6 agents: Lead, Quality, Security, Review, Merger, Feature |
What Maggy Adds
| System | What it does |
|---|---|
| 13-Tier Routing | Semantic blast score (1–10) routes to cheapest capable model. Local Qwen3 classifier → DeepSeek (~80% of tasks) → Kimi → Gemini → Grok → Codex → Claude. Budget-capped with auto-demotion. Routing details |
| Skill Protocols | YAML-defined workflows in maggy/skills/protocols/. "Push to git" → lint → test → stage → commit → push. Drop a .yaml to add your own |
| Telos | Testing beyond TDD. Three planes: Conformance × Validation × Integrity. A zero in any plane collapses the total score. Details |
| Cortex MCP | Code intelligence: 10 edge types, cyclomatic complexity, FTS5 search, bidirectional traversal. 15 tools, single SQLite DB. Benchmarks |
| Polyphony | Docker-isolated parallel agent execution. Second session auto-provisions a workspace. Spec |
| Engram | Cross-session memory. 7 amnesia types. Persists architectural knowledge across weeks |
| Council PR Review | Multi-model council reviews a GitHub PR from the dashboard — deterministic mega-PR chunking, a static gate (tsc/ruff) as ground truth, and an adversarial refute pass that kills false positives. Extensible per-language skills (Python/TS/Go/Rust/Java/C#/Ruby/PHP + drop-in more). pip install maggy-harness[review] |
| Plugins | Drop-in system. Ships with: Build-in-Public (auto-posts to LinkedIn/X), Telos, GitHub/Asana/Monday providers |
Model Routing
Every message is scored 1–10 for complexity and classified by task type. The cheapest capable model wins.
| Tier | Model | Role |
|---|---|---|
| T0 | Qwen3 (local) | Classification, triage, free bulk ops |
| T1 | Gemini Flash-Lite | Bulk extraction, CIG pipelines |
| T2 | DeepSeek Flash | Docs, tests, scaffolding |
| T3 | Gemini Flash | Multimodal, vision, audio |
| T4 | DeepSeek Pro | Complex coding, multi-file refactors |
| T5 | Gemini CLI | Multi-file agentic coding |
| T6 | AGY | End-to-end implementation (git + code + test) |
| T7 | Kimi | Long-context analysis, routing alt |
| T8 | Gemini Pro Search | Deep research, Google grounding, 2M context |
| T9 | Grok | Competitor intel, deep reasoning |
| T10 | Codex | Bulk generation, security-sensitive tasks |
| T11 | Claude Sonnet | Quality-critical code, complex debugging |
| T12 | Claude Opus | Architecture, security review, ADR decisions |
Routing is semantic (Qwen3 as local classifier), fatigue-aware, budget-capped, and cascading.
Gateway routing with srooter — www.srooter.ai
We've added first-class support for srooter, an Anthropic/OpenAI-compatible LLM gateway that routes your requests across models (Claude, MiniMax, DeepSeek, Kimi, Gemini, Grok, local Qwen) transparently — intent-based routing, budget caps, fallbacks, and a usage dashboard, without changing your tools.
Recommended with Maggy, Claude Code, or Codex. Point any of them at the gateway and your traffic is routed for you — no per-tool config:
# Claude Code (or Codex) → srooter
export ANTHROPIC_BASE_URL="https://www.srooter.ai/anthropic" # or your local gateway
export ANTHROPIC_API_KEY="<your-srooter-key>"
claude # now routed through srooter
Pick the model you "follow" once with /model-config — Maggy, the route-task hooks, and srooter all honor the same choice. Trivial asks stay on the cheap/local tier; real coding goes to your primary model (e.g. MiniMax-M2.5).
Context shunt — cheap reads, small context
Gateway routing picks the model for a turn. The context shunt trims what a
single tool call pulls in when the turn is legitimately on your main model: a
PreToolUse hook (context-shunt-gate) catches reads of large files — code or
logs/generated output — and steers them to bulk-read, which hands the files to
a cheap worker (default deepseek --flash) and returns a compact summary. The
raw bytes never enter context. For code symbols it points at the graph
(get_code_snippet) instead. Inspired by Spotify's "shunt" plugin.
Fully configurable in ~/.claude/shunt.conf (or env): SHUNT=on|off,
SHUNT_MIN_LINES (default 350), SHUNT_MODE=suggest|block|off (default
suggest — nudges, never blocks), SHUNT_MODEL. See the context-shunt skill.
bulk-read "how does token refresh work?" src/auth/session.ts src/auth/refresh.ts
Parallel Development (Polyphony)
Run several agents at once — each in its own Docker/OrbStack container with a full git clone on its own branch, so concurrent work never collides on files or branches.
- Auto-isolation — a second Claude Code session in the same project automatically provisions its own workspace (via the
polyphony-auto-isolatehook). No setup. /spawn-team— spawns a coordinated TDD agent team; container-isolated by default when Docker + thepolyphonyCLI are present, with a graceful fallback to native parallel agents.
polyphony init # one-time: create ~/.polyphony/ config
polyphony spawn "add auth" # create + route a task to an agent
polyphony status # running agents / task states
polyphony cleanup # remove completed workspaces
From Claude Code: /polyphony-init, /polyphony-spawn, /polyphony-status. Requires Docker or OrbStack. Full design: Polyphony spec.
Telos: Testing Beyond TDD
Standard TDD tells you if your code passes tests. Telos tells you if your code fulfills its intent.
IFS (Intent Fidelity Scale) = F1 × F2 × F3
F1 — Conformance: passed / total tests (pytest / vitest)
F2 — Validation: drift severity (Cortex drift_events)
F3 — Integrity: IF-3 orphan symbols (no reason edges)
IF-4 empty contracts (no pre/post/invariants)
IF-6 stale reasons (proposed >7d, never fulfilled)
IF-7 scope sprawl (reason scopes >10 files)
A zero in any plane collapses IFS to zero. 100% test pass rate with severe architectural drift = score of 0. This is intentional. See the Telos RFC.
Repo Structure
.claude/
skills/ # 73 skills — Python, TS, React, security, mobile, databases
hooks/ # TDD enforcement, quality gates, Mnemos lifecycle
rules/ # Conditional rules by file glob
templates/ # settings.json, CLAUDE.md, ADR template, PR template
maggy/
maggy/
pipeline/ # Unified ChatPipeline orchestrator
skills/ # Skill injection + YAML protocol engine
api/ # REST API (chat, routing, plugins, pipeline logs)
static/ # Web dashboard (vanilla JS, no build step)
services/ # Routing, memory, execution, Mnemos
cortex-mcp/ # Code intelligence MCP server
src/cortex/
structure/ # AST extraction, edge types, complexity
storage/ # SQLite graph store, FTS5 index
plugins/ # Drop-in plugins (build-in-public, telos, providers)
Tests
cd maggy && python3 -m pytest tests/ -x -q # 900+ tests
cd cortex-mcp && python3 -m pytest tests/ -q # 207 tests
What's New in v6.65
- Security audit — a default part of the harness —
skills/security-audit/runs a structured, adversarial, multi-phase audit (recon → coverage-led hunting → finder≠validator validation → machine-readablefindings.json→ target-neutral report). It's copied into every project at init (like the preventivesecurityskill), reusescouncil-review(adversarial validation),cpg-analysis(static taint), andagent-teams/polyphony(isolated parallel hunters), and ships a JSON schema + a zero-dependency integrity validator.
What's New in v6.64
- Visual validation is a default for web projects — the
demo-videoskill (captioned Playwright walkthrough → proof mp4 that doubles as a passing E2E test) now ships with the harness and is copied into every web project (React, Full Stack, PWA) at init. A user-facing web flow isn't "done" without it — it's part of the Definition of Done inbase, alongsidevisual-validation(screenshot-regression) andplaywright-testing(behavior).
What's New in v6.63
- DataForSEO skill —
skills/dataforseo/for keyword/SERP research (search volume, competition, CPC) to ground naming/SEO decisions in real data. Env-only auth.
What's New in v6.62
- Codex dual auth — Codex now works via your ChatGPT subscription (
codex login, forcodex execdelegation) or a real OpenAI API key (to run Codex as a Claude Code model through srooter).codex-statusshows what's detected;set-codex-auth auto|subscription|api_keypins it. (Subscription can't back a Claude Code model — OpenAI has no Anthropic endpoint — so that path is delegation-only.) - Direct-provider launchers —
/model-config deepseek --direct(orglm/kimi) writes~/bin/claude-deepseek/claude-glm/claude-kimi, each pointing Claude Code straight at the provider's native Anthropic endpoint — no srooter hop. Runclaude-deepseekinstead ofclaudeand that session runs directly on DeepSeek Pro. Plainclaudeis untouched, so you pick per terminal. (Codex isn't direct-capable — OpenAI has no Anthropic API — so it stays routed through srooter.) - Switch Claude Code's backend from inside Claude Code —
/model-config deepseek(orkimi,glm,codex) moves your real coding work onto DeepSeek Pro, Kimi K3, GLM 5.3, or Codex, routed through srooter. Restart srooter, start a fresh session, and coding runs on the chosen model while trivial asks stay on the fast local classifier. - Both coding routes follow the switch —
applynow rewriteslong_contextandsubstantiveinsrooter.yaml, so substantive traffic follows the chosen backend (not just long-context).
See CHANGELOG.md for full history.
Docs
| Getting Started | Installation, prerequisites, first session walkthrough |
| Architecture v5 | System design, routing, dashboard |
| CLI Reference | REPL commands, slash commands, routing |
| Telos RFC | Intent-grounded testing spec |
| Cortex docs | Code intelligence, edge types, MCP tools |
| Cortex benchmarks | Performance vs codebase-memory-mcp |
| Changelog | Version history (current: v6.66.0) |
Contributing
Skill PRs welcome. All skills run through the linter before merge:
PYTHONPATH=scripts python3 -m skill_lint --fail-on error skills/your-skill/
See CONTRIBUTING.md for the quality gate checklist.
License
MIT — See LICENSE
Need help scaling AI engineering in your org? LeanAI Ventures — Claude Code & MCP specialists
Collected info
- ★ 708 stars
- ⎇ 56 forks
- Language: Python
- Source updated: 9/19/2026
Config for your environment
Replace {MCP_ENDPOINT_URL} with this MCP’s endpoint URL (from its repo or docs above). No API key — you connect directly.
Tool
OS
Config file: ~/.cursor/mcp.json
{
"mcpServers": {
"mcp-server": {
"url": "{MCP_ENDPOINT_URL}"
}
}
}Paste into mcpServers in the config file. Restart Cursor after saving.
If this MCP is also published on mcpchannel.ai, you can subscribe from Browse and use the gateway config there instead.