← Discover MCPs and Agents
g
MCPAI & MLGitHub

guaardvark

The self-hosted AI studio: local video, image, music, voice, LoRA training, coding swarms, RAG and screen agents on one GPU, driven from the Studio or by your coding agent (Claude Code, Cursor, Codex, OpenClaw) through MCP and skills.

Links

README

From the repo.

Guaardvark AI System — www.guaardvark.com

Guaardvark

License: MIT CI PyPI Glama GitHub stars GitHub issues Sponsor

The self-hosted AI studio. Coding agents and 20-agent swarms in isolated git worktrees, screen agents with their own real desktop, self-tuning RAG, continuous voice chat — and a full media pipeline: video, image, full-song music, neural voice. One install, one GPU, everything on your machine. Your machine. Your data. Your rules.

Works with your coding agent. Claude Code, Cursor, Codex, OpenClaw and Gemini CLI drive every flow above through the built-in MCP server and fifteen agent skills: "make a music video from this song", "film this script", "train a LoRA of this character", "swarm this refactor" — the agent queues the job on your GPU and polls it to the finished file.

Runs on Linux with one NVIDIA card (16 GB for video). Apple Silicon is supported with GPU features arriving through Metal (what works today is in INSTALL.md); Windows through WSL2 is being verified.

Install (one command, then open the Studio):

curl -fsSL https://guaardvark.com/install.sh | bash

Add it to Claude Code (two lines, no clone):

/plugin marketplace add guaardvark/guaardvark
/plugin install guaardvark@guaardvark

See the VERSION file for the current release · guaardvark.com · Quick Start for manual install options.

For the exhaustive feature list, models, surfaces, and plugin details, see CAPABILITIES.md. This README focuses on the marquee experience, quick start, and what makes Guaardvark different.

What's in the box

See it
Media studio11 local video models across five families (Wan 2.2, CogVideoX, LTX, HunyuanVideo, MiniMax H3 with native audio), image generation, full-song music, neural voice with consent-gated cloning, 4K/8K upscalingEps 5 · 6 · 7
DirectorsA beat-synced music-video director, a 5-role Film Crew, an auto-editing video editor — and the walkthrough director that produced this README's own video seriesEps 8 · 9
Coding agent & code intelligenceMonaco editor, AST-aware analysis and dependency graphs, System Mapper: a live constellation of the whole codebaseEp 14
Agent swarmsUp to 20 parallel coding agents in isolated git worktrees with dependency-aware merging; fully-local backend via OllamaCAPABILITIES.md
Screen agentsA real Ubuntu/XFCE desktop of their own, vision + closed-loop servo clicking, live VNC viewer on any pageEp 4
KnowledgeHybrid RAG on pgvector with cross-encoder reranking, layout-aware document parsing with page-level citations, retrieval that shows its chunks and scores, Autoresearch that tunes retrieval overnightEp 3
Voice & channelsContinuous voice chat, a three-tier chat brain, Discord bot, supervised outreach, MCP in both directions, a 25-module CLIEp 2
Self-running platformSelf-improvement behind guardian review and kill switches, rules engine, jobs & scheduling, schema-aware backups, GPU orchestrator, multi-machine InterconnectorEps 11 · 12

The aardvark (/ˈɑːrd.vɑːrk/; Orycteropus afer) is a medium-sized, burrowing, nocturnal mammal native to Africa. The aardvark is the only living member of the genus Orycteropus, the family Orycteropodidae and the order Tubulidentata. It is found over much of the southern two-thirds of the African continent, avoiding areas that are mainly rocky. A nocturnal feeder, the aardvark subsists on ants and termites (myrmecophagy) by using its sharp claws and powerful legs to dig the insects out of their hills, and its long snout to sniff out food. It digs a burrow in which to live and rear its young.

— Wikipedia, CC BY-SA

The Guaardvark (/ˈɡwɑːrd.vɑːrk/; Workstationus selfhosticus) is a burrowing, nocturnal AI system native to consumer hardware. The Guaardvark is the only living member of the repository github.com/guaardvark/guaardvark, the family LocalAI, and the order AutonomousAgents. It is found across most of the modern desktop, avoiding regions that are mainly cloud. A nocturnal feeder, the Guaardvark subsists on prompts and unstructured data (promptophagy) by using its sixty-odd tools and a swarm of parallel coding agents to dig bugs out of their codebases, and a retrieval index to sniff out knowledge in the dark. It is fiercely territorial about its single GPU, admitting one process to the card at a time and evicting any language model found loitering there. It digs isolated git worktrees in which to work, and rears its images, video, music, and cloned voices entirely on your own machine.

▶ The Walkthrough Series — every feature, on camera

Short, unscripted-feeling screen recordings of the real system doing real work — narrated by a voice the system cloned itself (that's Episode 7). Twelve episodes are live: the first series covers the twelve subsystems, and a second series picks up what shipped since.

One Chat Box, Three Different Brains — Guaardvark Ep 2
Ep 2 — Chat Brain: one chat box, three speeds
Your Files Get a Desktop — Guaardvark Ep 3
Ep 3 — File Desktop: your files get a desktop, plus RAG that shows its work
The Agent Behind the Glass — Guaardvark Ep 4
Ep 4 — Screen Agent: its own desktop, eyes, and hands
Local Image Gen — Guaardvark Ep 5
Ep 5 — Image Gen: one prompt, a whole story
Local Video Generation — Guaardvark Ep 6
Ep 6 — Video Gen: seven models, one GPU
Local Voice Cloning — Guaardvark Ep 7
Ep 7 — Voice Clone: consent-gated, self-checking
Local AI Music Video — Guaardvark Ep 8
Ep 8 — Music Video: drop a song, get a film
Local AI Film Crew — Guaardvark Ep 9
Ep 9 — Film Crew: script, cast, storyboard, cut
The System That Fixes Itself — Guaardvark Ep 11
Ep 11 — Self-Repair: it fixes its own code, behind a gate you control
Command Center — Guaardvark Ep 12
Ep 12 — Command Center: see everything, gate everything, kill everything
The New Front Door — Guaardvark Ep 13
Ep 13 — The New Front Door: the Workspaces bar, and everything new since the first series
A Map of Everything — Guaardvark Ep 14
Ep 14 — System Map: every module, drawn from its real imports; findings that carry their own fix

▶ Watch the full playlist — Episode 1 (the full tour) and Episode 10 (the video editor) are on the way.

More demos

The beat-synced music video — watch on YouTube. One style prompt and a short narrative, then go: Guaardvark wrote every shot prompt, generated the storyboards, rendered the clips, and assembled the cuts to the beat it detected in the song. (Full disclosure: the glitch effect was the one manual touch, added in Shotcut; the song was made in Suno — Guaardvark's own music generation is being wired into this pipeline.)

Full visual gallery (dashboard, video generator, swarm planner, agents, plugins, media library, etc.) is available on guaardvark.com.


Marquee Capabilities

Generation & Editing (all local, no cloud APIs)

  • Text-to-Video / Image-to-Video (Wan 2.2 5B + 14B MoE, CogVideoX-5B, LTX-2.3, LTX-2.5) with an in-process batch queue, quality tiers, frame interpolation, prompt enhancement, and one-click jump to ComfyUI for custom workflows.
  • Audio Studio (Audio Foundry plugin): ACE-Step 3.5B music (vocals or instrumental, Suno-style chips + LLM polish), Stable Audio Open FX/ambience, Chatterbox + Kokoro neural TTS, Piper fallback, consent-gated voice cloning.
  • Image gen (Stable Diffusion + batch + face/anatomy controls) + powerful 4K/8K GPU upscaling (Real-ESRGAN family, HAT-L, NMKD, Foolhardy, two-pass, video frame-by-frame).
  • Built-in Video Editor (Shotcut-lite 3-lane timeline: video/text/audio, real ffmpeg drawtext overlays, drag-and-drop from media library, visual trims, undo, keyboard shortcuts).

Agents, Automation & Swarms

  • AgentBrain three-tier router (Reflex <100 ms pattern match, Instinct single-shot, Deliberation full ReACT).
  • Real-desktop screen agents (Xvfb + XFCE :99, Gemma4 vision + closed-loop servo, 45+ deterministic recipes, live per-iteration reasoning stream in chat, draggable VNC viewer everywhere).
  • Swarm: parallel agents in isolated git worktrees (Claude Code or fully local Cline/OpenClaw via Ollama), Flight Mode (offline), dependency-ordered merge, cost tracking, up to 20 concurrent.
  • Film Crew: 5 specialized agents that turn a logline into a finished video (script → casting with LoRAs → shots → keyframes → edit).
  • Self-improvement engine (test → agent fix → verify → broadcast) with guardian review and kill switches.
  • Supervised social outreach (Reddit fully working; others drafting+review ready) with persona, grading, cadence, audit, and global kill switch.
  • MCP server + client integration (Claude Desktop, Cursor, etc.).

Knowledge, Code & Workflow

  • Strong RAG (hybrid keyword + vector search on pgvector, cross-encoder reranking, AST code chunking, layout-aware PDF/DOCX parsing, entity extraction, RAG Autoresearch, per-project isolation).
  • Monaco code editor + Code Analyzer + per-repo indexing + dependency graphs + System Mapper (constellation view of the whole codebase).
  • Full desktop-grade file/project/client/website/notes/media management with cross-links and recursive indexing.
  • Task scheduler (Celery beat), Rules & Prompts (portable bundles), Interconnector for multi-machine clusters (master/client, approval gates, learning broadcast).
  • 10+ managed plugins with health checks, port orphan cleanup, and a real GPU Memory Orchestrator.

Platform & Ops

  • Everything stays on your machine by default. Flight Mode is real and end-to-end tested.
  • Plugin system + resource orchestrator so big models don't fight for VRAM.
  • Backup/restore (granular or full, schema-migration aware), advanced settings surfaced in UI, live GPU/CPU monitoring.
  • CLI (llx / PyPI guaardvark), browser UI, and MCP.

See CAPABILITIES.md for the complete enumerated list (models, exact tool counts, plugin manifests, page surfaces, etc.).


Why local?

Cloud platformsGuaardvark
Where your data livesTheir serversYour machine. Period.
Per-token / per-minute feesAlways on the meterFree. Generate all night if you want.
Content policyTheir rulesYour rules.
Custom models / LoRAsWhatever they exposeAny GGUF, any LoRA, any embedding model
Works offlineNoYes. Flight Mode tested end-to-end.
Agents drive a real desktopSandboxed browsersReal Ubuntu/XFCE on your hardware
Swarms of parallel agentsPer-task billing scales nastily20 agents in parallel; only cost is power
Multi-machine clusters"Talk to sales"Built-in. Master/client, approval gates
Lock-inMigrate at your own riskIt's your computer. Move it whenever.

How Guaardvark compares

The local-AI ecosystem has excellent tools for every slice: chat UIs, RAG second-brains, node-graph media pipelines, coding agents, assistant gateways. Guaardvark's bet is different — one install on one GPU that is the whole studio, with a single GPU orchestrator arbitrating all of it.

CapabilityChat UIsRAG appsNode graphsCoding agentsAssistant gatewaysGuaardvark
Local chat + RAGcorecore——via toolscore (Ep 3)
Media production (video · image · music · voice)——image/video graphs—via connected toolscore, with director engines (Eps 5–9)
Agents on a real desktop————browser/tool usecore (Ep 4)
Parallel coding swarms———usually one agent—up to 20 in git worktrees
Self-improvement behind human gates—————core (Ep 11)
One-GPU resource arbitration—————core (Ep 12)
Integration / plugin ecosystem breadthvariesvariesenormousgrowingenormoussmaller — 10 first-party plugins, plus MCP both ways
Hosted / mobile optionoftenoftenoftenoftenoftennone, by design — it's your machine

Columns describe the typical shape of each category, not any single project — several projects exceed their category in places. The Guaardvark column links to walkthrough episodes where you can watch the claim happen.

If all you need is one slice, use the excellent specialist: a chat UI like Open WebUI, a node graph like ComfyUI (Guaardvark hands off to it with one click), a RAG workspace like AnythingLLM. Guaardvark is for when you want the whole studio on one box.

Agent-driven media production, side by side

A newer category: the coding agent runs the studio. Facts checked 2026-09-11 from each project's repository; stars move, the shape does not.

OpenMontageNomiMaestroComfy MCPPromptus / LocalForge / SimpliGenGuaardvark
What it is12 video pipelines driven by Claude Code, Cursor, CodexDesktop video workbench with 25 MCP toolsLocal video, image, music, voice with a director modeOfficial MCP for ComfyUIOne-click local image + video apps, $30–$97 one-timeThe whole studio, driven by your agent or the Studio UI
Generation runsmostly cloud APIs, local models optionalyour ComfyUI or cloud providerslocal (a Wan2GP fork)your ComfyUI, or Comfy Cloud on subscriptionlocallocal
Agent driving ityes (skills + CLI)yes (MCP)no (in-app planner)yes (generation only)noyes (MCP + skills) and the built-in agent brain
Music, voice, voice clonevia cloud TTS/Suno—music + voiceaudio nodes—ACE-Step songs, three TTS engines, consent-gated clone
Film crew, music-video directorpipelines, storyboard boardstoryboard + timelinedirector mode——5-role Film Crew, beat-synced director, video editor
Coding swarm, screen agents, outreach, RAG—————core
LoRA training, upscaling, add any HF model by URL——LoRA browser—model managercore, in the Studio
OSmac, Linux, Windowsmac, WindowsNVIDIA via PinokioanyWindows, mac (LocalForge: Linux too)Linux; Apple Silicon partial (Metal); WSL2 in verification
LicenseAGPL-3.0AGPL-3.0WanGP non-commercialopen sourceproprietaryMIT
Repository stars, 2026-09-1157k0.5k0.5k(part of ComfyUI, 133k)—0.2k

Every cell is a claim you can check in the linked repositories; corrections welcome in an issue.


What Makes This Different

Security, Privacy & Local Guarantees

  • Everything runs locally by default. No telemetry or cloud phoning home unless you explicitly enable the Interconnector (master/client with approval workflows).
  • Flight Mode — fully offline operation with automatic network detection and local-model fallback. Swarm and agent tasks have been validated end-to-end without internet.
  • Self-improvement safety — three modes (Scheduled / Reactive / Directed). Every proposed code change can be reviewed by "Uncle Claude" (Anthropic API guardian) before application. Codebase lock toggle + Pending Fixes queue for human staging/approval. Fixes can be broadcast to connected family members.
  • MCP server uses a strong default-deny policy (backend/mcp/config.py): desktop control, agent execution, system/shell, browser automation, and test execution tools are hidden by default. Only safer tools + read-only guaardvark://outputs/ resources are exposed unless you explicitly allowlist.
  • Outreach is supervised by default (drafts queue; nothing posts without explicit Approve). Kill switch, per-platform cadence limits, full JSONL audit trail, and persona enforcement.
  • Voice cloning requires an explicit consent prompt. Reference clips stay under your control.
  • WordPress connectivity ships with security disclaimers and is treated as opt-in/beta until a final hardening pass.
  • Your data, models, LoRAs, and generated media never leave the machine unless you choose to push them.

AgentBrain — Three-Tier Neural Routing

Every message is routed through a three-tier decision engine that picks the fastest path to the right answer. Reflexes fire in under a millisecond. Instinct handles single-shot requests in one LLM call. Deliberation spins up a full ReACT reasoning loop when the problem demands it.

TierNameLatencyLLM CallsWhen It Fires
1Reflex<100ms0Greetings, farewells, media controls — pattern-matched, no inference
2Instinct1–3s1Single-shot questions, web searches, image generation, vision tasks
3Deliberation5–30s3–10Multi-step research, analysis chains, complex agent tasks
  • Automatic escalation — Tier 2 can signal complexity and hand off to Tier 3 mid-response.
  • Agent-screen gating — vision/desktop tools are only in scope when the virtual screen is active.
  • BrainState singleton + warm-up thread for zero-overhead routing and fast first-token times.

Autonomous Screen Agents

Guaardvark agents control a real Ubuntu desktop (Xvfb + XFCE at 1000×1000) — exactly what the model would see if you VNC'd into the box from another machine. Same Applications menu, same desktop icons, same taskbar. Agents see the screen through vision models, move the mouse, click buttons, type text, navigate browsers, and verify their own actions.

  • Real XFCE session — not a custom widget panel. xfce4-session runs on the virtual display via a scrubbed environment, with isolated XDG_DESKTOP_DIR and XDG_CONFIG_HOME so the agent's desktop, file manager, and configs never collide with the user's. Vision models recognize the layout instantly because it's standard Ubuntu.
  • Unified vision brain — Gemma4 sees the screen, decides the next action, and emits click coordinates (native box_2d) in a single inference call. Per-model scale factors are tracked and updated by the self-improvement loop.
  • Closed-loop servo targeting — three-attempt adaptive strategy: ballistic move → single correction with crosshair overlay → full corrections with zoom-cropped analysis around the cursor
  • Live per-iteration reasoning stream — every Think step (action, target, full reasoning, pivots when the loop gets stuck) streams into chat in real-time. No more 30-second blackouts followed by a single "completed" line. The trail persists in history so you can audit any run.
  • 45+ deterministic recipes — browser navigation, tabs, scroll, search, find, zoom, copy/paste — all execute instantly from a JSON recipe library, bypassing the vision loop entirely. Recipes carry optional preconditions (visibility checks) so they're skipped cleanly when their UI isn't on screen.
  • Obstacle detection — handles popups, permission dialogs, and notification bars with automatic thinking model escalation
  • Self-QA sweep — agent navigates every page of its own UI and reports what's working and what's broken
  • Live agent monitor — real-time SEE/THINK/ACT transcript of every decision the agent makes
  • Integrated screen viewer — draggable, resizable VNC viewer on any page with popup window mode

Supported Vision Models

ModelRoleCoordinate SystemNotes
Gemma4 (e4b)Sees + decides + clicksbox_2d normalized to 1000, [y1,x1,y2,x2]Unified brain — vision, reasoning, and coordinates in one call
MoondreamFallback eyes1024px internal widthFor text-only chat models (llama3, ministral-3) that need external vision

Swarm Orchestrator — Parallel Agent Execution

Launch multiple AI coding agents in parallel, each working in an isolated git worktree on its own branch. Results merge back with dependency-ordered conflict detection, optional test validation, and full cost tracking.

  • Two backends — Claude Code (cloud, cost-tracked at $0.015/$0.075 per 1K tokens) and Cline/OpenClaw (fully local via Ollama, zero cost)
  • Flight Mode — fully offline operation. Auto-detects network state, falls back to local models, serializes file conflicts automatically. No prompts, no internet required.
  • Git worktree isolation — each task gets its own branch and working directory. All worktrees share the .git directory (lightweight). Automatically excluded from git status.
  • Dependency-aware merging — topological sort ensures foundational changes land first. Dry-run conflict detection before real merge. Test suite validation before integration.
  • Built-in templates — REST API scaffold, refactor-and-extract, test coverage expansion, Flight Mode demo
  • Up to 20 concurrent agents — configurable limit with automatic slot management
  • Live dashboard — real-time status, per-task logs, cost breakdown, elapsed time, disk usage

Self-Improving AI

The system runs its own test suite, identifies failures, dispatches an AI agent to read the code and fix the bugs, verifies the fix, and broadcasts the learning to other instances. No human in the loop.

  • Three modes — Scheduled (every 6 hours), Reactive (triggered by repeated 500 errors), Directed (manual tasks)
  • Guardian review — Uncle Claude (Anthropic API) reviews code changes for safety before applying, with risk levels and halt directives
  • Verification loop — re-runs tests after every fix to confirm it worked
  • Pending fixes queue — stage, review, approve, or reject proposed changes
  • Cross-machine learning — fixes propagate to all connected instances via the Interconnector

RAG That Actually Works

Chat grounded in your documents. Upload files, build a knowledge base, and ask questions. The AI reads and understands your content — not just keyword matching.

  • Hybrid retrieval — Postgres full-text keyword + vector semantic search, fused with a per-query weighting that leans keyword-ward for identifier-like queries and semantic-ward for prose
  • Cross-encoder reranking — a reranker reads the query and passage together and reorders the candidates, which a bi-encoder cannot do; it is admitted against free VRAM and falls back to CPU rather than competing with image or video generation
  • Smart chunking — code files get AST-informed chunking (a function stays one chunk instead of being split mid-body), prose gets semantic splitting
  • Layout-aware document parsing — PDFs, DOCX and PPTX are parsed for reading order, section headers and page positions, so a retrieved passage can cite the page it came from. (Scanned documents need OCR, which is not installed by default — they report that rather than indexing as empty.)
  • Grounded citations — every chunk carries its source file, section breadcrumb and page, and that context is written into the embedded text as well, so a passage lifted out of the middle of a document still says where it came from
  • Postgres-backed vector store — embeddings live in pgvector alongside the rest of your data, with an ANN index and a persisted full-text index; they are covered by the same database backups
  • Index profiles — the same documents can be projected more than one way (fewer, larger passages for a small local model; finer-grained ones for an external client), switched in Settings
  • Corpus-level summaries — a recursive summarisation pass answers "what are the themes across all of this", which passage search structurally cannot
  • Multiple embedding models — switch between lightweight (300M) and high-quality (4B+) via UI
  • Entity extraction — automatic entity and relationship indexing
  • Per-project isolation — each project has its own knowledge base and chat context
  • Retrieval that shows its work — live retrieval tests display the actual chunks and scores behind an answer (Episode 3), instead of just asserting one. Every query can also return a trace of which retrieval legs actually ran, so a degraded answer is distinguishable from a bad one

Autoresearch — retrieval that tunes itself. An autonomous optimization loop runs overnight experiments on your corpus: it proposes changes to chunking and retrieval parameters, evaluates them with an LLM-as-judge harness, keeps wins, and reverts regressions — bounded by a wall-clock budget, a run ledger, and a circuit breaker (Episode 11). Your retrieval gets better while you sleep, and the morning report says exactly what changed and why.

Even the file manager is better

A client's Linux desktop player refused to play a video — the distro was missing the right codec plugin. They dropped the same file into Guaardvark's file desktop and it just played. No codec pack involved: the backend range-streams files inline (download_document in backend/api/files_api.py) into the browser's own decoders, which ship with H.264/VP9 support regardless of what the desktop has installed. The same mechanism means PDFs start painting before they finish downloading and video seeking works from the first byte.

Code Intelligence & the System Mapper

  • Monaco code editor with multi-file tabs and an AI assistant pane.
  • AST-aware code intelligence — repository maps, dependency graphs, and structure-aware code search (get_repository_map, read_ast_node, search_code — also exposed over MCP).
  • System Mapper — a live, force-directed constellation of the entire codebase computed from real imports (1,300+ modules on camera in Episode 14), with lifecycle tagging (active / dormant / auto-loaded / test / script / config), ranked findings you can dispatch to the self-improvement agent, and findings whose remedy is mechanical staged as an exact proposal for your review.
  • Guarded self-coding — every AI code write funnels through a single verified exact-replacement gate, behind the codebase lock.

Model Context Protocol (MCP)

Guaardvark speaks MCP both ways — exposes its tools to any MCP client (Claude Code, Cursor, Grok, Claude Desktop, Zed, Gemini, etc.) and can call tools from connected external MCP servers.

  • One-command setup — python -m backend.mcp install detects the agent clients on your machine and writes the guaardvark server entry into their configs (existing files are backed up, other entries untouched). python -m backend.mcp doctor diagnoses a broken setup: server self-test, a real stdio handshake, and a scan of client configs for stale paths.
  • Claude Code plugin — two lines, no clone: /plugin marketplace add guaardvark/guaardvark then /plugin install guaardvark@guaardvark. It asks for the path of your Guaardvark checkout, wires the MCP server from there, and loads every skill below as /guaardvark:<skill>.
  • Agent skills — .agents/skills/ ships one skill per flow (images, video, music video, Film Crew, voice, music, upscaling, Cast/LoRA training, Hugging Face model onboarding, swarm, knowledge, code, outreach, ops) in the Agent Skills format, so Claude Code, Cursor, Codex and OpenClaw know which tool or route to call for each job. python -m backend.mcp install --skills links them into ~/.claude/skills; other agents read .agents/skills/ from the checkout. Start with setup.
  • As a server — python -m backend.mcp (stdio, the default) or python -m backend.mcp http (streamable HTTP on 127.0.0.1:8788/mcp; loopback-only by default since there is no auth yet). Strong default-deny policy (see backend/mcp/config.py): categories such as desktop, agent_control, system, browser, test_execution, and mcp meta-tools are denied by default. Dozens of safer tools (chat, RAG, files, generation, memory, etc.) plus read-only guaardvark://outputs/ resources are exposed — python -m backend.mcp list-tools prints the live list. Generation tools queue by default over MCP and hand back a batch id (get_generation_status reads it); every call runs on a worker thread under an enforced timeout (GUAARDVARK_MCP_TIMEOUT, 120 s; 30 min when a caller asks to wait for a render). Verified end-to-end by an initialize/tools-list handshake in the smoke tests.
  • As a client — mcp_connect / mcp_execute + live tool inventory so the chat LLM can discover and use tools from other MCP servers by name.
  • Audit logging, timeouts, and circuit breakers are built in.

Outreach System — Supervised AI for Social-Media Engagement

A supervised, auditable framework for drafting and posting authentic comments on Reddit, Discord, Twitter/X, and Facebook — using your own indexed knowledge as the source of truth for citations and context. The point isn't volume. It's keeping up with engagement on your own products and topics, with the agent handling the legwork.

How it works:

  1. Discover — the agent scouts target threads either by URL (you paste one into the New Draft modal) or by walking platform-specific entry points (subscribed subreddits, Discord channels, Twitter feeds, Facebook groups).
  2. Context — for each candidate post, the agent fetches the OP body and top comments. Reddit goes through the JSON API (fast, no scrape). Discord, Twitter, and Facebook go through the agent's logged-in Firefox session over CDP/BiDi, with a vision-model fallback when DOM selectors drift after a platform redesign.
  3. Draft — your local LLM composes a reply grounded in the thread context plus citations from your indexed documents (clients, projects, products, examples — whatever you've fed the knowledge base).
  4. Grade — every draft is scored against a relevance + quality rubric. Anything below threshold is dropped before it reaches the queue. Generic "great post!" replies don't survive grading.
  5. Review — drafts land in a queue. In supervised mode (the default), nothing posts without your approval. Edit, save, approve, reject — your call on each one.
  6. Post — approved drafts post via the logged-in browser session (Reddit/YouTube servo) or Discord API, cadence-gated. Natural language from chat (/outreach …) or llx outreach "…" runs recon+draft; posting still needs approve while supervised. Twitter/Facebook drafting works; auto-post for those platforms is not wired.

Three layers of safety:

  • Kill switch at the system level. Flip it off and every outreach pipeline — drafting, queueing, posting — stops mid-flight. Nothing escapes.
  • Supervised mode is the default. Drafts queue, never auto-post. You approve each one explicitly.
  • Cadence gates — at most 1 post per 30 minutes per platform, configurable. Prevents bot-shaped behavior and respects platform anti-spam expectations.

Audit log — every action (scout, draft, grade, approve, reject, post, fail) is recorded in a JSONL audit trail with timestamps, draft IDs, and outcomes. Exportable for compliance or post-hoc review.

Persona system — a single configurable persona (voice, expertise areas, citation style, what to never say) shapes every draft for consistency. Your replies sound like you, not like an LLM.

Manual draft mode — paste a thread URL, the agent auto-scouts the context, the LLM seeds a draft, you edit and save. Full human control with the agent doing the legwork (scouting, context-fetching, citation suggestion).

On-demand passes — instead of waiting for the cron, fire a pass for a specific platform or subreddit on demand from the UI. Useful for active engagement around a launch or a thread you spotted.

Why it's not spam — outreach is anchored on your own knowledge base. Citations point at YOUR documentation, YOUR examples. The system grades drafts for genuine relevance and refuses to engage when it can't add value. The cadence gate keeps the volume human-paced. Supervised mode keeps the human in the loop. The result is closer to "an assistant that helps you keep up with engagement on your own products and topics" than "an outbound bot."

Film Crew — End-to-End Production Pipeline

Five specialized agents collaborate to turn a one-line idea into a finished video. Built on the Swarm Orchestrator, so every role runs in parallel where possible and merges back deterministically.

RoleWhat It Does
ScreenwriterGenerates the script + scene breakdown from a logline
CastingAssigns characters to LoRAs (via the LoRA Trainer plugin) or stock characters
CinematographerProduces a shot list with camera moves, framing, and lens choices
StoryboardGenerates keyframe images for every shot via the image pipeline
EditorAssembles the generated clips into a finished video via the Video Editor

The LoRA Trainer plugin ships alongside — train character/environment/prop LoRAs from reference images on your local GPU (bf16, ~46 MB per LoRA) and route them automatically to the Casting agent.

Music Video — Beat-Synced, Automatic

Give it a song (.mp3 / .wav), a style prompt, and a short narrative — Guaardvark does the rest:

  • Audio analysis & beat detection — the track is analyzed for tempo/beats so cut timing follows the music instead of an arbitrary clock.
  • Director — an LLM writes a distinct prompt for every cut (no mechanical repetition across a long song), keyed to your style + narrative.
  • Storyboards → video — a keyframe still is generated per cut (SDXL/FLUX, optional character LoRAs for identity), then animated with the chosen image-to-video model (Wan 2.2 I2V, etc.).
  • Beat-timed assembly — clips are stretched/filled to land on the beat (clip stretch, fill methods) and assembled into the final cut, with RIFE frame interpolation for smoothness.
  • Honest about the edges — native filters/transitions/effects aren't in yet (the demo's glitch effect was added manually in Shotcut); that's on the near-term roadmap.

Linux & macOS: The final assembly step needs melt (MLT) from Shotcut. ffmpeg is pre-installed by the platform bootstrap. Full commands (brew/apt/flatpak/snap) are in plugins/video_editor/README.md.

Video Generation Pipeline

State-of-the-art video generation running entirely on your GPU. No cloud APIs, no per-minute billing, no content restrictions.

ModelTypeMax DurationNative ResolutionVRAM
Wan 2.2 TI2V-5B (default)Text + Image-to-Video~5s (up to 121 frames @ 24fps)1280x704~11GB
Wan 2.2 (14B MoE)Text-to-Video5s (81 frames @ 16fps)832x48011GB
Wan 2.2 14B I2VImage-to-Video5s (81 frames @ 16fps)832x48011GB
CogVideoX-5BText-to-Video6s (49 frames @ 8fps)720x48016GB
CogVideoX-5B I2VImage-to-Video6s (49 frames @ 8fps)720x48016GB
LTX-2.3 Distilled FP8Text + Image-to-Video~10s (161 frames @ 16fps)768x512~14GB
LTX-2.5 Distilled Int8Text + Image-to-Video~10s (161 frames @ 16fps)768x512~14GB
HunyuanVideo 13B (GGUF Q5)Text-to-Video~3s (73 frames @ 24fps, up to 129)848x480~11GB
HunyuanVideo 13B I2V (GGUF Q5)Image-to-Video~3s (73 frames @ 24fps, up to 129)848x480~11GB
MiniMax H3 (pruned Int8)Text, first-frame, last-frame and first+last-frame → Video with its own stereo soundtrack (dialogue, ambience, score)~7s offered (175 frames @ 24fps; the model trains to 15s)864x480 default, 1344x768 max16GB card; measured 6.5 min for a 5s clip at 864x480, 20 steps, on a 16 GB RTX 40-series card
MiniMax H3 Reference (pruned Int8)Up to 9 images, 3 clips and 3 audio files → Video with soundtrack (identity, motion, voice, editing)samesamesame
  • Resolution options — 512px, 576px, 720px, 1280px, 1920px (1080p), and custom dimensions (aligned per model)
  • Quality tiers — Fast (10 steps), Standard (30), High (40), Maximum (50); a model declares the fewest steps it renders well at and a preset below that floor is raised to it (Wan and MiniMax H3: 20). MiniMax H3 also offers turbo speed profiles (8 and 4 steps) through its distilled LoRAs, installed from Manage Video Models.
  • MiniMax H3 — a video model that generates picture and sound in one pass. Its prompt is compiled into the model's structured format (numbered shots with cut times, speaker ids, tagged dialogue) by Guaardvark, and the Film Crew renders each scene as one spoken window on it. Licensed under the MiniMax H3 Community License, which names the EU, UK, South Korea and USA as territories that need MiniMax's application form; the Video Models modal shows the license and the link, and posts carrying H3 clips add a "Generated with MiniMax H3" line.
  • Frame interpolation — 1x raw, 2x doubled FPS, 2x + upscale for cinema-quality output
  • Prompt enhancement — Cinematic, Realistic, Artistic, Anime, or raw
  • Low VRAM mode — reduces resolution, frames, and inference steps to keep 16GB cards inside budget (mutually exclusive with High consistency); video generation itself needs a 16GB-class card — see docs/HARDWARE.md
  • Batch processing — queue multiple videos from a prompt list via an in-process worker (one batch at a time; ComfyUI primary, offline CogVideoX fallback)
  • ComfyUI integration — one-click launch to the node editor for custom workflows; Wan/LTX require ComfyUI. LTX-2.5 needs ComfyUI ≥ 0.32.0 and a one-time license accept on Lightricks/LTX-2.5 (HF_TOKEN in .env); after download, generation stays local.

Audio Studio — Music, FX, and Neural Voice

Three audio backends in one plugin with shared GPU-arbitration so they don't trample each other or fight Ollama for VRAM.

  • Music generation — ACE-Step v1 (3.5B) for full songs with vocals or instrumental-only mode. Suno-style chip-prompt UX (Genre / Mood / Instrument) with optional LLM "Polish" pass that translates plain English into ACE-Step's tag vocabulary plus a paired negative prompt. ~10 GB VRAM at fp16.
  • FX Lab — Stable Audio Open for sound effects and short ambient pieces. Light, fast, runs alongside other models.
  • Neural Voice — Chatterbox as the primary TTS backend, Kokoro as a fast fallback, Piper for narration with 6 voice profiles included. Used for chat narration, voiceover for videos, and the voice-chat conversational mode.
  • Voice Cloning — opt-in, gated behind an explicit consent prompt before any clone is created or used. Reference clips are kept under your control; the system never auto-clones from incidental audio.
  • Built-in audio player — generated WAVs and MP3s open in an in-app player modal instead of triggering a browser download. Documents page surfaces audio rows with prompt, model, duration, and a waveform.
  • Suno export — bulk-export a Suno library into the local DocumentsPage for use with the other generators.

Video Editor — Shotcut-lite Timeline

A built-in non-linear editor for stitching generated clips, layering text, and rendering finished videos — without leaving the app.

LaneHoldsSource
Videoone clip per timeline (multi-clip tracking on the roadmap)Media Library — drag-and-drop
Textunlimited overlays, draggable on the preview, properties-panel for size/color/rotationAdd-Text button + properties editor
Audioone music or voice clipMedia Library — Audio tab
  • Visual trim slider — Material UI range slider bound to source duration, two thumbs for start/end, live monospace readout. No more typing seconds into number inputs.
  • Tabbed icon-grid library — three tabs (Video / Audio / Images) with counts in the tab labels. 36px tiles, drag from tile to matching timeline track.
  • Real text overlay rendering — backend uses ffmpeg drawtext (9 named positions, optional outline + translucent box, proper escaping for colons/quotes/commas). Original is preserved.
  • Keyboard shortcuts — space to play/pause, arrow keys to scrub, t to add text, del to remove selected, cmd+z for one-step undo.
  • JobOperationGate — render path checks the gate before grabbing the GPU, so a render won't trample an active video generation or upscaling job.
  • Standalone Video Text Overlay tool — for the simple one-off case where you don't need a timeline.

Linux & macOS prerequisites: See plugins/video_editor/README.md ("Linux & macOS Setup") for melt + Shotcut install (ffmpeg is already handled by core platform scripts on brew/apt).

GPU Image Upscaling — 4K and 8K Output

Upscale images and video frames to 4K (3840px) or 8K (7680px) with GPU-accelerated super-resolution — eight models from HAT-L (maximum-quality restoration) to anime-tuned Real-ESRGAN variants, two-pass mode, FP16/BF16 precision with torch.compile, frame-by-frame video upscaling, and an optional watch folder. The full model table is in CAPABILITIES.md.


For the complete, enumerated reference (every tool category, exact model support, plugin manifests, page routes, RAG details, self-improvement internals, vision pipeline, dependency reconciler, backup format, advanced settings, etc.) see CAPABILITIES.md.

The sections above cover the experience and differentiators. The rest of this README focuses on getting started, requirements, architecture notes, operations, and contributing.


Quick Start

Python 3.12 is required for the ML stack. Ubuntu 26.04 ships Python 3.14 by default — ./start.sh installs 3.12 automatically (deadsnakes or uv). Manual installs: use a 3.12 interpreter only.

curl -fsSL https://guaardvark.com/install.sh | bash

This clones to ~/guaardvark (override with GUAARDVARK_HOME=/path) and launches ./start.sh. Re-running it updates an existing install. Prefer doing it by hand? Same thing:

git clone https://github.com/guaardvark/guaardvark.git
cd guaardvark
./start.sh

First run handles everything: Python 3.12, venv, Node dependencies, PostgreSQL, Redis, Ollama, Whisper.cpp, database migrations, frontend build, and all services. Requires your system password once for PostgreSQL setup (and optionally apt packages on fresh Linux installs).

ServiceURL (defaults; see .env for VITE_PORT / FLASK_PORT)
Web UIhttp://localhost:5173
APIhttp://localhost:5000 (macOS: 5055)
Health Checkhttp://localhost:5000/api/health (macOS: 5055)
./start.sh                    # Full startup with health checks
./start.sh --fast             # Reuse venv + node_modules as they are: no installs, no frontend build, no preflight
./start.sh --test             # Health diagnostics
./start.sh --plugins          # Start all enabled plugins
./start.sh --external-ollama  # You run Ollama yourself; never started or stopped by these scripts
./stop.sh                     # Stop Guaardvark (and only the Ollama that start.sh launched)
./stop.sh --keep-ollama       # Stop Guaardvark, leave Ollama running whoever started it
./stop.sh --all               # Also stop your own `ollama serve` and the systemd service

Ollama already running before ./start.sh is adopted, not restarted, and left running on ./stop.sh. To make either policy permanent, flip the switches in Settings → Product Profile → Ollama, or set GUAARDVARK_OLLAMA_KEEP_RUNNING=1 / GUAARDVARK_OLLAMA_EXTERNAL=1 in .env.

Pick a profile

The first start asks what Guaardvark is for here. Creator lists the media workflow — image, video, audio, Film Crew, LoRA, upscaling — and leaves agents, the knowledge index, outreach and automation installed but out of the way; Workstation is everything. Either is a starting point, not a ceiling: switch in Settings → Product Profile, or ./start.sh --profile creator. Details in backend/profiles/README.md; building a distribution of your own is docs/EXTENSIONS.md.

Install via PyPI

pip install guaardvark

The package is the guaardvark command. It talks to a running backend on the configured port, or starts one with start.sh from a checkout it finds through GUAARDVARK_ROOT or the current directory; with no checkout it stops with "Guaardvark installation not found". An MCP client can start the server with guaardvark mcp serve from a pip install plus a checkout.


CLI

Typer + Rich + prompt_toolkit. The PyPI package and the command are guaardvark (llx is a deprecated alias). Tab completion works with or without a leading /; /help imagine shows one command; unknown commands suggest a close match.

guaardvark                              # Interactive REPL
guaardvark status                       # System dashboard
guaardvark chat "explain this codebase" # Chat with RAG context
guaardvark search "query"               # Semantic search
guaardvark files upload report.pdf      # Upload and index
guaardvark plugins list                 # ComfyUI / Ollama / …
guaardvark gpu status                   # VRAM and owner lock
guaardvark mcp install --client cursor  # Wire Guaardvark into Cursor
guaardvark completion zsh               # Shell completion script

Config: ~/.guaardvark/cli.json (legacy ~/.llx/config.json is still read). Themes: default, teal, musk, hacker, vader, guaardvark, day, auto. Short terminals get a compact aardvark banner.

REPL Slash Commands (examples)

/imagine <prompt>       Generate an image (inline preview in Kitty/iTerm)
/video <prompt>         Generate a video from text
/voice <text>           Text-to-speech (plays locally)
/agent [on|off|shot]    Screen-agent mode; shot = desktop screenshot
/web [images|chat]      Open the web UI on the real frontend port
/ingest <path>          Index files or directories for RAG
/plugins list|start     GPU / service plugins
/gpu status|release     VRAM and owner lock
/audio tts|music|sfx    Audio Foundry
/swarm run <prompt>     Parallel agents in worktrees
/lessons begin|end      Lesson pearls
/skills                 List SKILL.md files
/help [query]           Full command reference, or one command

Requirements

DependencyVersionNotes
Python3.12 onlyBackend. 3.13/3.14 not yet supported — the ML stack (numpy<2.0, mediapipe, basicsr/gfpgan) has no wheels for them.
Node.js20+Frontend build
PostgreSQL14+Auto-installed
Redis5.0+Auto-installed
OllamalatestLocal LLM inference
CUDA GPU8GB+ VRAM16GB recommended for video generation

Which tier is your machine? See docs/HARDWARE.md for what runs CPU-only, on 8–12 GB, on the 16 GB design target, and with 24 GB+ of headroom.

GPU Memory Guide

FeatureMinimumRecommended
Chat + RAG4GB8GB
Image generation6GB12GB
Wan 2.2 video16GB*16GB
CogVideoX-5B video16GB20GB
Upscaling0.5GB2–4GB

* Wan's weights fit in ~11GB, but the generation preflight requires a 16GB-class card for every current video family — see docs/HARDWARE.md.

Making It Fast

Chat and agent latency are dominated by a few settings, all in Settings unless noted. Defaults favor visibility while you learn the system; flip these once you trust it:

  • Thinking mode off — extended reasoning (/thinking, or the chat-thinking default under Settings) adds a long deliberation pass to every turn. Off, simple turns answer in a second or two.
  • Developer toggles off — RAG Debug, Verbose Logging, and LLM Debug each add per-request work. Leave them off outside debugging sessions.
  • Pick one reliable model and stay on it — every model switch evicts and reloads weights on the GPU (seconds to a minute). A single mid-size model that stays resident beats a bigger one that thrashes.
  • Mind the VRAM neighbors — renders wait politely for the card, but idle services holding VRAM (voice models, image pipelines) slow everything's admission. The Plugins page shows who's holding what.

Architecture (simplified)

Browser / CLI (PyPI: guaardvark) / MCP Client (Claude Desktop, Cursor, etc.)
    | HTTP + WebSocket / stdio MCP
    v
Flask (~90+ API modules, auto-discovered) + GraphQL + Socket.IO
    |
    +-- AgentBrain (3-tier routing: Reflex → Instinct → Deliberation)
    |
Service Layer (many modules; plugin sidecars for heavy GPU work)
|-- Agent Executor (ReACT + ~60 tool classes + BrainState)
|-- Screen Control (See-Think-Act-Verify + live reasoning stream)
|-- RAG + Autoresearch + Entity extraction
|-- Self-Improvement (detect/fix/verify/broadcast + guardian)
|-- Generation (image/video/audio/voice/content)
|-- Swarm + Film Crew (isolated worktrees + 5-role pipeline)
|-- Servo + Vision Pipeline
|-- System Mapper / Repo intelligence (AST dependency graphs)
|-- GPU Memory Orchestrator + Plugin runner (CUDA sidecar safety)
\-- Interconnector (multi-machine sync + cluster)
    |
+---+---+---+---+---+
v   v   v   v   v   v
PostgreSQL  Redis  Ollama  Agent Display (:99, on-demand)  ComfyUI / Audio Foundry (plugins)
            Celery

Notes:

  • Many components (blueprints, tool registry, plugins) are discovered or declared at runtime.
  • Exact counts drift between releases; see source and CAPABILITIES.md.
  • backend/mcp/config.py controls the default-deny policy for the MCP server.

Frontend: React 18 · Vite · Material-UI v5 · Zustand · Apollo Client · Monaco Editor · Socket.IO
Core models & engines: Gemma4 / Llama-family / Moondream (vision) · Stable Diffusion · Wan 2.2 / CogVideoX · ACE-Step / Chatterbox / Kokoro / Piper · Real-ESRGAN family + HAT · Whisper.cpp


Roadmap (high-level signals)

See the more detailed view in the project plans and CAPABILITIES for status.

Near term / in flight

  • Polish + full platform support for supervised outreach (Discord, X, Facebook posting).
  • Stronger tier-gated memory and conversation context.
  • Continued video/music pipeline unification and Film Crew robustness.
  • Plugin GPU auto-orchestration (intent-driven start/stop based on route + VRAM).
  • Repo intelligence surfaces and more AST-precise agent tools.

Longer term / research

  • Singing voice cloning (Applio-style) with consent + watermarking.
  • Cluster metrics + better multi-node UI bridge.
  • Video editor multi-clip + advanced timeline UX.
  • Embeddings-backed semantic memory recall (via the gpu_embedding plugin).

Not on the roadmap

  • Cloud-by-default or SaaS-hosted primary experience. Local-first is the product.

Release & Docs Maintenance (for contributors)

  • Version source of truth: root VERSION file. backend/app.py, the CLI, and setup.py read it. Avoid hard-coding the version string in README.md, CAPABILITIES.md, or README_zh.md.
  • On release: verify that public screenshots in docs/screenshots/ are up to date, spot-check counts (blueprints via discovery, exposed MCP tools, plugin manifests, CLI catalog), and make sure the "See VERSION" line and CAPABILITIES link are current. Most visual assets live in a separate non-public directory.
  • npm run build in frontend/ before trusting JSX-related docs or claiming UI completeness (the production Rollup build is strict).
  • AGENTS.md + CLAUDE.md + GROK.md are the orientation files for AI coding sessions in this workspace.

Support the Project

Guaardvark is built with love by a solo developer. If it's useful to you:

Star the repo if you find it interesting — it helps with visibility.

Questions, install trouble, or feedback: support@guaardvark.com. Press, partnerships, and business: info@guaardvark.com.


Get Involved

The Guaardvark engine is open source (MIT) and built in public; paid layers and client products built on it are separate and not covered by that license. Whether you want to try the bot, ship a small PR, or hang out with other local-AI builders — here is the short path.

1. Join the community

WhereWhat
DiscordThe Discord bot ships as a plugin — connect it to your own server for local chat, /imagine images, and search against your own install (see plugins/discord/).
GitHub IssuesBugs, features, and labeled starter work
GitHub DiscussionsLonger-form questions if enabled

2. Run it (≈ two commands)

git clone https://github.com/guaardvark/guaardvark.git && cd guaardvark
./start.sh

Web UI → http://localhost:5173 · API → http://localhost:5000 (macOS: 5055)
Details: INSTALL.md · agent mental model · full feature list: CAPABILITIES.md

3. Pick a good first issue

Starter issues carry the good first issue label, each with acceptance criteria and a clear out of scope list. When none are open, the safe zones below are the best place to start.

We aim to review serious PRs within 24–48 hours.

4. Safe vs high-risk contribution zones

Safe (great first PRs)Ask first / high risk
Agent recipes (data/agent/recipes.json)Agent loop, servo, vision targeting
Docs, INSTALL, mental-model guidesSelf-improvement auto-apply paths
CLI polish & offline commandsMCP default-deny / security policy
UI copy, empty states, error messagesCore GPU fork/CUDA plugin runner
Tests for pure helpersProduction auth / credential handling

Full setup, style, and PR expectations: CONTRIBUTING.md

For AI coding agents and heavy contributors: read AGENTS.md (required reading order), CLAUDE.md, and GROK.md. They document the self-coding chokepoint (guarded_code_service.py::apply_exact_replacement), project conventions, dead-code handling, and verification habits.

5. Other ways to help (no code required)

  • Star the repo and share a short demo (screen agent, Film Crew, or Discord /imagine)
  • Report install friction with GPU model + logs from logs/ (an issue, or email support@guaardvark.com)
  • Suggest recipes or workflows you wish worked out of the box
  • Support development: Ko-fi · Sponsors · PayPal

License

MIT License — Copyright (c) 2025-2026 Albenze, Inc.

"Guaardvark"™ and the Guaardvark logo are trademarks of Albenze, Inc. The MIT License covers the code, not the name; see TRADEMARK.md for what you may do with the name without asking.

Guaardvark mascot

Collected info

  • ★ 231 stars
  • ⎇ 48 forks
  • Language: Python
  • Source updated: 9/24/2026

Config for your environment

Replace {MCP_ENDPOINT_URL} with this MCP’s endpoint URL (from its repo or docs above). No API key — you connect directly.

Tool

OS

Config file: ~/.cursor/mcp.json

{
  "mcpServers": {
    "mcp-server": {
      "url": "{MCP_ENDPOINT_URL}"
    }
  }
}

Paste into mcpServers in the config file. Restart Cursor after saving.

If this MCP is also published on mcpchannel.ai, you can subscribe from Browse and use the gateway config there instead.