← Discover MCPs and Agents
c
AgentAI & MLGitHub

codex-autoresearch

Codex Autoresearch Skill — A self-directed iterative system for Codex that continuously cycles through: modify, verify, retain or discard, and repeat indefinitely. Inspired by Karpathy’s autoresearch concept.

Links

README

From the repo.

Codex Autoresearch

Aim. Iterate. Arrive.

Autonomous, measurable experimentation for Codex.

Codex Skill GitHub Stars MIT License

English · 中文 · 日本語 · 한국어 · Français · Deutsch · Español · Português · Русский


Tell Codex what measurable result you want. Codex inspects the repository, confirms the experiment with you, changes one thing, verifies it, keeps improvements, reverts failures, and repeats until the target is reached.

Autoresearch works for test failures, coverage, type errors, warnings, latency, binary size, reproducible security findings, and any other outcome a command can measure.

Quick Start

Install from Codex:

$skill-installer install https://github.com/leo-lilinxiao/codex-autoresearch

Open a clean Git repository with Full Access:

codex --dangerously-bypass-approvals-and-sandbox

Then invoke the skill:

You:   $codex-autoresearch
       Reduce `python3 scripts/score.py` error_count to 0.

Codex: Baseline: 5
       Target: 0 (lower is better)
       Scope: src/
       Verify: python3 scripts/score.py, JSON key error_count
       Guard: python3 -m pytest -q
       Run in foreground or background?

You:   Background. Go.

Codex launches the confirmed run. No Codex configuration changes or special prompt syntax are required.

See Installation for manual and development installs.

The Loop

inspect evidence
      |
change one focused thing
      |
commit and measure
      |
      +-- improved + guard passes --> keep
      |
      +-- otherwise ---------------> revert
      |
append an audit event
      |
repeat until target

The control script owns commits, verification, rollback, and state. Codex owns the hypotheses and code changes.

Foreground And Background

ForegroundBackground
Runs inCurrent Codex taskDetached controller
ContinuationOfficial Codex GoalOne codex exec worker per iteration
Best forWatching and steering liveLong or overnight runs
ControlCodex Goal pause/resumeAsk $codex-autoresearch for status, stop, or resume

Foreground and background use the same experiment rules. A run uses one mode at a time. Foreground continuation uses a Codex Goal; background continuation belongs to the detached controller.

What Gets Confirmed

Before the first write, Codex shows:

  • the goal and numeric target;
  • repository-relative paths it may change;
  • the metric command and explicit parser;
  • an optional regression guard;
  • foreground or background mode;
  • an optional iteration limit.

Initialization requires a clean named Git branch. One run manages one repository.

Results

Run artifacts live in autoresearch-results/ and stay uncommitted:

PathPurpose
run.jsonImmutable confirmed configuration
events.jsonlAppend-only baseline, iteration, stop, and completion history
logs/Full metric, guard, and background worker output
runtime.jsonBackground process state
runtime.logBackground controller lifecycle events
report.htmlOptional, regenerated visual snapshot

events.jsonl is the state history. Missing, malformed, contradictory, or partial state is an error; the skill never guesses a result from old files or conversational memory.

Review Results

Ask the skill to show the validated experiment history:

$codex-autoresearch show experiment history
Codex Autoresearch
Run: 0a516883  Status: complete  Mode: foreground
Metric: error_count  2 -> 0  Target: 0 (lower is better)

SEQ  ITER  EVENT     PREVIOUS  TRIAL  RETAINED  DESCRIPTION
---  ----  --------  --------  -----  --------  ------------------------------------
  0     0  baseline         -      -         2  Initial measurement
  1     1  discard          2      3         2  Broaden parser fallback
  2     2  keep             2      1         1  Fix nested parser branch
  3     3  keep             1      0         0  Remove final parser error
  4     3  complete         -      -         0  retained metric satisfies the target

The same validated events can be exported as TSV or rendered as a self-contained static report:

$codex-autoresearch export experiment history as TSV
$codex-autoresearch generate an HTML report

The report is written to autoresearch-results/report.html. It is a replaceable snapshot, not runtime state.

Codex Autoresearch HTML report showing metric trajectory and experiment history

Safety Model

  • Every trial is a Git commit.
  • A non-improving trial or failed guard is reverted with git revert.
  • Out-of-scope edits, branch changes, HEAD drift, malformed metrics, command failures, timeouts, and generated byproducts stop the run with an exact error and log path.
  • Autoresearch artifacts are never staged.
  • A run reports complete only when the retained metric reaches the confirmed target.
  • A genuine external blocker is reported explicitly; a difficult or unsuccessful hypothesis is not treated as blocked.

This strictness is intentional. Silent recovery makes long autonomous runs impossible to trust.

Good Metrics

The verify command must exit successfully and place one finite number on its final non-empty stdout line. It may instead print a JSON object on that line when Codex names one numeric key explicitly.

7
{"error_count": 7, "passed": 12}

Use a guard for behavior the metric does not protect, such as a test suite around a latency benchmark. The guard must pass at baseline.

Documentation

GuideContents
InstallationInstall, update, and verify the skill
User GuideConfiguration, lifecycle, state, and troubleshooting
ExamplesPractical prompts and metric patterns
ContributingArchitecture and validation for contributors

FAQ

Does installation change my Codex settings?

No. Installation copies the skill files. Use a current Codex release so foreground runs can use the built-in Goal capability.

Why Full Access?

Each iteration creates or reverts a Git commit. Restricted sandboxes may block writes under .git. Background runs therefore default to Full Access; workspace-write remains an explicit option when its limitations are acceptable.

Can I stop and resume?

Yes. Interrupt or pause a foreground Goal. For background, invoke $codex-autoresearch and ask for status, stop, or resume with a new direction.

Can it run without Git or across several repos?

No. Git is the experiment memory and rollback boundary. Use one run per repository so commit ownership and metrics remain unambiguous.

Is this only for small changes?

No. One experiment should test one coherent hypothesis. Its size should match the hypothesis, while still being independently measurable and reversible.

Acknowledgments

Inspired by Karpathy's autoresearch, generalized for Codex and software repositories.

Citation

@misc{codex-autoresearch,
  author = {Li, Linxiao},
  title = {Codex Autoresearch: Autonomous Goal-Driven Experimentation for Codex},
  year = {2026},
  publisher = {GitHub},
  url = {https://github.com/leo-lilinxiao/codex-autoresearch}
}

GitHub also reads CITATION.cff for its Cite this repository menu.

Star History

Star History Chart

License

MIT, see LICENSE.

Collected info

  • 1,987 stars
  • 115 forks
  • Language: Python
  • Source updated: 7/20/2026