codex-autoresearch
Codex Autoresearch Skill — A self-directed iterative system for Codex that continuously cycles through: modify, verify, retain or discard, and repeat indefinitely. Inspired by Karpathy’s autoresearch concept.
Links
README
From the repo.
Aim. Iterate. Arrive.
Autonomous, measurable experimentation for Codex.
English · 中文 · 日本語 · 한국어 · Français · Deutsch · Español · Português · Русский
Tell Codex what measurable result you want. Codex inspects the repository, confirms the experiment with you, changes one thing, verifies it, keeps improvements, reverts failures, and repeats until the target is reached.
Autoresearch works for test failures, coverage, type errors, warnings, latency, binary size, reproducible security findings, and any other outcome a command can measure.
Quick Start
Install from Codex:
$skill-installer install https://github.com/leo-lilinxiao/codex-autoresearch
Open a clean Git repository with Full Access:
codex --dangerously-bypass-approvals-and-sandbox
Then invoke the skill:
You: $codex-autoresearch
Reduce `python3 scripts/score.py` error_count to 0.
Codex: Baseline: 5
Target: 0 (lower is better)
Scope: src/
Verify: python3 scripts/score.py, JSON key error_count
Guard: python3 -m pytest -q
Run in foreground or background?
You: Background. Go.
Codex launches the confirmed run. No Codex configuration changes or special prompt syntax are required.
See Installation for manual and development installs.
The Loop
inspect evidence
|
change one focused thing
|
commit and measure
|
+-- improved + guard passes --> keep
|
+-- otherwise ---------------> revert
|
append an audit event
|
repeat until target
The control script owns commits, verification, rollback, and state. Codex owns the hypotheses and code changes.
Foreground And Background
| Foreground | Background | |
|---|---|---|
| Runs in | Current Codex task | Detached controller |
| Continuation | Official Codex Goal | One codex exec worker per iteration |
| Best for | Watching and steering live | Long or overnight runs |
| Control | Codex Goal pause/resume | Ask $codex-autoresearch for status, stop, or resume |
Foreground and background use the same experiment rules. A run uses one mode at a time. Foreground continuation uses a Codex Goal; background continuation belongs to the detached controller.
What Gets Confirmed
Before the first write, Codex shows:
- the goal and numeric target;
- repository-relative paths it may change;
- the metric command and explicit parser;
- an optional regression guard;
- foreground or background mode;
- an optional iteration limit.
Initialization requires a clean named Git branch. One run manages one repository.
Results
Run artifacts live in autoresearch-results/ and stay uncommitted:
| Path | Purpose |
|---|---|
run.json | Immutable confirmed configuration |
events.jsonl | Append-only baseline, iteration, stop, and completion history |
logs/ | Full metric, guard, and background worker output |
runtime.json | Background process state |
runtime.log | Background controller lifecycle events |
report.html | Optional, regenerated visual snapshot |
events.jsonl is the state history. Missing, malformed, contradictory, or partial state is an error; the skill never guesses a result from old files or conversational memory.
Review Results
Ask the skill to show the validated experiment history:
$codex-autoresearch show experiment history
Codex Autoresearch
Run: 0a516883 Status: complete Mode: foreground
Metric: error_count 2 -> 0 Target: 0 (lower is better)
SEQ ITER EVENT PREVIOUS TRIAL RETAINED DESCRIPTION
--- ---- -------- -------- ----- -------- ------------------------------------
0 0 baseline - - 2 Initial measurement
1 1 discard 2 3 2 Broaden parser fallback
2 2 keep 2 1 1 Fix nested parser branch
3 3 keep 1 0 0 Remove final parser error
4 3 complete - - 0 retained metric satisfies the target
The same validated events can be exported as TSV or rendered as a self-contained static report:
$codex-autoresearch export experiment history as TSV
$codex-autoresearch generate an HTML report
The report is written to autoresearch-results/report.html. It is a replaceable snapshot, not runtime state.
Safety Model
- Every trial is a Git commit.
- A non-improving trial or failed guard is reverted with
git revert. - Out-of-scope edits, branch changes, HEAD drift, malformed metrics, command failures, timeouts, and generated byproducts stop the run with an exact error and log path.
- Autoresearch artifacts are never staged.
- A run reports
completeonly when the retained metric reaches the confirmed target. - A genuine external blocker is reported explicitly; a difficult or unsuccessful hypothesis is not treated as blocked.
This strictness is intentional. Silent recovery makes long autonomous runs impossible to trust.
Good Metrics
The verify command must exit successfully and place one finite number on its final non-empty stdout line. It may instead print a JSON object on that line when Codex names one numeric key explicitly.
7
{"error_count": 7, "passed": 12}
Use a guard for behavior the metric does not protect, such as a test suite around a latency benchmark. The guard must pass at baseline.
Documentation
| Guide | Contents |
|---|---|
| Installation | Install, update, and verify the skill |
| User Guide | Configuration, lifecycle, state, and troubleshooting |
| Examples | Practical prompts and metric patterns |
| Contributing | Architecture and validation for contributors |
FAQ
Does installation change my Codex settings?
No. Installation copies the skill files. Use a current Codex release so foreground runs can use the built-in Goal capability.
Why Full Access?
Each iteration creates or reverts a Git commit. Restricted sandboxes may block writes under .git. Background runs therefore default to Full Access; workspace-write remains an explicit option when its limitations are acceptable.
Can I stop and resume?
Yes. Interrupt or pause a foreground Goal. For background, invoke $codex-autoresearch and ask for status, stop, or resume with a new direction.
Can it run without Git or across several repos?
No. Git is the experiment memory and rollback boundary. Use one run per repository so commit ownership and metrics remain unambiguous.
Is this only for small changes?
No. One experiment should test one coherent hypothesis. Its size should match the hypothesis, while still being independently measurable and reversible.
Acknowledgments
Inspired by Karpathy's autoresearch, generalized for Codex and software repositories.
Citation
@misc{codex-autoresearch,
author = {Li, Linxiao},
title = {Codex Autoresearch: Autonomous Goal-Driven Experimentation for Codex},
year = {2026},
publisher = {GitHub},
url = {https://github.com/leo-lilinxiao/codex-autoresearch}
}
GitHub also reads CITATION.cff for its Cite this repository menu.
Star History
License
MIT, see LICENSE.
Collected info
- ★ 1,987 stars
- ⎇ 115 forks
- Language: Python
- Source updated: 7/20/2026