Open-source agent skills for Claude Code and Codex

Clear requirements in. Tested software out.

The Enso Method is a way of building production software with AI coding agents, and the skills that carry it out. It starts with the requirements: you say what you want in plain language, the open questions get answered up front, and then the agents build it while your attention goes elsewhere.

Tell AI what you want, not what to do.

  • 20+ skills, idea to production
  • Claude Code and Codex
  • Refined daily on production client work
  • Free and open source

The problem

Whatever you don’t tell the AI, it guesses.

AI writes most of the code now, and it usually starts the same way: a sentence or two in the chat, then whatever comes back. Sometimes the guesses are good. When they aren’t, you get software that runs, but that nobody meant to build.

01

Every gap gets a guess.

“Add refunds to checkout” leaves the real questions open. Partial refunds? Store credit? What if the payment provider is down? The agent answers each one itself, without telling you, and you find out later which answers were wrong.

02

Steering it minute by minute.

Then the corrections start: do this, now fix that, no, not like that. You’re telling it what to do instead of what you want, so you can’t step away, and the agent never sees the whole picture.

03

Code nobody can explain.

A month later something breaks, and someone asks what it was supposed to do. The code can’t say, and neither can the next AI session. Without the intent written down, a fix is just another guess.

The fix isn’t a better prompt. It’s doing the thinking before the build: requirements your stakeholders can read, a design, and every open question answered or written down as an assumption. Then the agent builds straight through, phase after phase, and your attention goes back to the decisions only you can make.

The method

Six ideas, one discipline.

The skills are how the method runs. These are what it believes, and they hold whichever skill is doing the work.

  1. 一

    Requirements lead the build.

    Work starts from a PRD and a technical design, and a meaningful change amends them before it is built. Small fixes will still happen straight in the code, as they always do, but the intent stays written down, so each new session starts from what was decided rather than guesses about the code.

  2. 二

    Assumptions are green lights.

    When a business question comes up, the agent researches it, proposes an answer, writes it into the assumptions document with its confidence, and keeps building. Stakeholders confirm or correct it on their own schedule. Nothing waits.

  3. 三

    Research, don’t ask.

    Technical questions are answered against the real code, data and APIs. Only genuine business decisions go to people, and they arrive as proposals, never as open questions.

  4. 四

    Outcomes, not instructions.

    Describe the outcome, not the steps. The method turns it into acceptance criteria, a design and a plan, so the agent works from intent instead of a stream of instructions, and can tell you when your first idea is the wrong place to solve the problem.

  5. 五

    Bite sized, not half built.

    No agent can build a whole system in one go, so large work is cut along natural seams into thin vertical slices, each small enough for one AI session and working on its own. When a phase turns out bigger than planned, it splits again.

  6. 六

    Every acceptance criterion, proven.

    Before work is called finished, each acceptance criterion in the PRD is checked against a test that exercises it, against real systems wherever they exist, and that would fail if the behavior were wrong. Finished means every acceptance criterion is proven, not that the agent says it’s done.

How it works

From idea to production, in one continuous line.

Each step is a skill your agent runs. Each one hands the next a document it can trust, and none of them stops to ask what the documents already answer.

Decide what to build

  1. Roadmap prd-roadmap

    Cuts an idea too big for one PRD into an ordered roadmap of PRDs, each ending in something that ships.

  2. Requirements prd-writing-standards

    A PRD in business language, with testable acceptance criteria and an explicit out of scope.

  3. Design technical-design-writing-standards

    Architecture, interfaces, data models and integration points, held to your house standards.

  4. Refinement autonomous-requirements-refinement

    Pass after pass, finds every gap, researches the answers and folds them into the documents. Checkpointed, so it resumes instead of restarting.

Build it in slices

  1. Phasing phase-split

    Cuts the work into independently shippable phases, each sized to one AI session.

  2. Build implement-from-requirements

    Loads every document, sizes the work, builds one phase and tests it against real resources. No one has to say which files to read.

  3. Continue continuation-prompt

    Hands the next phase to a fresh session that knows exactly which PRD and phase it serves.

Prove it works

  1. Review deliverable-review

    A fresh-eyes agent re-reads every changed file and traces each requirement to the output.

  2. Harden test-hardening

    Maps every acceptance criterion to evidence and upgrades mocks to real sandbox resources.

  3. Commit pre-commit-validation

    Checks the change before it’s committed: only this work’s files, no secrets, and nothing another session committed in the meantime.

The whole flow, step by step, and every skill →

Change requests become a diff, not a detour.

When the business changes its mind mid-project, the change starts in the requirements. They are amended and committed on their own, and the build works from that diff, so the change is reviewable and the requirements keep up with what was built.

  1. 1Amend the PRD and design
  2. 2Commit the amendment alone
  3. 3Build from the diff

The assumptions register

Stakeholders answer proposals, not questions.

Don’t bring people problems; bring them solutions. Before anything is built, the method asks: if we were building this today, what questions would we have? Then it answers each one as best it can, so a stakeholder never gets a bare question. They get a proposed answer, with its evidence, a confidence level, the impact if wrong, and a place to confirm or correct it.

Technical questions rarely need a person: the method researches them against the real data and APIs until it knows. Business questions get the same digging, and whatever stays below 95% confidence goes into assumptions.md for stakeholders to confirm or correct.

An assumption is still a green light, not a blocker. Implementation starts right away, on the requirements and the assumptions together, so nobody waits for answers. If the business says one was wrong, no sweat: the answer goes into the requirements and that part is rebuilt. Implementation is cheap.

How the research works
  • It runs before anything is implemented. autonomous-requirements-refinement finds every question the PRD and the technical design don’t already answer, then researches them in one long session, folding each answer back into the requirements, so it’s all sorted out before the build starts. implementation-readiness-check runs the same check once, for you to review.
  • Technical questions: it digs until it knows. It talks to the data sources, writes temporary test programs against the APIs, and tries the third-party components and integration points until it gets the answer it needs. Most of these end at 95% confidence and never become assumptions. The rest become sharper ones, with a higher confidence.
  • Business questions: sharper assumptions. Research can’t decide what the business wants, but it can raise the confidence of the proposed answer: one that started at 55% might end at 85%, which makes it easier for a stakeholder to respond to. At 95% or more it isn’t declared at all.
  • When the business answers. Confirmations are recorded. Corrections go into the requirements, and the work built on them is re-implemented. The method stops to ask only when the evidence favors no answer at all.

Example entry

RET-004

Can a refund go to store credit after 30 days?

Proposed answer
Yes. After 30 days, refunds go to store credit only, never back to the original card.
Why we think so
The published returns policy says “store credit after 30 days,” and 214 of the last 220 late refunds in the order history went to store credit.
Confidence 80% Impact if wrong Medium
One entry from assumptions.md, as a stakeholder sees it.

The build

Decide up front. Then step away.

Once the requirements are ready, phase-split cuts the work into phases, each big enough to matter and small enough for one Claude Code or Codex session to finish reliably. A large PRD often becomes a dozen of them.

Point your agent at each phase with implement-from-requirements, and it builds, tests and closes it. The decisions were made before the build began, so nobody has to sit in the chat making them. If a phase turns out bigger than planned, it splits again, 2 into 2a and 2b, without stopping to ask.

docs/prds/customer-returns/

  • customer-returns-prd.md
  • technical-design.md
  • assumptions.md
  • phase-1-return-intake.mdClosed
  • phase-2-refund-routing.mdSplit into 2a, 2b
    • phase-2a-card-refunds.mdClosed
    • phase-2b-store-credit.mdImplementing
  • phase-3-returns-dashboard.mdPlanned

Built into the skills

What the agent won’t do.

Eight rules written into the skills, so nobody has to type them into the chat. Each one names the skill that carries it.

Wander out of scope.

implementation-lifecycle

Before it edits a file the plan didn’t call for, it has to name the acceptance criterion that needs the edit. “While I’m here” doesn’t count. When an edit nobody asked for starts breaking things downstream, it undoes that edit instead of fixing everything after it.

Lose the thread in a long chat.

continuation-prompt implement-from-requirements

Each phase starts in a fresh session that loads the PRD, design and phase document itself. Before a handoff, anything the next session needs is written into those documents, so nothing important lives only in a chat that’s about to end.

Bet anything irreversible on a guess.

implementation-lifecycle engineering-principles

An assumption lets the build continue, never an action that can’t be undone. Charging money, deleting data and writing to production wait for a person, however confident the agent is. It builds in dev and sandbox, and reads production only when nothing else can answer the question.

Overrule your stakeholders.

assumptions-document-writing

When new evidence contradicts something a stakeholder already decided, the agent reopens the question under its original entry, with their ruling and the evidence side by side. It never edits their decision to match, and a test that checks it stays red until they rule again.

Take a polished design at its word.

review-technical-design

An AI-written design reads confident whether or not it’s right. For one with hard-to-reverse decisions, a senior review sketches its own design first, checks the load-bearing claims against the real code and infrastructure, and points you to the decisions that most need your judgment.

Ignore how your team builds.

engineering-principles

One line in your repo’s agent instructions names your team’s engineering standards, and they load alongside the method’s own principles. On stack, style, naming, configuration and deployment, your standards win.

Stop at the first fix that works.

engineering-principles

When a fix needs a new lock, queue, retry or special mode, it first asks why the system needs one. It ships the safe fix and writes the deeper design question into the technical design, instead of quietly piling on machinery.

Fake a passing test.

engineering-principles

A failing test is understood before anything changes, never forced green by weakening an assertion, skipping the test or hardcoding the answer. And when data or configuration is missing, the code fails loudly instead of quietly running on a default.

Evidence, not claims

Finished means proven.

An agent saying it’s done doesn’t count. These skills check the work before a phase is called finished, and again once the whole PRD is.

Fresh-eyes review

deliverable-review

A separate agent re-reads every changed file from disk and traces each requirement to what was built. When the new work falls short of the PRD, the work gets fixed; the review never trims a requirement to match the code.

Test hardening

test-hardening

Every acceptance criterion gets credible evidence. Mocks are converted to real sandbox resources, weak assertions become exact ones, and no test is added unless it can name the failure it would catch.

Test audit

test-audit

Once a phase or a whole PRD is finished, an independent audit breaks the code on purpose and watches which tests notice. Its reviewers take the criteria from the PRD, not from the agent that wrote the code, and it keeps fixing until every criterion is proven.

No loose ends

implementation-lifecycle

Everything a session notices is dismissed, done now, or filed where the next session will find it. The wrap-up says what landed, not a list of caveats for you to sort out.

Who it’s for

For when the person who decides isn’t the one at the keyboard.

Most AI methods are written for a solo builder who is developer and product owner at once. That works for a side project. On client and company work, the decisions belong to someone else, a client, a business owner, a department head, and their time is scarce. The Enso Method keeps the build moving on proposals they can confirm or correct later, and leaves them documents they can actually read.

Engineering leaders

Give the whole team one method it can run, and documents anyone can read, instead of a dozen personal prompting styles.

Consultants and agencies

Keep a dozen builds moving at once while clients answer on their own schedule, and hand over systems that come with their own requirements and design.

Teams changing existing systems

Write the requirements for the part that’s changing, and a design of just the difference: what is kept, what is reused, what is switched off, and which behavior must not change.

Why ensō

One stroke. The circle closes.

An ensō is a circle drawn in a single brushstroke in Zen calligraphy. There is no going back over the line, and it is finished when the circle closes. That is the discipline the method asks of every session: decide deliberately, move without hesitation, and close the loop before you call it done.

How it compares

Great methods. Different beliefs.

The leading AI development methods, from spec-driven toolkits like GitHub Spec Kit to skill libraries like Superpowers, are excellent at what they set out to do. Most tackle the same problem; they differ in what they believe: who decides, how far ahead to plan, and what proves the work is done. Here is where the Enso Method stands.

Swipe the table sideways to see every column.

The Enso Method compared with Superpowers, the BMAD Method, GitHub Spec Kit, OpenSpec and other AI development methods
Method Known for Changes amend the requirements first Business decisions reach stakeholders without stalling the build Every acceptance criterion proven against real systems
The Enso Method Intentional systems, built for someone else Built in Built in Built in
Superpowers Test-first rigor Not a focus Not a focus Partly
Matt Pocock’s skills Design interviews Not a focus Not a focus Not a focus
agent-skills (Addy Osmani) Engineering checklists Partly Not a focus Not a focus
GSD Fresh-context execution Not a focus Not a focus Partly
Compound Engineering A learning loop Not a focus Not a focus Partly
BMAD Method Agile team roles Partly Not a focus Partly
AWS AI-DLC Enterprise approval gates Not a focus Not a focus Partly
HumanLayer QRSPI Context engineering Not a focus Not a focus Partly
GitHub Spec Kit Spec-first features Not a focus Not a focus Not a focus
OpenSpec Change-based specs Partly Not a focus Not a focus
pstack Verified parallel agents Not a focus Not a focus Partly
Agentic Coding Flywheel Multi-model plans, agent swarms Not a focus Not a focus Partly

Built in Partly Not a focus Based on each project’s published skills and documentation, September and October 2026.

What each method believes, and where the Enso Method chose differently →

Get started

Install it, then three steps to your first phase.

Install the skills

In Claude Code and Codex the skills install as a plugin, which names each one enso-method:<skill> so none can clash with skills from anywhere else. New to plugins? See how they work in Claude Code and Codex. Other agents get plain copies through the skills CLI. To update, or to work from a clone of the repo, see the setup guide.

claude plugin marketplace add EnsoDynamics/enso-method
claude plugin install enso-method@enso-method
  1. Write the PRD

    In your agent, run /prd-writing-standards (in Codex, pick it from the $ menu) and describe what you want in your own words. Talk it through; the messy version is better than a tidy instruction.

  2. Design and refine

    Run /technical-design-writing-standards to design how it will be built, then /autonomous-requirements-refinement, which researches the open questions in both documents and folds the answers in.

  3. Build

    Run /implement-from-requirements. It splits the work into session-sized phases with phase-split when it needs to, then builds, reviews and tests the first one.

Who’s behind it

Forged on real production work.

The Enso Method is created and maintained by Doug Kerwin, author of The Enterprise Vibe Coding Playbook, at EnsoDynamics. Every rule in it exists because a real client project needed it, and it is refined every day on production systems.

Work with EnsoDynamics →

Questions

Frequently asked.

What is the Enso Method?
The Enso Method is an open-source software development methodology packaged as more than 20 agent skills for Claude Code and OpenAI Codex. It takes work from idea to production through a PRD, a technical design, autonomous refinement, session-sized phases, implementation, review and test hardening, with written requirements, not a chat history, driving the build.
Which AI coding agents does it work with?
Claude Code and OpenAI Codex, where they install as a plugin. The skills follow the open Agent Skills format, a folder with a SKILL.md file, so other agents that read that format can install them with the standard skills CLI.
Do I have to answer the agent’s questions while it builds?
Rarely. Technical questions are researched against your code, data and APIs. A business question the agent is less than 95 percent sure about becomes a written assumption, with its best answer and its confidence, that your stakeholders confirm or correct later while the build continues. It stops to ask only when the evidence favors no answer at all, and then it names the options.
Does it work on an existing codebase?
Yes, and most of its use is on existing systems. The PRD covers the change being made, and the technical design describes only the difference: what is kept, reused or switched off, and which behavior must not change. Technical questions are researched against the real code, so the documents start from the system as it is.
How is it different from spec-driven tools like Spec Kit?
Spec-driven tools usually write a spec for each feature and then leave it behind. In the Enso Method the PRD and technical design carry on: a meaningful change amends them first and is built from that amendment, so the intent behind the system stays written down. It adds an assumptions register for stakeholders, phases sized to one AI session, and evidence against real systems for every acceptance criterion.
What does “ensō” mean?
An ensō is a circle drawn in a single brushstroke in Zen calligraphy. It is complete when the circle closes, which is the idea behind the method’s rule that every session ends with no loose ends. The red seal beside it reads 円 (en), “circle.”