Mission: Save tokens for the machine. Save orientation for the human.
π Website: bemyagent.md
BEMYAGENT.md is a lightweight, self-bootstrapping protocol that bridges the gap between humans and AI agents. Instead of forcing alignment through code reviews or rigid procedures, it creates a shared workspace where the machine thinks in structured files and the human validates at the right level of abstraction.
When working with AI agents on complex projects, three things break down:
BEMYAGENT.md provides a single markdown file
(BEMYAGENT.md) that acts as a bootstrap prompt. When
fed to an AI assistant, it generates a structured
.bemyagent/ workspace:
.bemyagent/docs/ β Permanent
project memory (architecture, code map, tech stack, decisions)..bemyagent/work/ β Tactical,
volatile memory organized as a Hierarchical Task Network (HTN).| Concept | What it does |
|---|---|
| TTEV Workflow | Think β Task β Execute β Verify. A four-phase cycle where the agent strategizes, plans atomic steps, executes, and self-validates before notifying the human. |
| Lazy Loading | The agent never reads specs, drafts, or decisions during context restoration unless the current task explicitly requires them. Saves tokens by default. |
| Fractal Decomposition (HTN) | If a task is too large, the agent decomposes it into sub-tasks
(e.g., work/1/1.1/, work/1/1.2/). Each
leaf node gets its own TTEV cycle. |
| Context Saturation Check | Before executing, the agent verifies it has enough context (target files, expected behavior, constraints, dependencies). If too much is unclear, it asks instead of guessing. |
| Contextual DNA Mapping (CDM) | During planning, the agent embeds validation criteria directly into each task β scaled by complexity, in three tiers: Micro tasks get none, Standard tasks get Validation criteria, and Heavy tasks get the full set β Drift sensors, Validation criteria and Pivot triggers. |
| Symbiotic Validation | After execution, the agent evaluates its own output against the CDM criteria and produces a verdict (PASS / PASS_WITH_CAVEATS / FAIL) before presenting results. The human validates the sense, the agent has already validated the form. |
| Self-Registration | The agent configures the projectβs native rule files
(.cursorrules, AGENTS.md, etc.) to read
00-ai-rules.md before every task. |
The human controls how much autonomy the agent has:
Independently of pacing, autoModelSwitching lets the
agent use a stronger model for THINK and VERIFY and cheaper tiers
for mechanical EXECUTE steps. It composes with either mode rather
than being a third one.
BEMYAGENT.md into the root of your
project..bemyagent/ directory
structure and templates.BEMYAGENT.md and start a fresh chat session
(the bootstrap context is no longer needed).Thatβs it. From this point on, the agent reads
.bemyagent/docs/00-ai-rules.md before every task and
knows how to operate.
These live here rather than in 00-ai-rules.md so
they cost nothing at session restore β they are for you to run, not
for the agent to carry in context every turn.
Every procedural rule in the protocol forces an artifact to exist; none checks that it is true. This prompt is the reconciliation pass. Paste it into a session roughly monthly:
βCompare
03-code-map.mdvs the real file structure; report drift. Check01-overview.mdenv vars vs actual config. Verify.gitignorecoverage. Check test coverage vs recent changes. List recent decisions missing from05-decisions-and-issues.md. Flag placeholder sections and language inconsistencies in docs/.β
.bemyagent/ is tracked in git by default. Teams
preferring a clean VCS history may .gitignore
work/ β the audit trail is kept locally and lost in
VCS.
In the worktree workflow (00-ai-rules.md Β§8) the
human dispatches one session per worktree, merges via PR, and
resolves conflicts. There is no automated orchestrator by
design.
.bemyagent/
βββ docs/ # Permanent project memory
β βββ 00-ai-rules.md # The protocol itself (agent reads this first)
β βββ 01-overview.md # What the project does, quick start
β βββ 02-architecture.md # System diagram, component roles
β βββ 03-code-map.md # Routes, key functions, data schemas
β βββ 04-tech-stack.md # Technologies, versions, external services
β βββ 05-decisions-and-issues.md # Decision log and known issues
β βββ 06-implementation-plan.md # Milestones and task index
β βββ decisions/ # Complex ADRs (loaded on-demand)
β βββ specs/ # Feature specifications (loaded on-demand)
β βββ drafts/ # Unscoped ideas (loaded on-demand)
βββ work/ # Tactical memory (volatile)
βββ {milestone}/{task}/ # One folder per atomic task
βββ 01_think.md # Strategy & context check
βββ 02_tasks.md # Checklist with CDM criteria
βββ 03_execute.log # What happened (retrospective)
βββ 04_verify.md # Self-validation report
harness/
β measuring whether a rule actually worksNot part of the protocol, and not shipped to
you. BEMYAGENT.md is the only thing you copy
into your project. harness/ is never referenced by it,
never lands in your repo, and its tooling β Node, sqlite, Python β
is not a requirement for using BEMYAGENT: the
protocol is plain markdown and assumes no runtime, no package
manager and no particular operating system. harness/ is
the test environment used to develop the protocol, kept here for
anyone who wants to reuse the method.
The problem it solves: a rule written for an AI agent is a claim about behaviour, and reasoning about that claim predicts the outcome badly. Across three milestones here, most proposed rules did not survive measurement β several turned out inert, and one made the agent measurably worse before it was reworked.
The method β one variable, two arms, N=3 each:
fixture/ into 6 isolated directories.Whatβs inside:
fixture/ β a small tic-tac-toe app (Express +
sqlite + vanilla client, ~120 lines) with real layer separation:
schema β store β API β client β tests. It is deliberately layered so
a feature request cuts through everything at once, which is what
makes decomposition and scoping rules observable. Swap in your own
codebase if you prefer.tokens.py β per-arm cost from agent session
transcripts, cache-weighted (raw token sums mislead: cache reads are
~10Γ cheaper than fresh input).README.md β the method, plus what the harness
cannot measure and the traps that cost real
experiment rounds: consent-shaped rules are unmeasurable because
subagents never treat a coordinator as the user; directory names
leak the hypothesis to the agents; a fixture that advertises itself
as a test changes behaviour; a planted defect that doesnβt actually
exist turns every arm into a different experiment.Useful for anyone tuning agent instructions β prompts, skills, rule files β who wants evidence instead of intuition.
This repository uses the BEMYAGENT.md protocol to develop itself.
The .bemyagent/ directory contains the live workspace
where the protocol is planned, documented, and evolved β using its
own rules.
Explore .bemyagent/work/ to see real TTEV cycles,
CDM annotations, and verification reports in action.
This project is licensed under the MIT License β see the LICENSE file for details.