Rebuilding my local AI development environment
A working log: auditing a Claude Code setup that had grown by accretion, pruning it against what Boris Cherny, Addy Osmani and Anthropic's own docs actually recommend, and rebuilding it as a harness — minimal context, deterministic guardrails, visible state.
§ 1Why rebuild at all
My setup had grown the way most do: one install at a time, each reasonable in isolation. By October 2026 the agent loaded two overlapping skill libraries, seven skills twice, a deployment plugin exposing hundreds of tools, and an MCP server that failed to connect on every start. My global settings also described one specific project as if it were the whole machine. None of this broke anything loudly. It just sat in context, every session, before I typed a word.
Anthropic's own guidance puts it bluntly: context is "the most important resource to manage" (Claude Code best practices). Addy Osmani makes the same point about tools: "Loading every tool and MCP into context at startup degrades performance before the agent takes a single action" (Agent Harness Engineering).
§ 2What the practitioners converge on
Before touching anything I collected what the people closest to these tools have written recently. The overlap is striking.
Keep it vanilla. Boris Cherny, who created Claude Code, describes his own setup as "surprisingly vanilla" (Business Insider, May 2026), and reportedly keeps a personal CLAUDE.md of about two lines; when a shared one bloats, the advice is to delete it and add back only what is needed (YC Lightcone transcript).
Prune on a schedule. Osmani: "Installing a useful skill and keeping it forever are separate decisions." Audit every few weeks; archive, then delete unused skills, MCP servers, plugins and slow hooks (Audit your Agent files).
Rules that must hold are not prose. "Every repeated correction is a missing piece of the harness" — encode it as a permission, hook or test (Brownfield Agentic Engineering). Keep CLAUDE.md under about 60 lines, a "pilot's checklist, not a style guide" (Audit your Agent files).
Verify, with a separate checker. Cherny calls verification "probably the most important thing" (summary of his January 2026 thread). Osmani: "Done is a claim, not a proof" — separate the maker from the checker (Loop Engineering).
Prefer CLIs over MCP. The official docs call CLIs like gh "the most context-efficient way" to reach external services (best practices).
§ 3The audit
The audit was read-only: versions, installed plugins, skills, MCP servers, settings, and the Homebrew and npm globals. Findings, in order of how much they mattered:
- Policy without enforcement. My global CLAUDE.md required every plan to go through Plannotator, a browser review UI. The marketplace was registered but the plugin was never installed, so the rule could not fire. A rule the harness can't execute is worse than no rule: it teaches you to ignore the file.
- Duplicated skills. Seven caveman skills existed both as a plugin and as standalone copies, so each appeared twice in the skill list.
- Overlapping libraries. Superpowers and Osmani's agent-skills both ship test-driven development, planning, debugging and review skills.
- Dead weight. A deployment plugin with roughly 400 deferred MCP tools and 40 skills; a broken MCP server (
caveman-shrink, connection closed on every start); a database MCP registered at home scope that was already registered in the one project using it. - Wrong scope. The global
autoMode.environmentblock described one project's repo, secrets and stack — applied to every project on the machine. - Drift. The caveman plugin was 50 commits behind upstream; 78 Homebrew formulae were outdated; the terminal config was an empty file.
§ 4Pruning log
Rule for this pass: verify before deleting, back up config first, and stop when evidence contradicts the plan.
- Removed the seven standalone caveman skill copies; the plugin remains the single source.
- Uninstalled the Vercel Claude Code plugin. The audit showed four projects with
.vercellinks, so the Vercel CLI stays — it is how those projects deploy, and the CLI is the context-cheap path anyway. - Removed the broken
caveman-shrinkMCP server and the home-scope duplicate of the database MCP. - Deleted stale
.bakconfig files. - On macOS: removed iTerm2 (replaced by Ghostty), autojump (to be replaced by zoxide), an unused MongoDB 7.0 alongside the running 8.0, and unused gdown, yarn, rbenv, madge and firebase-tools. About 680 MB freed.
One planned step failed usefully. I meant to move the project-specific autoMode block into that project's settings, but the settings reference lists autoMode as user-or-managed scope only: in a project file it would be silently ignored. The fix is to make the global block generic instead. When the agent tried to rewrite its own permission config, the auto-mode classifier blocked it as self-modification — which is the correct behaviour, so that edit is mine to make by hand.
§ 5Glossary
Terms that recur in everything above, defined the way I'm using them.
- Harness. The model plus everything around it: prompts, tools, context policy, hooks, permissions, sandboxes (Osmani). The thing you are actually configuring.
- Context engineering. Deciding what enters the model's window and when, so it has what it needs and nothing it doesn't (Osmani).
- Agent skill. A
SKILL.mdwith frontmatter. Only its name and description load up front; the body loads when invoked. Cheap individually, expensive in bulk. - Subagent. A separate context window with its own tools — a context firewall for research and review, so the main session keeps the conclusion and not the file dumps.
- Hook. A deterministic script at a lifecycle point (before a tool call, after an edit, on stop). CLAUDE.md is advisory; a hook is guaranteed.
- Plan mode. Read-only exploration and planning before any edit. Cherny reportedly starts most sessions in it.
- Verification loop. The agent runs a pass/fail check and iterates until it passes; the strongest form uses a separate checker.
- Worktree. An isolated git checkout per session, so parallel agents don't collide. "Worktrees isolate changes, not behavior" (Osmani).
- MCP. A protocol for external tool servers. Every server's tool descriptions are prompt text — a context cost and a trust boundary.
- Compaction. Summarising older context near the limit. Useful, lossy; durable state belongs on disk.
- Loop engineering. Building the system that prompts the agent for you —
/goal,/loop, scheduled routines — until a verifiable condition holds (Osmani).
§ 6Next
Still to do, each to be logged here as it lands:
- Generic global
autoModeblock; decide between Superpowers and agent-skills rather than running both. - Fix the unenforced plan-review rule: install Plannotator or delete the rule.
- A status line that shows context usage, rate limits, cost, model, effort and branch.
- An observability layer for agent sessions: what ran, what it cost, where it went wrong.
- Ghostty configuration.
- The macOS engineering toolset, installed from a Brewfile so the whole environment is reproducible.
