MCP, agents, loop, testing

MCP and sub-agents

1AI2MCPexternal tools3Browser / database/ files…4Sub-agentsin parallel

MCP and Playwright

MCP (Model Context Protocol) is the standard interface that lets an AI use external tools. Install an MCP for Claude Code and it gets a new group of "tools". The most common example is the Playwright MCP: it opens a browser, clicks buttons and takes screenshots by itself, to check that a web page was built right (official project: microsoft/playwright-mcp).

# Add it (run in your own command line; what follows -- is the command that starts this MCP)
claude mcp add playwright -- npx -y @playwright/mcp@latest

claude mcp list            # see what is installed
claude mcp remove playwright
# In a conversation type /mcp to see the status and do login authorisation

When adding you can choose a scope: local (default, only you, only this project), project (written to .mcp.json in the project root, shareable through git), user (all your projects). Configuration is stored in ~/.claude.json for local/user and in the project's .mcp.json for project. Official pages: MCP, modelcontextprotocol.io.

Honest note: read what an MCP can do before installing it. An MCP that can drive your computer's browser or read files has wide powers. Install only trustworthy ones, and after installing check which tools you have allowed in the permissions.

Sub-agents and loop

Sub-agents

A sub-agent is a "clone" sent out by the main conversation: it has its own separate context and hands back only the conclusion. Good for: ① wide searches (you do not want dozens of files stuffed into the main conversation); ② several unrelated jobs in parallel. Official page: sub-agents. T00's rule: do small tasks yourself, send clones only for big tasks over 5 files; parallel clones must not edit the same files.

loop (scheduled repetition)

The built-in /loop repeats something at an interval, e.g. /loop 5m check whether the deployment is finished. Without an interval the AI decides when to come back. Principle: use it only for things that need repeated checking; do not use it for one-off jobs.

Testing and confirming

Letting an AI write code is fast; the hard part is confirming it did it right. This is the step I miss most often myself, so several layers are built into the flow:

1. Write "acceptance criteria" first

Use the given-when-then style to turn "what counts as done and correct" into a few yes/no clauses covering the normal path, errors, boundaries and permissions. The matching Skill in T00: t00-acceptance-criteria.

2. Really run it

After changing a web page, use browser automation to open it once at phone width and once at PC width and check: any horizontal overflow, any console errors, whether the key flow works, and keep screenshots as evidence. Matching Skills: t00-webapp-testing, t00-deploy-verify. Working locally does not mean working online: after deploying, verify once more on the real URL.

3. Let another AI be the reviewer (LLM-as-a-Judge)

When the output is free text (answers, summaries) and cannot be compared exactly, an LLM can score another LLM's output against a rubric. Key points: make the rubric concrete; regress every time on a fixed set of "standard questions + reference answers"; spot-check the judge's results by hand, never trust them blindly. Introductory articles (in Japanese):

A cheaper way: on your own machine use your subscription's claude -p as the judge instead of a pay-per-use API (T00's evaluation project runs this way).

Honest note: a judge model favours answers in its own style and can be "charmed" into high scores by long answers. It is a "cheap first filter", not the final referee.