Skip to content
Founder & team guideUpdated 2 Oct 2026

Claude Code Best Practices: AI-Driven Software Development, With a Real Example

We took an app an AI builder made in an afternoon and used Claude Code to make it production-ready — tests, security, accessibility and a new design — logging every prompt, minute and dollar. Here are the eight practices that made it work, and the places where a human still had to step in.

Free: download the starter pack — our prompts, CLAUDE.md, permission rules and screenshot script.

  • ~24 minof Claude Code working time, 11 sessions
  • $8.86API-equivalent cost (covered by our plan)
  • 10 → 42automated tests, all passing
  • 4real defects found and fixed
TL;DR

The short answer

Claude Code can do a remarkable share of real engineering work — but only inside a process that keeps it honest. The process that worked for us: give it a project memory → set permission rules → let it explore and audit before it edits → make the tests trustworthy → fix one problem per fresh session, test-first → have a second session review the result → check it yourself like a user → commit after every step.

  • It found real problems. A race condition, a leak of students’ emails, broken date rules and tests that only passed by accident.
  • It needed checks it could run. Tests, a type-check and screenshots turned “looks done” into “proven done”.
  • A fresh session caught what the first one missed. Three real issues, after the implementer said it was finished.
  • A human still mattered. Product decisions, a check Claude wasn’t permitted to run, and a screen-reader gap no tool flagged.
The real build

The real project behind this guide

In our Manus guide a founder built “Tempo” — a lesson-booking app for a music teacher — with an AI app builder in about 75 minutes of agent time. It works, but it is openly a prototype. That is the most common starting point we see: something that runs, built fast with AI, that nobody has checked for production.

So we gave the same code to Claude Code (version 2.1.286, running Anthropic’s Opus 5.5 model) and asked it to make Tempo production-ready, one step at a time. Each step was a fresh Claude Code session with a written prompt and only the permissions it needed. Between steps we read every change, ran the checks and committed.

Basics

What Claude Code is (and who it’s for)

Claude Code is Anthropic’s AI coding agent. Unlike a chatbot that answers and waits, it reads your project, runs commands like your tests, edits files and works through a problem — asking for permission according to rules you control. You talk to it in plain English.

Where it fits in the three ways to build software today:

  • Bootstrap (free): a founder builds a prototype with an AI app builder. See our Manus guide.
  • AI with a human in the loop (under $20,000): AI writes most of the code; experienced engineers direct, review, secure and ship it. This is where Claude Code shines.
  • Enterprise ($50,000 and up): regulated or large-scale systems, where AI speeds the work but architecture, audits and integration dominate.

More on these worlds in what an app costs in 2026.

10 minutes

Set it up

  1. Choose a plan. Claude Code needs a paid Claude plan or an API account: Pro is $20 a month ($17 a month billed annually), Max starts at $100 a month, and Team seats include it. The free plan does not.
  2. Install it. Open a terminal and run the official installer (or download the desktop app if you prefer not to use a terminal).
  3. Start it in your project. Open the folder your code lives in, type claude, and follow the browser login once.
Install (macOS, Linux, WSL)
$ curl -fsSL https://claude.ai/install.sh | bash
$ claude --version
$ cd your-project
$ claude
Install (Windows PowerShell)
$ irm https://claude.ai/install.ps1 | iex

Everything in this guide also works non-interactively — claude -p "your prompt" — which is how we logged every step. In a normal session you type the same prompts into the chat.

Practice 1

Give it a project memory (CLAUDE.md)

Claude Code starts each session knowing nothing about your project. A file called CLAUDE.md in the project fixes that: it is read at the start of every session. Run /init and Claude writes a first version by studying the code.

Step 1 — the whole prompt
/init

Keep it short and true. Anthropic’s advice is to ask of every line: “Would removing this cause Claude to make mistakes?” If not, cut it. Check it into git so your team shares it. Our real file is in the starter pack.

Practice 2

Set guardrails before it touches code

Permission rules decide what Claude may do without asking. Put them in .claude/settings.json and commit them. Rules are checked in the order deny, ask, allow — so a deny always wins.

.claude/settings.json — what we used
{
  "permissions": {
    "allow": [
      "Bash(npx vitest *)",
      "Bash(npx tsc *)",
      "Bash(git diff *)",
      "Bash(git status)"
    ],
    "deny": [
      "Read(./.env)",
      "Bash(git push *)",
      "Bash(rm *)"
    ]
  },
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [
          {
            "type": "command",
            "command": "jq -r '.tool_input.file_path' | xargs npx prettier --write"
          }
        ]
      }
    ]
  }
}
  • Allow the safe, repetitive things: running tests and the type-check.
  • Deny reading .env, so your database password is never sent to the model. Our tests loaded it at runtime instead.
  • Deny git push and rm: publishing and deleting stay human decisions.
  • A hook auto-formats every file Claude edits. Hooks run every time; instructions in CLAUDE.md are only advice.
Practice 3

Explore and audit before changing anything

Anthropic’s recommended workflow is explore → plan → implement → commit. In plan mode(press Shift+Tab until the status bar shows “plan mode on”) Claude reads and reasons but doesn’t edit. We used it twice: once to have the app explained in plain English, once for a production-readiness audit.

Step 3 — the audit prompt
Act as a senior engineer doing a production-readiness review of this prototype. Do not change any files.
You may run the type-check and the test suite.
Produce a prioritised list of findings in three groups — Must fix before real users, Should fix soon, Can wait.
For each finding give: the file and line, what is wrong, why it matters for the business in one sentence, and how you would verify the fix.
Check at least: authentication and authorisation on every teacher route, input validation, secrets and config, error handling,
data integrity (double submissions, status transitions), the test suite, and anything that only works because of the prototype environment.
Only report problems you have confirmed in the code — say so if something is a guess.

In 2.5 minutes it returned 20 findings grouped into “Must fix”, “Should fix soon” and “Can wait” — each with a file and line, why it matters for the business, and how to verify the fix. It also said what was fine:

Claude Code replied · step 3, audit

Authentication on teacher routes is fine. All five teacher.* procedures use teacherProcedure (server/routers.ts:217-262), and so does system.notifyOwner. The server rejects requests without a valid session, and checks expiry (server/db.ts:66).

Practice 4

Make the tests trustworthy first

Anthropic’s most important tip is to give Claude a way to verify its work — tests, a build, a screenshot. Without one, “looks done” is the only signal. So before fixing anything we fixed the tests.

The prototype’s suite failed 5 of 10 tests — not because the app was broken, but because the tests were stale and one depended on whatever happened to be running on port 3000. Claude Code made the suite start its own server, load secrets without reading them, and refuse to run against any database not clearly named as a test one.

Step 4 — the prompt
Before we fix anything, make the test suite trustworthy. Right now 5 of 10 tests fail.
Goal: `npx vitest run` passes on its own, with one command, every time.
- The two stale unit tests (logout, cookie sameSite): check git log to see whether the code or the test is wrong, then fix the right one. The Lax cookie was a deliberate change.
- The acceptance test must not depend on something already running on port 3000. It should start the app itself on a free port and stop it afterwards.
- It must refuse to run unless the database name contains "test" or "walkthrough", so it can never touch production data.
- `.env` holds a throwaway DATABASE_URL and TEMPO_TEACHER_PASSWORD. You are not allowed to read it; load it with dotenv in the test setup.
- Don't change app behaviour in this step. Don't touch approved copy.
Run the suite until it is green, then tell me what you changed and why each failure happened.
Claude Code replied · step 4, tests

One thing I couldn’t check: I didn’t see the database-name guard actually refuse. The command I tried (DATABASE_URL="mysql://…/tempo_prod" npx vitest run server/tempo.test.ts) wasn’t approved under the permission rules. You can run it yourself…

We ran it. It refused, with a clear message. That honesty — saying what it could not verify — is exactly what you want.

Practice 5

One problem per session, test-first

For each real defect we opened a fresh session (the equivalent of /clear), described the problem and the rule, and asked Claude to write a failing test first, show it failing, then fix it.

Step 5 — the race-condition prompt
Fix audit finding 2: accepting and declining a request can race, because the status check and the update are separate queries.
Work test-first:
1. Write a test that sends an accept and a decline for the same Submitted request at the same time and expects exactly one to succeed and the other to get a clear conflict error. Run it and show me it fails.
2. Make the status change a single conditional UPDATE (only when status is still Submitted) and check the affected-row count.
3. Return a proper tRPC CONFLICT error instead of a plain Error, and make the request-detail page show the error instead of failing silently. Use this exact copy for the message: "This request was already answered. Refresh to see its current status."
4. Run the whole suite and the type-check.
Do not change the double-booking rule; that is a product decision for later.
Claude Code replied · step 5, race fix

Because the race depends on timing, I changed the test to run 10 pairs at once. The next run reproduced the actual bug: expected … to have a length of 1 but got 2, meaning both accept and decline returned success for the same request.

  • Race condition: a teacher could accept and decline the same request at once and both “succeeded”. Now exactly one wins and the other sees a clear message.
  • Privacy: the public confirmation endpoint returned the student’s email and message to anyone with the link. Now it doesn’t; the page still greets the student by email from their own browser.
  • Dates: 20 table-driven tests, 13 of which failed on the old code — impossible dates, dates centuries away, and “today” calculated in the wrong time zone, which rejected real evening bookings.
Practice 6

Get a second opinion from a fresh session

Anthropic notes that a fresh context reviews better “since Claude won’t be biased toward code it just wrote”. After the fixes we opened a new, read-only session and asked it to review the whole branch like a senior engineer reviewing someone else’s pull request.

Step 8 — the review prompt
You are a senior engineer reviewing someone else's pull request. Review every change on this branch since commit 129ed74 (`git diff 129ed74`).
Look for bugs, missed edge cases, security problems, tests that don't really test what they claim, and anything that changed behaviour or approved copy without being asked.
Run the tests and the type-check. Report findings ranked by severity with file and line. Say clearly if you found nothing serious. Do not change any files.
Claude Code replied · step 8, review

I found nothing serious. The three changes the branch is meant to make all look correct… Below that, there are a few medium and low issues, mostly behaviour or copy that changed without being asked for, and some gaps in the tests.

“Nothing serious” — but it was still worth it. It found that after a conflict the teacher’s page stayed stale with live buttons, that the “test database only” safety check would also have accepted names like latest or contest_prod, and that one date boundary was untested. A follow-up session fixed all three, test-first.

Practice 7

Verify like a real user: screenshots and a browser

Tests check logic; they don’t see the screen. For the redesign we gave Claude a small script that screenshots every page at phone and desktop width, and told it to look at the results and keep fixing until every page looked right.

Step 10 — the redesign prompt
This is now the production version of Tempo, and it needs its own look, clearly different from the AI-builder prototype.
Product decision (approved by the owner): replace the painterly prototype style with a calm, professional "studio" style.

Design brief
- Palette: warm paper background #F7F4EE, cards #FFFFFF with a 1px #E4DED3 border, ink text #1F2A44, muted text #5B6475,
  one accent: deep teal #0F766E (hover #115E59). Keep status colours distinct and at least WCAG AA on their tints.
- Type: headings in a serif system stack (ui-serif, Georgia, "Iowan Old Style", serif); body in the system sans stack
  (system-ui, -apple-system, "Segoe UI", Roboto, sans-serif). Remove the Google Fonts link from client/index.html.
- Remove the decorative paintings everywhere. Instead use one subtle CSS-only motif: thin horizontal staff lines (like a music staff)
  in #E4DED3 behind the page header area. No images.
- Flat, crisp: radius 10px on cards and 8px on inputs/buttons, no heavy shadows. Primary buttons solid teal with white text.
- Change the demo banner text to: "Production demo. Hardened with Claude Code for the NerdHeadz guide. Test data only."
  Keep the rest of the user-facing copy exactly as it is.
- Do not change routes, behaviour, data, validation or tests' expectations.

Work method
1. Look at the current screens first: run `python3 scripts/shoot.py before-claude` and read a few of the PNGs in tmp/shots/before-claude/.
   The app is running at http://127.0.0.1:5191 with hot reload.
2. Put the new values in client/src/lib/brandTokens.ts and the CSS variables, then update the components.
3. Run `python3 scripts/shoot.py after` and READ the screenshots at 390 px and 1280 px. Fix anything that looks broken, cramped,
   low-contrast or inconsistent, and shoot again until every page looks right.
4. Update CLAUDE.md's product-constraint lines about tokens and artwork to describe the new style.
5. Run the test suite and the type-check.
Tell me what you changed and anything you were unsure about.
Claude Code replied · step 10, redesign

My first version broke the lines around each nav item and looked wrong in the screenshots, so they now run unbroken behind the text.

Then we ran our own checks in a real browser: 26 scenarios — booking a lesson on a phone, the two-tab accept/decline race, sign-in and sign-out, impossible dates sent straight to the server, and accessibility scans on every page at two widths. The first run passed 24. The two “failures” looked fine on screen but exposed a real gap: error messages were never announced to screen-reader users. No automated scan flags that, because it only appears after a failed submit. Step 11 fixed it; the re-run passed 26 of 26.

Tempo request page showing Accepted badge and the message: This request was already answered. Refresh to see its current status.
The race fix in a real browser: the second tab tried to decline an already-accepted request and got a clear, announced message.
Practice 8

You decide product questions, then commit

The best moments in the run were when Claude Code didn’t guess. When removing the student’s email from the public page also removed their note from the confirmation screen, it said so and asked. When the new 180-day booking limit reused an imperfect error message, it flagged the wording instead of inventing copy.

  • Put rules for copy and design in CLAUDE.md (“use the approved copy verbatim”) — it respected them all run.
  • Answer the decisions it raises; don’t let them pile up in a long session.
  • Read the diff, run the checks, then commit — Claude Code’s checkpoints (Esc Esc or /rewind) are handy, but they are not a replacement for git.
Results

Before and after

Same app, same copy, same data — two very different products.

Manus prototype

Tempo prototype home page with a large impressionist painting of a piano

After Claude Code

Tempo production home page: calm paper background, serif heading, teal button, availability card
Home page. The production version drops the paintings and Google Fonts — and loads much faster.

Manus prototype

Prototype request form on a phone

After Claude Code

Production request form on a phone
Request form on a phone. Behind the new look: real date rules, linked error messages and focus on the first problem.

Manus prototype

Prototype teacher inbox

After Claude Code

Production teacher inbox
Teacher inbox. Same features; race-safe answers and a clearer status palette.
Measured on the same server, production builds, mobile Lighthouse
MeasureManus prototypeAfter Claude Code
Automated tests passing5 of 1042 of 42
Known defects (race, data leak, dates, screen readers)40
Performance score (home)6485
Largest content appears after10.6 s3.5 s
Page weight (home)2,408 KB437 KB
Accessibility score100100

Accessibility scores were already 100 — yet the screen-reader gap was real. Scores are a floor, not a finish line.

Full log

Every step, time and cost

Each row is one fresh Claude Code session. Time is Claude’s working time; cost is Claude Code’s API-equivalent estimate.
StepWhat it didTimeCost
1 · /initWrote a 40-line project memory file (CLAUDE.md)45s$0.52
2 · ExplorePlain-English tour of the app, read-only51s$0.50
3 · Audit20 findings with file and line, read-only2m 34s$0.94
4 · TestsTest suite from 5 failing to all passing, safe by design1m 25s$0.58
5 · Race fixAccept and decline could both “win” — fixed test-first2m 13s$0.76
6 · Privacy fixPublic page stopped exposing student emails1m 09s$0.46
7 · Date rulesImpossible and past dates rejected; teacher’s time zone2m 01s$0.70
8 · ReviewFresh session reviewed all changes, read-only1m 38s$0.84
9 · Review fixesFixed the three gaps the reviewer found1m 17s$0.46
10 · RedesignNew visual identity, checked with its own screenshots7m 44s$2.50
11 · AccessibilityScreen readers now hear form errors1m 56s$0.59
Total11 sessions23m 33s$8.86

Every prompt above is in the starter pack, in order. The redesign was the most expensive step — design takes many look-fix-look rounds.

Honest note

What we did by hand

Between Claude Code’s sessions somebody had to read the changes, run the checks and decide. In this walkthrough that review pass was run by NerdHeadz’s own QA tooling — scripted test runs and browser checks driven by our agent — rather than a person clicking through each screen. On client projects, a senior engineer owns that step.

  • Set up a separate test database and the permission rules.
  • Checked three audit findings in the code before acting on them.
  • Ran the one check Claude wasn’t permitted to run.
  • Deleted the unused artwork it wasn’t allowed to remove.
  • Ran the 26 browser scenarios and Lighthouse; found the screen-reader gap.
  • Made the product calls: privacy over the confirmation note, the new visual direction.

Claude’s working time was about 24 minutes. The review, checks and decisions around it took longer — and that is the point.

Honest boundary

When to bring in engineers

Claude Code multiplies an engineer. It doesn’t replace the judgement. Bring in experienced people when:

  • Mistakes cost money or trust: payments, health or financial data, children’s data.
  • You can’t judge the diff: if you can’t tell whether a change is right, you can’t review it.
  • It touches real systems: your CRM, accounting, older company software.
  • It needs to scale and stay up: monitoring, incident response, maintenance.

Related reading: vibe-coding security risks, a smarter way to prototype AI products and how long an app takes to build in 2026.

Ready to ship it properly?

NerdHeadz turns AI-built prototypes into production software.

We build with Claude Code and keep senior engineers in the loop on every change. Send us your prototype or your brief — we’ll come back with a scoped proposal within 48 hours.

FAQ

Frequently asked questions

What is Claude Code?

Claude Code is Anthropic’s AI coding agent. It runs in your terminal (and in a desktop app and IDE extensions), reads your project, runs commands such as tests, and edits files — asking for permission according to rules you set. You describe what you want in plain English; it explores, plans and implements.

How much does Claude Code cost?

It is included in Claude’s paid plans: Pro is $20 a month ($17 a month billed annually), Max starts at $100 a month, and Team seats include it too; the free plan does not. In our walkthrough, Claude Code reported an API-equivalent cost of $8.86 for all 11 sessions — on our Max plan that was covered by the subscription, not billed separately.

Can a non-technical founder use Claude Code?

You can run it — every prompt in this guide is plain English — but the value comes from judging its work: reading what it changed, knowing what to test, and spotting what it did not check. That is why we position it as AI with a human in the loop. Non-technical founders get the most from it alongside an experienced engineer, or by starting with an AI app builder for the prototype.

How is Claude Code different from an AI app builder like Manus?

App builders generate and host a whole app from a description, which is ideal for a prototype. Claude Code works inside a real code repository with your tests, tools and git history, so it fits the next stage: hardening, extending and maintaining software. In this guide we took an app Manus built and used Claude Code to make it production-ready.

Is code written by Claude Code safe to ship?

Not automatically. In our run it found and fixed a race condition, a data leak and date bugs — and a fresh review session still found three issues the implementing sessions had missed, and our own browser checks found a screen-reader gap. Ship it only after tests, an independent review and real-user checks, with permission rules that keep secrets and destructive commands out of reach.

What is CLAUDE.md?

A plain text file in your project that Claude Code reads at the start of every session: commands, conventions and gotchas it cannot work out from the code. Run /init to generate a first version, then keep it short and correct — Anthropic suggests asking of each line whether removing it would cause mistakes.

How long does it take to see results?

In our walkthrough Claude Code worked for about 24 minutes across 11 sessions to audit the app, fix four real defects, make the tests trustworthy, redesign the interface and fix an accessibility gap. Reviewing each change and running the checks between sessions took longer than Claude’s own working time — and that review is the part you should not skip.

Aleksandr Kamenev

Aleksandr Kamenev

Founder & CEO, NerdHeadz

Aleksandr founded NerdHeadz in 2022 to ship software in production, not decks. The team has delivered 60+ products and uses AI agents and Claude Code on every project — with senior engineers owning each decision.

Sources

Claude Code documentation and pricing were checked on 2 October 2026; products change, so re-check before relying on a detail.

Related: build an app with Manus AI · app costs in 2026 · AI-assisted development · MVP development