Taking AI-driven developmentinto production
A practical guide to building with coding agents such as Claude Code and Codex — where to start as a non-engineer, and how to design the boundaries of production, billing, and operations. Grounded in Nihonbashi AI Lab’s experience running its own AI products.
Last updated October 6, 2026. Tool details checked against the official Anthropic and OpenAI documentation.
What stays the same when the tools change
The order we hand work to AI
- 01
People decide first
The goal, the source of truth, what must not change, and the acceptance tests
- 02
AI builds
Claude Code or Codex, one reviewable unit at a time
- 03
A different AI reviews
Not the tool that did the building
- 04
People make the call, then verify
A clear go or no-go, checked on the running product
What are Claude Code and Codex?
Claude Code, from Anthropic, and Codex, from OpenAI, are coding agents. They read your codebase, edit files, run commands, and check the result. The common description — write code from natural language in a terminal — is accurate but misses the point.
The essence is that AI enters the development process itself: reading specs, examining existing code, implementing, testing, and explaining its changes. Think of an implementation agent embedded in development, not autocomplete with extra steps. For non-engineers, the desktop apps are the friendliest way in — you can review changes on screen and work on several tasks in parallel.
The two tools come from different companies but are built from very similar parts: a standing instruction file (CLAUDE.md in one, AGENTS.md in the other), skills loaded on demand, hooks that run at fixed points, subagents, MCP connections, and a sandbox with permission controls. That is why this guide covers practices that hold for either tool, not the menus of one.
One misconception worth correcting: neither is a casual no-code tool. Both reward people who can take design responsibility — who can articulate what to build, what must not change, and what done means. The clearer your intent, the denser the implementation you get.
What AI-driven development actually changes
The headline is usually speed. But what matters to management is not velocity itself — it is how decisions change. When a missed requirement could cost millions of yen, you planned defensively. With AI-driven development you can build small, test, discard what fails, and grow what works. What changes is not development speed; it is the number of times you can afford to be wrong.
By 2026 the basis of competition has shifted from coding ability to agent direction: not whether you have a smart AI, but whether you can design the boundaries — context, guardrails, review gates — that let it work autonomously without breaking things. Which tool is ahead changes within months. The boundary design carries over.
And business context beats coding experience. The judgment of what to build — which lives with executives and operators, not just engineers — is the fuel of AI-driven development. What loses value is not the ability to write code, but the ability to write code without business context.
What we have built this way
Rather than listing feature categories, we measure each example by how far past the demo stage it reached in production. Three examples from our own operations, shared at the level of design lessons rather than internals:
Kokai AI — a data product that converts Japanese public information (corporate registrations, subsidies, laws, statistics, procurement) into AI-ready data with citations, served to external AI agents over MCP. The design problem was not integration for its own sake, but shaping data so AI can reference it safely.
GoOCR — an AI OCR service that converts PDFs, images, and scans into structured text, including vertical Japanese and complex layouts. The demo is easy; the hard part is how customer documents are handled: what is never stored, where files are discarded, how billing works, and what happens on failure. No-retention is a design philosophy there, not a feature.
Our own CRM and email infrastructure — lead management, consent trails, unsubscribe flows, notifications. The unglamorous plumbing behind the site you are reading. AI-driven development shows its real strength in how fast it hardens exactly this kind of un-AI-looking work.
How to start, and how to proceed
Start with something real but low-stakes: an internal tool, a report generator, a prototype of an idea you have been sitting on. Describe the goal, the constraints, and the definition of done in writing — the same discipline you would use briefing a contractor. That document becomes the backbone of every session.
Treat the first output as a conversation opener, not a deliverable. Correct it against reality, in small steps, and keep the instructions and decisions on record. Teams that keep this history build a reusable playbook; teams that skip it repeat the same conversations every session.
Which tool to start with? The one you already pay for. If your team has a paid Claude plan, start with Claude Code; if it uses ChatGPT, start with Codex, which is included in ChatGPT plans with usage that varies by plan. The practices are shared, so adding the second tool later requires little additional learning.
The essence: designing boundaries for agents
What makes agentic development work in production is boundary design: what the agent may touch, what it must never touch, what evidence counts as success, and where a human must review. Concretely this means curated context (project instructions and memory), hooks and gates that enforce rules mechanically, and acceptance tests the agent must pass.
One rule we hold to: the AI that built something does not grade it. We build with both Claude Code and Codex, and whichever one did the work, the other reviews it. The point is a reviewer that starts from a clean context and does not carry the habits of the model that wrote the code. Running two tools makes that possible.
Technical debt from vibe coding — building by mood, accepting whatever runs — is real, but it is not inherent to AI. It comes from delegating without design discipline. In our experience running production products, what compounds is not initial speed but the operating patterns that do not collapse under real use.
Doing it yourself vs. working with Nihonbashi AI Lab
If you have someone curious and a low-stakes use case, start yourselves — the learning is the asset. Where Nihonbashi AI Lab adds value is the production boundary: data structure and permission design, billing, failure handling, and the review discipline that keeps quality stable as the system grows.
We offer training that embeds AI-driven development as a team practice — requirements, implementation, review, and operations — and hands-on delivery for systems that need to be dependable from day one. Both are informed by running our own AI products in production, in Japanese business contexts.
Claude Code is a product of Anthropic, and Codex is a product of OpenAI. For current features, supported platforms, and pricing, see the official documentation at code.claude.com and developers.openai.com/codex (we last checked both on October 6, 2026). This page is Nihonbashi AI Lab’s independent guidance, not official content from either company.
Bring your business context — we bring the production discipline
Whether you want your team trained in AI-driven development or a production system built, the first conversation is the same: tell us about the work.
