The tools and framework behind ReviewMix, a SaaS built with AI

ReviewMix is built by AI coding agents inside a written working agreement. This is the stack, the agents, and the rules that make their output safe to ship.

By Sorin Maxim · Founder, ReviewMix5 min read

Title card reading "The tools behind ReviewMix, and the rules that keep AI honest"
In short

ReviewMix is built with Claude Code writing the code, Codex reviewing it, and a written working agreement that sorts every change by risk. AI-written code stays safe when risky changes get two independent reviews, tests are proven to fail when the guard breaks, and every claim of success carries its evidence.

The tools to build a SaaS with AI are only half of the answer; the other half is the rules around them. ReviewMix is written by an AI coding agent, reviewed by a second one, and governed by a working agreement that decides how carefully each change is checked. This article covers all three: the agents, the stack, and the framework.

What tools and framework can a non-developer use to build a SaaS with AI?

ReviewMix uses two AI agents and a conventional web stack: Claude Code writes the code, OpenAI's Codex reviews it, and the product runs on TypeScript, Next.js and PostgreSQL. The framework that ties them together is a plain text file of rules in the repository, which the coding agent reads at the start of every turn.

The two agents have different jobs on purpose. Claude Code works in the terminal next to the code: it reads files, proposes a plan, edits, and runs the checks. Codex runs in a read-only sandbox and is asked to read a change it did not write, which makes it a useful first reviewer and a fast drafter of long text. The builder never grades its own work.

The stack is deliberately ordinary:

  • TypeScript in strict mode, so the type checker catches a whole class of mistakes before anything runs.
  • Next.js 15 with the App Router, and React 19 with Tailwind CSS 4 and shadcn/ui for the interface.
  • PostgreSQL with the Drizzle ORM, chosen because its queries read like SQL, which matters when you audit them.
  • Stripe for payments and an EU-routed email service for transactional mail.
  • Hetzner Cloud in Germany, deployed with Coolify, a self-hosted platform where a push to the main branch is the deploy, with Cloudflare in front for CDN and DNS.
  • Self-hosted error tracking with GlitchTip and uptime monitoring with Uptime Kuma.

One app and one database, no microservices. A boring stack is an advantage when the person in charge is still learning to read it: every part has years of documentation and the agent knows it well. Where ReviewMix's data physically goes is set out in where a review widget sends your visitors' data.

What the framework is

The ReviewMix framework is a working agreement: a file of rules the agent loads on every turn, plus a set of documents that are treated as the truth about the product. When the code and a document disagree, the document is fixed first and the code follows. The agreement grew to 46 rules and was cut to 22 on 2 September 2026.

Its centre is a tier system that matches care to blast radius:

  • Tier 1 covers sign-in, payments, database schema and migrations, row-level security and anything that deletes data. It lands as two commits, the documents first, which I review, then the code, and it gets two independent review passes, the second one adversarial.
  • Tier 2 covers business logic and features that destroy nothing. One review pass.
  • Tier 3 covers documentation and copy. One consistency pass, checking that references resolve and nothing contradicts.

Other rules are smaller but earn their place: no new dependency without approval and four supply-chain checks, a weekly dependency audit, and one canonical home for every fact that changes, with every other mention pointing to it.

How do you keep AI-written code safe and reviewable as a solo founder?

A solo founder keeps AI-written code safe by making sure no change is judged only by the agent that wrote it, and by demanding evidence instead of reassurance. In ReviewMix that means independent review passes sized to the risk, tests that are proven to catch the failure they claim to catch, and automatic checks that run whether or not anyone remembers them.

The guard rails, from the moment a change is written to the moment it is live:

  • Hooks on every commit. The type checker runs before a commit is accepted, and a commit-message tripwire refuses a Tier-1 change that bundles documents and code together.
  • Mutation-tested tests. A new test is only trusted after the guard it protects has been deliberately broken, the test has failed for the right reason, and the guard has been restored.
  • Reviews that read the change itself. A reviewer reads the actual diff, never the author's summary. On risky work the second reviewer's job is to try to break the change.
  • Evidence beside every claim. "Nothing else changed", "all tests pass" or "no other files use this" must carry the command output or the list that proves it.
  • A live check after every deploy. A script checks each production route for the response it should give, because some failures only appear on the real server.

None of this assumes the agent is careless. It assumes the agent is confident, which is a different problem: a confident wrong answer looks exactly like a confident right one.

What the rules cost

The rules cost time, and review is the largest share of it: review passes became the single largest cost of the ReviewMix build process. That is the reason the tiers exist. A spelling fix does not need the scrutiny a payment change does, and spending Tier-1 effort on everything would have stopped the project.

Rules also go stale. Some were written for problems the tools later solved, so the agreement is pruned as well as extended, and a new rule has to replace or merge an old one. The working agreement is small enough to be read in full every time, which is the only way it actually gets followed.

How a marketer ended up running this process in the first place is the story in how I built a SaaS without coding.

Frequently asked questions

Which AI tools were used to build ReviewMix?
Claude Code, Anthropic's terminal coding agent, writes and changes the code. OpenAI's Codex reads changes as an independent first reviewer and drafts long text. An adversarial second review on risky changes runs separately.
Is a written framework necessary for AI-built software?
For anything that handles payments or personal data, yes. Without written rules the agent's defaults decide what gets checked. A short file of rules, read on every turn, makes the important checks happen every time.
What does it cost to review AI-written code properly?
Review became the largest single cost of the ReviewMix build process. That is why the review depth is matched to the risk of each change instead of applied evenly.

Related articles

More on building reviewmix and building reviewmix, ai coding.

All articles

Your customers would review you. Ask them.

Start free — no card. Import your Google reviews and embed them today; upgrade when you want the printables kit, replies, daily refresh, or more businesses — or don't.

Free to start · Setup takes minutes · Hosted in the EU