Build & Ship: It Works on Your Screen. That's Not the Same as Ready.

AI can build a working application from a description, and a lot of them are genuinely good. What AI does not do is ask the questions nobody asked it: who is allowed to see this record, what happens when the payment succeeds but the page never loads, which key ended up in the repository, and whether the thing you demoed to a customer can survive that customer's second user. This section is about the gap between working and ready.

TL;DR

  • The gap is not code quality. It is unasked questions: access, money, data, and what happens when something fails halfway.
  • A demo proves it can work. It does not prove it cannot break, and those are different claims.
  • The expensive failures are quiet: a logged-in user seeing someone else's record, a charge with nothing delivered, a key sitting in git history.
  • Bottom line: the builder did what it was asked. The gap is the question nobody thought to ask.

The uncomfortable part

Here is the moment that catches people. Someone asks a fair question about the app you built: "can a customer see another customer's data?" And you realize you do not actually know. Not because you were careless. Because you never opened that particular door, and the tool that built it never mentioned there was a door.

That feeling is not a verdict on you. Building software by describing it is a genuine skill and it produces genuinely working products. But it changes where the risk sits: from can I make this to do I know what I made. And the second question does not answer itself, no matter how well the first one went.

What's inside this section

Ship Your App is the readiness half, written for people who built something real without a traditional engineering background: client versus server checks, row-level security, the money path from checkout to fulfillment, logged-in versus allowed, keys and rotation, test mode versus live, data model sanity, knowing what personal data you now hold, accepting or reverting an AI change, and being able to explain your own app to someone who asks.

Agents & AI Engineering is the deeper half: multi-agent orchestration, prompt evaluation and regression testing, token economics, code review discipline, prompt injection, and agent security. This is where the work stops being about one app and starts being about systems that act on their own.

How to use this section

  1. Start with access. Being signed in is not the same as being allowed. This is the single most common and most damaging gap.
  2. Follow the money end to end. Not the happy path: the refund, the failed card, the duplicate notification, the payment that succeeds while the page dies.
  3. Find out what you are holding. You cannot protect or delete data you do not know you collected.
  4. Then make it explainable. If you cannot describe what your app does with data, you cannot answer a customer, a partner, or a buyer.

The honest line: nothing here is a penetration test, a security certification, or a legal review, and none of it makes an app "safe" as a finished state. These are read-off checks. You look at your own screen and answer honestly, and where something is short the result tells you what to look at next -- and says plainly when nothing is. That is a genuinely useful thing and it is not the same as an audit.

FAQ

What usually breaks first in an AI-built app?

Access control. The app checks whether you are logged in but not whether this particular record belongs to you, so any signed-in user can reach another user's data by changing a number in the address bar. It is invisible in a demo because a demo only ever has one user.

Is vibe-coded software inherently unsafe?

No. It is software that was built quickly by describing intent, so the risk concentrates in whatever the description did not mention. The code is often fine. The gaps are in access rules, failure paths, and money handling, the parts nobody thinks to specify.

Why do payments need more than a success page?

Because the browser can close, the network can drop, and the notification can arrive twice. A payment path that only works when the customer stays on the page will eventually take money and deliver nothing, and that failure looks like fraud to the person it happens to.

I found a key in my repository. What now?

Treat it as exposed and rotate it, rather than deleting the file and hoping. Removing a key from the current version does not remove it from the history, and rotation is the only step that actually revokes access. Rotate first, tidy the history second.

Do I need to know what data my app stores?

Yes, and it is usually more than you chose. Logs, analytics, error reports, and AI prompts all accumulate copies of things you never decided to keep. Knowing what you hold is the prerequisite for every other decision about it, including whether you should be holding it at all.

How do I check any of this without being an engineer?

By reading your own screens and answering specific questions honestly: open two accounts and try to reach one from the other, follow a real payment through to delivery, look at what your database actually stores. The checks are observational, not technical.

Find out what your app does when you're not watching

The Vibe-Coded App Hardening Kit ($79) walks the gaps that survive a good demo. Or take the whole readiness slate with Past the Prototype ($459), ten drills covering access, money, data, and change control. One-time, instant download, yours to keep.

Harden your app ($79)

Latest in this section

Tools for this →

AI Reliability: How to Know Your AI Actually Works

AI reliability means proving your AI feature works — not hoping. The four quiet failure surfaces, deterministic checks, and how to gate every release.

Sep 28, 2026 · 9 min read

RAG Evaluation: How to Grade Your Retrieval

RAG evaluation in plain English: why retrieval feels right but pulls wrong chunks, precision and recall explained, gold sets, and grading it in CI.

Sep 28, 2026 · 8 min read

Model Deprecation: When the Model Changes Under You

Model deprecation breaks AI features nobody touched. What changes when a provider retires or updates a model, and how to find out before your users do.

Sep 28, 2026 · 8 min read

AI Agent Reliability: Catch Quiet Failures

AI agent reliability is about the failures you don't see: happy-path passes, quiet drift, wrong tool calls. What to monitor and how to catch them early.

Sep 28, 2026 · 8 min read

LLM Evaluation: How to Test AI Outputs

LLM evaluation, step by step: what an eval is, how to build a test set, and how eval-driven development replaces "it looked good" with a real verdict.

Sep 28, 2026 · 8 min read

Jev AI Explained: The New Model Changing Automation

Jev AI is TypeSafe's first System One model. It returns typed decisions, not prose. What it does, what it costs, and the two claims worth checking.

Sep 20, 2026 · 10 min read

How Jev AI Works: System One Models and RLCD

How Jev AI works: state in, typed decisions out, evaluated in parallel. The three primitives, the real context budget, and what RLCD does and doesn't prove.

Sep 20, 2026 · 7 min read

How to Set Up Jev AI With Claude Code, Cursor & Codex

Jev AI setup, step by step: install the official TypeSafe skill in Claude Code, Cursor or Codex, add your key safely, and build your first typed decision.

Sep 20, 2026 · 7 min read

Jev AI vs LLMs: Speed, Cost, Limits and Best Uses

Jev AI vs LLMs is the wrong contest. Where a decision model wins, where a language model still has to do the work, and the arithmetic at a million calls.

Sep 20, 2026 · 6 min read

When Not to Trust Jev: 9 Documented Failure Modes

Jev AI limits, from TypeSafe's own documentation: nine published failure modes, what each one breaks, and the calibration figure nobody has released yet.

Sep 20, 2026 · 7 min read

DeepSeek V4.1 Flash Changes the Economics of AI Agents

DeepSeek V4.1 Flash runs agent work for $0.15 in and $0.60 out per million tokens off-peak. What the benchmarks show and what it means for your costs.

Sep 11, 2026 · 10 min read

DeepSeek V4.1 Flash Pricing: What a Task Really Costs

DeepSeek V4.1 Flash pricing is $0.15/$0.60 per million tokens off-peak, and US business hours are off-peak. The per-task math, retries and review included.

Sep 11, 2026 · 6 min read

DeepSeek V4.1 Flash vs Opus 5: Where It Wins and Loses

DeepSeek V4.1 Flash vs Opus 5 and GPT-5.6 Sol on 19 vendor benchmarks: it wins older agent tests and trails the newest, hardest ones. The full table.

Sep 11, 2026 · 6 min read

Always-On AI Agents: Cheap Tokens Still Run Away

Always-on AI agents got cheap with DeepSeek V4.1 Flash, but cheap per token is not capped per month. How to run agents all day without an open-ended bill.

Sep 11, 2026 · 6 min read

Is DeepSeek V4.1 Flash Safe for Business Data?

DeepSeek V4.1 Flash data privacy: DeepSeek's privacy policy says it stores personal data in China. What open weights change, and four questions to ask first.

Sep 11, 2026 · 6 min read

Vibe-Coding Security: Making an AI-Built App Safe for Real Users

Vibe coding security means closing the gaps AI builders skip — exposed secrets, client-only auth, no rate limits, no cost caps, and leaky errors.

Sep 10, 2026 · 8 min read

Is Your App Actually Production-Ready? A Pre-Launch Check

Is my app production ready? Run six yes/no checks — secrets out of the browser, server-side auth, rate limits, cost ceilings, sanitized errors, monitoring.

Sep 10, 2026 · 5 min read

The Vibe-Coding Mistakes That Bite You After Launch

Vibe coding mistakes repeat in a predictable set — secrets in the client, frontend-only auth, no rate limits, editable paywalls, and leaky errors.

Sep 10, 2026 · 5 min read

How to Secure an App You Didn't Fully Write

Secure an AI-built app by treating the AI as a junior developer — check secrets, server-side auth, rate limits, error handling, and monitoring.

Sep 10, 2026 · 5 min read

Exposed API Keys and Secrets: The Vibe-Coder's First Fix

Exposed API keys hide in vibe-coded apps because the browser can't read server env vars — move every secret server-side, then rotate the old key.

Sep 10, 2026 · 5 min read

Agentic AI Security: Securing Autonomous Agents

Agentic AI security means gating an agent's four attack surfaces — install, input, memory, and credentials — before you trust it in production.

Sep 1, 2026 · 7 min read

Prompt Injection: What It Is and How to Stop It

Prompt injection hides malicious instructions inside content an AI agent reads. Learn direct vs. indirect prompt injection and how to gate the kill-chain.

Sep 1, 2026 · 5 min read

MCP Security: How to Vet a Server Before You Install It

MCP security means grading a Model Context Protocol server or agent skill before you trust it. Here's what to check before you install — and why.

Sep 1, 2026 · 5 min read

Non-Human Identity Security: Securing AI Credentials

Non-human identity security means grading the API keys, tokens, and service accounts an AI agent runs on. Here's why NHI is the credential attack surface.

Sep 1, 2026 · 6 min read