~/portfolio
← back to all posts
★ featured•Aug 1, 2026•6 min read

Why Deterministic JSON Changed Our Entire AI Pipeline

Large language models introduce pure entropy into software systems. Discover how strict Zod schema enforcement, AST sanitization, and deterministic JSON modes transformed our extraction pipeline from a brittle demo into a resilient, zero-downtime kernel.

#AI Systems#Zod#Type Safety#LLMs#System Architecture

Imagine you're an ancient architect in Mesopotamia. You're building an aqueduct to keep a city alive. If the stone blocks don't fit with mathematical precision, the whole canal collapses, water floods the streets, and people die. You don't ask the stone carver to "give you something with good vibes." You demand exact geometric dimensions.

For fifty years, software engineering operated on this exact principle. We built compilers, typesystems, and foreign key constraints because we realized early on: CHAOS IN THE FOUNDATION DESTROYS EVERYTHING ABOVE IT.

Then, the world discovered Large Language Models.

Almost overnight, the entire software industry forgot five decades of hard-won engineering discipline. People started writing prompts like: "Please return your answer as a JSON object, pretty please, and don't include markdown fences."

And what did the model do? It returned a broken trailing comma. Or wrapped the output in ```json. Or inserted an unescaped double quote right in the middle of a title string, blowing up JSON.parse() in production at 3:00 AM.

When we started building Vault—transforming ephemeral, noisy social video into durable personal knowledge—we hit this wall immediately.

Here is why killing prompt-based optimism and moving to strict, deterministic JSON contracts completely revolutionized our entire AI pipeline.


1. The Myth of the "Fast MVP" Chatbot Wrapper

In the startup world, there is a loud consensus: be fast, ship fast, don't worry if it's janky, just get it out.

My take has always been the polar opposite: I'd rather be slow and tasteful than be fast and trash.

If you are building a toy chatbot, a 5% parsing failure rate feels like a minor annoyance. The user just clicks "regenerate." But if you are building an autonomous background pipeline—where a user sends an Instagram Reel via DM, a media worker downloads the video, transcribes the audio, analyzes visual frames, extracts key insights, generates vector embeddings, and links it to existing user collections—A SINGLE PARSING FAILURE BRICKS THE ENTIRE WORKFLOW.

[User DMs Reel] 
   ↳ [Worker Resolves Media] 
      ↳ [Multimodal Signals Extracted] 
         ↳ [LLM Generates Knowledge JSON] 
            ↳ 💥 JSON.parse error: Unexpected token ' in JSON at position 142
               ↳ [Job Stalled in Dead Letter Queue]

When an automated ingestion pipeline fails silently, user trust evaporates. The user thinks your product is broken. And they are right: it is broken.


2. The Anatomy of LLM Chaos: Why Traditional Repair Scripts Fail

Early on, like everyone else, we wrote regex sanitizers. We stripped markdown code blocks. We tried string replacements for trailing commas. We wrote regex patterns to find { and }.

// The "Hope and Pray" pattern that fails in production
export function naiveParseJson(raw: string) {
  const cleaned = raw.replace(/```json/g, "").replace(/```/g, "").trim();
  return JSON.parse(cleaned); // 💥 Fails when the LLM hallucinates an internal quote
}

Why does this fail? Because natural language is inherently messy:

  1. Unescaped Quotes: If a reel discusses "The 'Atomic Habits' framework", the LLM frequently outputs "title": "The "Atomic Habits" framework". Standard parsers crash immediately.
  2. Schema Drift: The model decides to rename key_points to bullet_points or returns a single string instead of an array of strings.
  3. Type Hallucinations: confidence should be a number between 0 and 1, but the LLM returns "0.95" (string) or "High".

When your downstream database has strict Postgres constraints (NOT NULL, foreign keys, vector dimensions), stochastic JSON output is poison.


3. The Paradigm Shift: Deterministic Schema Enforcement

We stopped treating the LLM as a "creative writer" and started treating it as an unreliable RPC endpoint that must conform to a strict binary interface.

We re-architected the extraction layer around three non-negotiable pillars:

graph TD
    A[Raw Multimodal Signals] --> B[LLM with Structured Output Mode]
    B --> C[Raw String Output]
    C --> D[Multi-Stage Deterministic Parser]
    D -->|Valid| E[Zod Type Schema Validation]
    D -->|Malformed| F[AST Sanitizer & Quote Balancer]
    F --> E
    E -->|Pass| G[Postgres / pgvector Storage]
    E -->|Fail| H[Self-Healing LLM Repair Prompt]
    H --> E

Pillar 1: Native Structured Outputs & Tool Calling

Instead of begging the model via free-form system prompts, we enforce strict schema mode (response_format: { type: "json_schema" } or Gemini structured output definitions). The LLM's decoding loop literally constrains token probabilities so it cannot emit tokens that violate the JSON grammar.

Pillar 2: The AST Sanitizer (Pre-Zod)

For providers or fallback models where strict grammar masks aren't guaranteed, we run an AST-level sanitizer that normalizes control characters, repairs truncated brackets, and handles internal unescaped string delimiters.

Pillar 3: Zod as the Source of Truth

We defined every domain entity in Vault as a strict Zod schema:

import { z } from "zod";

export const ExtractedInsightSchema = z.object({
  title: z.string().min(3).max(100),
  summary: z.string().min(10).max(300),
  why_it_matters: z.string().min(5).max(150),
  key_points: z.array(z.string().max(120)).min(1).max(5),
  topics: z.array(z.string()).min(1).max(3),
  entities: z.array(z.string()),
  confidence: z.number().min(0).max(1),
  assigned_collection_ids: z.array(z.string().uuid()),
  new_collections: z.array(
    z.object({
      name: z.string().max(50),
      summary: z.string().max(150),
    })
  ).max(1),
});

export type ExtractedInsight = z.infer<typeof ExtractedInsightSchema>;

4. What Changed Across Our Entire System

Once JSON output became 100% deterministic, everything downstream unlocked:

  1. Zero-Downtime Database Writes: Every payload ingested from the worker cleanly maps to our Supabase tables without schema-mismatch runtime errors.
  2. Instant Frontend Optimistic UI: The Next.js frontend can trust the shape of the data completely. No defensive insight.key_points && Array.isArray(insight.key_points) ? ... rendering hacks littered across UI components.
  3. Automated Vector Categorization: Because assigned_collection_ids and topics conform to an exact vocabulary, our automated graph linking and embedding generation run in sub-second background jobs without human babysitting.
  4. Massive Token & Cost Reduction: With strict schema limits on string lengths (max(150)), the model stops rambling. We slashed inference token consumption by over 40%.

5. The Core Philosophy

Software is a game of managing entropy.

When you introduce AI into your stack, you are inviting pure entropy into your system. If you try to fight entropy with sloppy hacks and rushed duct tape, your product will fall apart the moment real users stress-test it.

DETERMINISTIC JSON IS NOT A NICE-TO-HAVE FEATURE. IT IS THE CONSTITUTION OF YOUR AI ARCHITECTURE.

If you want to build an AI product that endures, stop accepting sloppy outputs. Build slow, build tasteful, and enforce determinism at the gate.