~/portfolio
← back to all posts
★ featured•Aug 4, 2026•5 min read

Building a Self-Healing LLM Retry Engine with Zod

Naive retries that re-send the same prompt into the void are gambling, not engineering. Learn how to convert structured ZodError metadata into targeted cognitive feedback loops that enable LLMs to self-correct validation errors in real time.

#Reliability#TypeScript#Zod#LLMs#Self-Healing

Imagine a biological organism whose cells cannot heal. The moment it gets a minor scratch on its skin, the entire body shuts down, enters shock, and dies.

That would be an evolutionary dead end. Life survived for billions of years because biological systems developed recursive feedback and repair mechanisms. When a cell encounters a damaged protein, it doesn't crash the organism; it tags the defect, repairs it, or recycles the components.

Now look at how 95% of modern AI applications handle runtime failures.

An LLM emits a response with one missing key or a string instead of an array. The JSON validator throws a runtime error. The application logs an uncaught exception, displays a blank white screen or stalls a worker queue, and gives up.

Or worse: developers implement a "naive retry loop" that simply sends the exact same prompt three times into the void, hoping that stochastic probability will miraculously fix itself.

That is not engineering. That is gambling.

When building Vault, where hundreds of multimodal extractions run unattended in background workers every hour, we needed our system to behave like an immune system.

Here is how we built an autonomous, self-healing LLM retry engine with TypeScript and Zod.


1. The Anatomy of an LLM Failure

When an LLM fails in production, the failure falls into one of three distinct categories:

  1. Transport Failure: HTTP 429 (Rate Limit), 503 (Provider Overload), Network Socket Timeout.
  2. Grammar Failure: Invalid JSON syntax, trailing commas, unescaped quotes inside strings.
  3. Semantic / Schema Failure: Syntactically valid JSON that violates domain contracts (e.g., confidence is 1.5 instead of 0.0-1.0, or assigned_collection_ids contains non-UUID strings).

Transport failures require exponential jitter backoff. Grammar failures require AST sanitization. But Schema failures require cognitive feedback.

Traditional Retry:
Prompt ──> LLM ──> Bad Output ──> [Wait 1s] ──> Same Prompt ──> LLM ──> Bad Output (Dead)

Self-Healing Retry:
Prompt ──> LLM ──> Bad Output ──> Zod Inspection ──> Feedback Injection ──> LLM ──> Correct Output

2. Step 1: The Zod Contract

Everything begins with a declarative, runtime-validated schema:

import { z } from "zod";

export const KnowledgeCardSchema = z.object({
  title: z.string().min(5).max(100),
  summary: z.string().min(20).max(300),
  why_it_matters: z.string().min(10).max(150),
  key_points: z.array(z.string().min(5).max(120)).min(2).max(5),
  topics: z.array(
    z.enum([
      "Productivity",
      "Finance",
      "Mindset",
      "Business",
      "AI & Tech",
      "Health & Fitness",
      "Content Creation",
      "Self Improvement"
    ])
  ).min(1).max(3),
  confidence: z.number().min(0).max(1),
  assigned_collection_ids: z.array(z.string().uuid()),
});

export type KnowledgeCard = z.infer<typeof KnowledgeCardSchema>;

Zod gives us something priceless: structured error metadata. When parsing fails via KnowledgeCardSchema.safeParse(data), Zod returns a detailed ZodError containing the exact property path, expected type, and received value.


3. Step 2: Turning Zod Errors into Cognitive Feedback

When validation fails, we don't discard the model's work. The model likely understood 95% of the reel; it just hallucinated a single property or violated a constraint.

We format the ZodError into a human-readable, high-signal diagnostic prompt:

export function formatZodErrors(error: z.ZodError): string {
  return error.errors
    .map((err) => {
      const path = err.path.join(".");
      return `- Field '${path}': ${err.message} (Code: ${err.code})`;
    })
    .join("\n");
}

If the model returned:

{
  "title": "Short",
  "topics": "AI & Tech",
  "confidence": 1.2
}

The error formatter generates:

- Field 'title': String must contain at least 5 character(s)
- Field 'topics': Expected array, received string
- Field 'confidence': Number must be less than or equal to 1

4. Step 3: The Recursive Self-Correction Loop

We pass this diagnostic directly back to the model in a second-pass turn:

export async function executeSelfHealingExtraction<T>(
  prompt: string,
  schema: z.ZodSchema<T>,
  options: { maxRetries?: number; model?: string } = {}
): Promise<T> {
  const maxRetries = options.maxRetries ?? 3;
  let attempt = 0;
  let conversationHistory: Array<{ role: "user" | "assistant"; content: string }> = [
    { role: "user", content: prompt },
  ];

  while (attempt < maxRetries) {
    attempt++;

    // 1. Invoke LLM with adjusted temperature
    // Lower temperature on retries to reduce stochastic variance
    const temperature = Math.max(0.1, 0.7 - attempt * 0.2);
    const rawOutput = await callLLM(conversationHistory, { temperature });

    // 2. Multi-stage JSON sanitizer
    const parsedJson = parseLLMJson(rawOutput);
    if (!parsedJson) {
      conversationHistory.push({ role: "assistant", content: rawOutput });
      conversationHistory.push({
        role: "user",
        content: `Your response was not valid JSON. Return ONLY a valid, parsable JSON object with no markdown fences.`,
      });
      continue;
    }

    // 3. Zod validation
    const validationResult = schema.safeParse(parsedJson);

    if (validationResult.success) {
      return validationResult.data;
    }

    // 4. Build targeted healing prompt
    const errorDetails = formatZodErrors(validationResult.error);
    
    conversationHistory.push({ role: "assistant", content: rawOutput });
    conversationHistory.push({
      role: "user",
      content: `Your output had validation errors:\n${errorDetails}\n\nFix ONLY these fields and return the complete corrected JSON matching the required schema.`,
    });
  }

  throw new Error(`Self-healing extraction failed after ${maxRetries} attempts.`);
}

5. The Production Impact: Moving from 92% to 99.8% Reliability

Before implementing self-healing loops with Zod, our extraction pipeline had an 8% edge-case failure rate. At 10,000 reels ingested per week, that meant 800 broken extractions landing in dead letter queues requiring manual debugging.

With self-healing retries:

  • First-pass success rate: 92.4%
  • Second-pass (Self-Healed) success rate: 99.81%
  • Third-pass recovery: 99.96%

The model almost always fixes its own mistake on the first retry when given the exact Zod diff.


The Takeaway

Building production AI is not about hoping your model never hallucinates. It's about designing a deterministic harness around an inherently non-deterministic engine.

When you pair strict Zod types with feedback-driven retry loops, your pipeline stops being a fragile demo and starts operating like an industrial machine.