~/portfolio
← back to all posts
Aug 8, 2026•5 min read

How We Evaluate AI Quality Before Shipping to Users

Vibe checks in playground environments are the single greatest cause of broken AI products. Discover our 150-reel Golden Dataset, multi-tier evaluation rubric, and automated CI quality gates that block prompt regressions before deployment.

#AI Evals#Quality Assurance#Golden Datasets#CI/CD#Prompt Engineering

Imagine a master watchmaker in Switzerland crafting a mechanical tourbillon. Before that watch is placed into the velvet box and handed to a customer, it doesn't just get a quick glance. It goes into a temperature-controlled vault. It is tested under shock, submerged in water, and measured against atomic clocks for six weeks.

If it gains or loses two seconds a day, it is disassembled and rebuilt from scratch.

Now look at how the majority of modern AI startups test their products.

A developer tweaks a system prompt in a playground. They test it on three hand-picked examples. The output looks reasonably smart. They say: "Looks good to me, vibes are immaculate, ship it."

Two days later, real users start feeding the system chaotic, unpredictable inputs. The prompt breaks. The AI begins summarizing background pop music, hallucinating facts that were never mentioned, or generating bloated 600-word essays when the user wanted a 2-sentence actionable insight.

"Vibe checks" are the cancer of modern AI development.

When we built Vault, we knew that if our product was going to replace a person's notebook and earn their long-term trust, our knowledge extraction pipeline had to meet a watchmaker's standard of precision.

Here is the exact evaluation framework we built to measure, score, and gate every AI model and prompt change before a single line of code reaches production.


1. The Golden Dataset: Curating Chaos

The first mistake people make in AI evaluation is testing on clean, artificial synthetic data.

In the real world, user data is nasty. On Instagram, creators actively try to game the algorithm with confusing visual hooks, background noise, ironic captions, and rapid transitions.

We constructed an immutable Golden Dataset of 150 Real-World Reels spanning ten distinct edge-case categories:

┌─────────────────────────────────────────────────────────────┐
│                    THE 150-REEL GOLDEN DATASET              │
├──────────────────────────────┬──────────────────────────────┤
│ 1. Dense Talking-Head Masterclasses │ 6. Fast Slides & Text Cards    │
│ 2. Pop Music + Educational Overlays │ 7. Multi-Speaker Interviews    │
│ 3. Silent Cooking & Visual Tutorials│ 8. Non-English / Accented Audio│
│ 4. Sarcastic Hooks with Real Meat   │ 9. Low-Resolution Audio Tracks │
│ 5. Heavy Code/Technical Demos       │ 10. Abstract Philosophical Reels│
└──────────────────────────────┴──────────────────────────────┘

For every single reel in this dataset, human domain experts wrote the "Ground Truth Grounding":

  • The exact core concept (Title).
  • The 1-2 sentence core takeaway (Summary).
  • The atomic action steps (Key Points).
  • The correct topic classifications.

2. The Multi-Tier Evaluation Rubric

Every candidate prompt or model change is run against the entire Golden Dataset and evaluated across four distinct dimensions:

graph TD
    A[Reel Test Case] --> B[Extraction Pipeline]
    B --> C[Candidate Knowledge Output]
    
    C --> D1[Tier 1: Deterministic Assertions]
    C --> D2[Tier 2: Information Fidelity]
    C --> D3[Tier 3: Density & Taste Metric]
    C --> D4[Tier 4: Graph & Topic Accuracy]
    
    D1 --> E[Automated Evaluation Report]
    D2 --> E
    D3 --> E
    D4 --> E
    E -->|Score >= 95%| F[Approved for Production]
    E -->|Score < 95%| G[Blocked in CI]

Dimension 1: Deterministic Assertions (Programmatic Pass/Fail)

Before we even look at semantic meaning, we run deterministic TypeScript assertions:

  • Schema Conformance: Did it validate against the Zod schema on the first pass?
  • Length Constraints: Is the summary strictly under 40 words? Is each key point under 15 words?
  • No Markdown / Preamble Leakage: Is the output clean JSON with zero conversational filler ("Here is your summary...")?

Dimension 2: Information Fidelity & Hallucination Rate (LLM-as-a-Judge)

We use a high-capacity reasoning model (e.g., Claude 3.5 Sonnet or GPT-4o) configured with a strict judging rubric:

  • Hallucination Penalty: Did the output claim anything that was NOT stated in the source transcript, OCR text, or visual notes? Any hallucination drops the score to zero.
  • Signal Distillation: Did it identify the actual lesson, or did it get distracted by the intro hook/metaphor?

Dimension 3: The "Taste & Density" Metric

A good summary is not just accurate; it is dense and punchy. We penalize fluff words ("In this video, the speaker discusses how you can easily..."). We compute an Information Density Ratio: $$\text{Density Ratio} = \frac{\text{Atomic Facts Extracted}}{\text{Total Word Count}}$$

Dimension 4: Classification & Topic Precision

Did the model classify a productivity hack under "Productivity" or hallucinate a random category like "Lifestyle & Daily Routines"?


3. Our Automated CI Evaluation Harness

We turned this entire evaluation framework into an automated CLI tool that runs in our continuous integration pipeline:

$ npm run test:evals -- --dataset=golden-150 --model=gemini-2.5-flash

Running Vault AI Evaluation Suite...
Processed: 150/150 test cases

RESULTS:
--------------------------------------------------
Deterministic Schema Pass Rate:  100.0%  (Target: 100%) [PASS]
Information Fidelity Score:       97.4%  (Target: >95%) [PASS]
Information Density Score:        92.1%  (Target: >90%) [PASS]
Topic Classification Accuracy:    96.8%  (Target: >95%) [PASS]
Average Ingestion Latency:        3.42s  (Target: <4.0s)[PASS]
--------------------------------------------------
OVERALL SCORE: 96.5% -> PASSED FOR DEPLOYMENT

If anyone on the team alters a system prompt, swaps a model temperature, or introduces a new extraction parser, the CI pipeline runs the 150-reel eval suite. If the overall fidelity score drops by even 1.5%, the pull request is automatically blocked.


4. The Philosophy: Quality is a Moral Choice

In software, quality is never an accident. It is always the result of intelligent effort and the refusal to compromise.

When you're building in AI, you are dealing with a non-deterministic medium. The only way to tame non-determinism is to surround it with rigid, unforgiving evaluation boundaries.

If you don't evaluate your AI systematically before shipping, your users will become your unpaid QA team—and they will quietly leave you.