Imagine you're the flight director at NASA monitoring a rocket launch to Mars.
As the rocket ascends through the atmosphere, the only telemetry data on your monitor is a live webcam showing the rocket's exterior and a single text box that prints: "Engine status: Feeling good!"
Ten seconds later, an engine explodes. You look at your console. You have no chamber pressure data, no fuel flow rates, no vibration frequencies, no valve temperatures. You have no idea what failed, why it failed, or how to stop the next rocket from exploding.
That is how most software teams monitor their production AI applications.
They install an APM tool, log an unstructured text dump of the prompt and completion into CloudWatch or Datadog, and call it a day.
When a user complains that a knowledge card is completely inaccurate, the engineering team has to manually scroll through 5,000 lines of messy console logs trying to figure out whether the failure was caused by bad audio, an OCR glitch, a prompt regression, an LLM hallucination, or a database constraint timeout.
In Vault, we treat AI observability not as a passive logging afterthought, but as an exact science of telemetry and diagnostics.
Here is the exact observability schema and tracing architecture we use in production.
1. The 5 Telemetry Dimensions of Every AI Run
For every single knowledge extraction job that runs in Vault, our worker emits a structured, strongly typed telemetry payload containing five distinct operational dimensions:
graph TD
A[Ingestion Execution Span] --> B1[1. Input Signal Metrics]
A --> B2[2. Model & Version Provenance]
A --> B3[3. Latency & Token Economics]
A --> B4[4. Structural & Schema Diagnostics]
A --> B5[5. Downstream User Drift Signals]
export interface IngestionTelemetryEvent {
// 1. Trace Identity
jobId: string;
reelId: string;
userId: string;
timestamp: string;
// 2. Input Signal Quality
inputSignals: {
videoDurationSeconds: number;
audioWordCount: number;
sttConfidenceScore: number;
ocrKeyframeCount: number;
declaredExtractionPath: "audio_primary" | "visual_fallback" | "multimodal";
};
// 3. Model & Version Provenance
modelProvenance: {
provider: "openai" | "gemini" | "anthropic";
modelName: string;
promptVersionHash: string; // SHA-256 of system prompt
temperature: number;
};
// 4. Latency & Token Economics
performance: {
mediaDownloadDurationMs: number;
transcriptionDurationMs: number;
llmDurationMs: number;
totalEndToEndDurationMs: number;
promptTokens: number;
completionTokens: number;
estimatedCostUsd: number;
};
// 5. Schema Validation & Healing
validation: {
firstPassSuccess: boolean;
retryCount: number;
healedErrors: string[]; // Zod error paths if repaired
};
}
2. Why Each Dimension is Critical
1. Prompt Version Hashes (Tracing Prompt Drift)
Whenever someone edits a system prompt, git tracks the change—but production logs often don't. By logging a SHA-256 hash of the exact system prompt template (promptVersionHash), we can instantly correlate whether a spike in customer complaints was caused by a specific prompt commit deployed at 2:15 PM.
2. Input Signal Metrics (Isolating the Culprit)
When an extraction fails or produces a bizarre output, the LLM is often the innocent scapegoat.
By logging sttConfidenceScore and ocrKeyframeCount, we can immediately see:
- If
sttConfidenceScore == 0.12, the microphone audio was unintelligible static. - If
ocrKeyframeCount == 0, the media worker failed to sample keyframes during a video transition.
3. Zod Healing Diagnostics (healedErrors)
By tracking which fields trigger Zod validation failures before self-healing, we pinpoint exactly which schema constraints the model finds confusing. If 30% of runs trigger a validation error on assigned_collection_ids, we know our prompt instructions for collection UUID formatting need clearer few-shot examples.
4. Microcent Cost & Token Tracking
We track estimatedCostUsd down to four decimal places on every run:
$$\text{Cost} = (\text{Prompt Tokens} \times P_{\text{input}}) + (\text{Completion Tokens} \times P_{\text{output}})$$
This enables real-time alerts if a runaway recursive loop or prompt injection attempt spikes token consumption.
3. Real-Time Distributed Tracing
Because an extraction spans multiple systems (Meta Graph Webhook → Railway Media Worker → Deepgram API → OpenAI/Gemini → Supabase Postgres → SSE Client), we pass a unified trace_id header across every boundary:
[Meta Webhook] ──(trace_id: abc-123)──> [Media Worker]
↳ [Deepgram STT] (trace_id: abc-123)
↳ [LLM Extraction] (trace_id: abc-123)
↳ [Supabase Write] (trace_id: abc-123)
↳ [SSE Push to Browser] (trace_id: abc-123)
If a job stalls, one single query in our logging dashboard surfaces the entire lifecycle of that reel in chronological order:
SELECT timestamp, stage, duration_ms, status, error_details
FROM system_telemetry_logs
WHERE trace_id = 'abc-123'
ORDER BY timestamp ASC;
The Philosophy
You cannot optimize what you cannot measure. And you cannot fix what you cannot see.
If you treat LLMs as magic black boxes and rely on blind faith, production will eventually humble you.
BUILD RIGID TELEMETRY INTO YOUR AI FOUNDATION ON DAY ONE. OBSERVE THE SIGNALS, MEASURE THE COSTS, AND HUNT DOWN ENTROPY RELENTLESSLY.