~/portfolio
← back to all posts
Aug 2, 2026•6 min read

We Benchmarked 5 Extraction Strategies on Instagram Reels

Short-form social video is the most chaotic medium humans have ever built. We benchmarked 1,000 real-world reels across 5 extraction architectures to find the true Pareto frontier of latency, cost, and signal accuracy.

#Multimodal AI#Benchmarks#Computer Vision#Audio Processing#System Architecture

Imagine walking into an ancient Roman forum. Five hundred people are shouting at once. A street performer is playing a lute in the corner, a philosopher is delivering a speech on stoicism, a vendor is holding up a sign with today's bread prices, and a guy in the back is doing handstands for attention.

Now, someone hands you a notebook and says: "Distill the core knowledge from this scene in five seconds."

That is the exact reality of an Instagram Reel.

Short-form video is the most chaotic medium humans have ever engineered. Information does not sit politely in a single stream. It is scattered across spoken dialogue, background tracks, sarcastic captions, rapid-fire text overlays, facial expressions, and visual demonstrations.

When building Vault, our goal was simple yet audacious: never produce a generic "summary." Convert chaotic social video into high-fidelity, durable personal knowledge.

To get there, we had to answer a fundamental question: How do you accurately extract signal from a medium engineered for pure sensory overload?

We built and benchmarked five distinct extraction strategies on over 1,000 diverse reels (educational lectures, silent coding tutorials, gym technique breakdowns, recipe slides, and lyric-heavy meme clips).

Here is what we discovered.


The 5 Strategies We Tested

graph TD
    subgraph S1 [Strategy 1: Audio-Only]
        A1[Reel] --> B1[Audio Extraction] --> C1[Whisper / Deepgram] --> D1[LLM Insight]
    end

    subgraph S2 [Strategy 2: Frame OCR]
        A2[Reel] --> B2[Fixed Frame Sampler] --> C2[OCR Engine] --> D2[LLM Insight]
    end

    subgraph S3 [Strategy 3: Binary Fallback]
        A3[Reel] --> B3[Whisper]
        B3 -->|Word Count > Threshold| D3[LLM Insight]
        B3 -->|Word Count Low| C3[OCR on Frames] --> D3
    end

    subgraph S4 [Strategy 4: Pure Multimodal Video]
        A4[Reel] --> B4[Full Video Upload] --> C4[Gemini 2.5 / GPT-4o Video API] --> D4[LLM Insight]
    end

    subgraph S5 [Strategy 5: Vault Hybrid Engine]
        A5[Reel] --> B5[Worker Media Extraction]
        B5 --> C5a[Audio Separation + Whisper]
        B5 --> C5b[Scene Change Keyframes + OCR]
        B5 --> C5c[Caption & Creator Meta]
        C5a --> E5[LLM Multimodal Reconciler]
        C5b --> E5
        C5c --> E5
        E5 --> D5[Structured Knowledge Card]
    end

Strategy Breakdown & Benchmark Results

1. Strategy 1: Audio-Only Pipeline (Whisper / Deepgram)

  • How it works: Strip audio track with ffmpeg, run speech-to-text, send raw transcript to an LLM.
  • Cost: ~$0.003 / reel
  • Median Latency: 2.1s
  • The Fatal Flaw: The "Trending Audio" Disaster. Over 42% of educational reels on Instagram use trending pop songs as background audio while delivering 100% of their substance through text overlays or on-screen slides. The audio-only pipeline faithfully transcribed Taylor Swift lyrics while completely missing a breakdown on distributed systems.

2. Strategy 2: Frame OCR Pipeline (Tesseract / Vision OCR)

  • How it works: Sample 1 frame every 2 seconds, run optical character recognition on every frame, concatenate text.
  • Cost: ~$0.006 / reel
  • Median Latency: 4.8s
  • The Fatal Flaw: Frame Blur & Temporal Chaos. Instagram compression creates severe artifacting during fast pans. Furthermore, if a speaker talks continuously for 60 seconds without text overlays, OCR returns empty strings.

3. Strategy 3: The Naive Binary Fallback (Audio First, OCR if Silent)

  • How it works: Run Whisper. If word count > 15 words, assume audio is king and ignore video. If word count < 15 words, fallback to OCR.
  • Cost: ~$0.005 / reel
  • Median Latency: 3.2s
  • The Fatal Flaw: False Positives. A creator talks for 5 seconds saying "Check out this crazy tip!" and then points to 55 seconds of dense text on screen. Because the word count crossed the arbitrary threshold (15 words), the system discarded the visual stream and produced a meaningless 1-sentence summary of the hook.

4. Strategy 4: Direct Native Video LLM (Gemini 2.5 Flash / GPT-4o Video)

  • How it works: Stream the entire .mp4 video buffer directly to a multimodal model endpoint with a structured extraction prompt.
  • Cost: ~$0.018 / reel
  • Median Latency: 7.9s
  • The Pros: Incredible comprehension. The model understands visual demos, tone, and text simultaneously.
  • The Catch: Huge payload overhead, cold starts, payload size limits on serverless functions, and costly token consumption on 90-second 1080p clips.

5. Strategy 5: Vault's Hybrid Multi-Signal Engine

  • How it works:
    1. Media worker extracts audio track + computes video scene changes (keyframe detection via ffmpeg select='gt(scene,0.3)').
    2. Parallel execution: Whisper STT on speech + Vision OCR on high-variance keyframes + Metadata capture (creator, caption).
    3. Signal Reconciler Prompt: We pass all three streams into our structured LLM with a declared path (audio_primary, visual_fallback, or multimodal).
  • Cost: ~$0.007 / reel
  • Median Latency: 3.8s
  • Accuracy Score: 96.4%

The Hard Numbers: 1,000 Reel Benchmark

Strategy Accuracy Rate Median Latency Cost / 1k Reels Critical Failure Mode
1. Audio Only 54.2% 2.1s $3.10 Summarizes pop song lyrics instead of slides
2. Frame OCR Only 61.8% 4.8s $6.40 Completely blind to pure spoken masterclasses
3. Binary Fallback 69.5% 3.2s $4.90 Trapped by intro banter; skips core visuals
4. Direct Video API 93.8% 7.9s $18.20 Latency spikes, payload timeouts, expensive
5. Vault Hybrid 96.4% 3.8s $7.10 Dynamic signal reconciliation with zero blind spots

What We Learned

1. In Social Video, No Single Modality is the "Source of Truth"

Sometimes the spoken voice is 100% of the value. Sometimes the caption is the actual essay and the video is just a 5-second aesthetic loop. Sometimes the on-screen slides are the recipe.

A production AI system cannot make early binary assumptions about human intent.

2. Scene-Change Keyframe Detection Beats Fixed Interval Sampling

Sampling 1 frame every second yields 60 redundant frames for a static speaker, wasting OCR compute. Using ffmpeg scene-detection algorithms captures the exact millisecond a slide transitions, cutting frame volume by 70% while boosting text capture accuracy.

3. Taste Over Shortcuts

The easy way out was to ship Strategy 1 or Strategy 3 in a weekend, declare victory, and watch users churn as their vault filled up with song lyrics.

Taking the time to build a robust, multi-signal pipeline wasn't the fastest route. But it was the only route that honored the craft.