~/portfolio
← back to all posts
★ featured•Aug 11, 2026•5 min read

The Architecture Behind an Instagram-to-Knowledge Vault

The full end-to-end systems architecture behind capturing content from a closed, hostile platform and turning it into structured personal assets: edge gateways, Railway media workers, Deepgram transcription, and Supabase pgvector.

#System Architecture#Distributed Systems#PostgreSQL#pgvector#Next.js

Imagine designing a water filtration plant for a river that carries pure mountain water mixed with city sewage, plastic debris, and toxic runoff.

Your plant has to take in that chaotic, muddy torrent, filter out the silt in real time, extract the pure mineral essence, bottle it in sterile glass containers, and deliver it to people's doorsteps in under five seconds—without ever spilling a drop or mixing clean water with sludge.

That is what building an Instagram-to-Knowledge engine actually looks like.

Instagram is one of the most hostile, closed, and noisy platforms on the internet. It is engineered for infinite algorithmic retention, not for open data export or structured knowledge retrieval.

Yet, millions of hours of world-class knowledge—podcasts, coding tutorials, investment breakdowns, culinary techniques, and philosophical lectures—are uploaded exclusively as Instagram Reels every day.

To capture this knowledge and convert it into a permanent personal vault, we had to architect an end-to-end distributed system capable of handling media resolution, multimodal extraction, vector embeddings, and real-time streaming at scale.

Here is the complete architectural blueprint behind Vault.


1. High-Level System Architecture

flowchart TD
    subgraph Capture [1. Capture & Webhook Layer]
        IG[User DMs Reel on Instagram] -->|Meta Webhook / Graph API| API[Next.js Webhook Ingestion API]
        WEB[User Pastes URL on Web UI] --> API
        API -->|Idempotency Verification| DB_JOB[(Supabase Jobs Queue)]
    end

    subgraph Worker [2. Media Worker Cluster]
        DB_JOB -->|Poll / Webhook Dispatch| MW[Railway Media Worker Node]
        MW -->|Stream Download| CDN[Instagram CDN / HikerAPI]
        MW -->|ffmpeg Audio Demux| AUD[16kHz Mono WAV Buffer]
        MW -->|Scene Detection ffmpeg| FRAMES[Keyframe Image Buffers]
    end

    subgraph Extraction [3. Multimodal AI Pipeline]
        AUD -->|Fast STT| STT[Deepgram / Whisper API]
        FRAMES -->|Optical OCR| OCR[Vision OCR Engine]
        STT & OCR --> RECON[LLM Signal Reconciler & Zod Engine]
        RECON -->|Deterministic JSON| INSIGHT[Structured Knowledge Card]
        INSIGHT -->|text-embedding-3-small| EMB[1536-dim Vector Embeddings]
    end

    subgraph Storage [4. Durable Storage & Search]
        INSIGHT & EMB --> PG[(Supabase Postgres + pgvector)]
        PG --> GRAPH[Dynamic Auto-Collections & Topic Graph]
    end

    subgraph Delivery [5. Real-Time Delivery & Client]
        PG -->|SSE Stream / Optimistic UI| UI[Next.js App Router Client]
        PG -->|Spaced Repetition Trigger| DIGEST[Weekly Knowledge Digest]
    end

2. Deep Dive: The 5 Core Subsystems

Subsystem 1: The Capture & Webhook Gateway

  • Meta Graph API Webhooks & HikerAPI Fallbacks: When a user DMs a reel to @vault_ai, Meta fires a webhook payload containing the sender ID and the media link.
  • Idempotency & Deduplication Engine: Meta webhooks are notorious for duplicate deliveries. We compute a SHA-256 hash of the (user_id, reel_shortcode) and execute an atomic UPSERT ... ON CONFLICT DO NOTHING in Postgres. If the job already exists, subsequent duplicate webhooks return 200 OK instantly and terminate.

Subsystem 2: The Media Extraction Worker (Railway / Node + ffmpeg)

Extracting raw video streams on serverless platforms (like AWS Lambda or Vercel) is an anti-pattern due to strict execution time limits and ephemeral /tmp storage caps.

We offloaded all heavy media operations to a dedicated Railway Media Worker:

  1. Media Resolution: Resolves authenticated Instagram CDN streaming URLs via HikerAPI proxy layers.
  2. Audio Demuxing: Uses ffmpeg to strip audio into an optimized 16kHz mono .wav stream, reducing payload size by 90% before speech transcription.
  3. Scene-Change Keyframe Extraction: Instead of dumb interval sampling, ffmpeg -vf "select='gt(scene,0.35)'" isolates only the exact frames where on-screen slides, text overlays, or camera angles shift.

Subsystem 3: The Multimodal Reconciliation Engine

Once speech transcripts and visual keyframes are generated, they pass into our Multimodal Signal Reconciler:

  • Declared Extraction Path: The worker flags the reel as audio_primary (clear speech), visual_fallback (silent reel or text cards), or multimodal (both active).
  • Zod Deterministic Validation: Enforces exact JSON schema bounds, preventing hallucinations and schema drift.
  • Self-Healing Retries: Automatically feeds parsing errors back to the model with lowered temperatures if validation ever fails.

Subsystem 4: Vector Storage & Knowledge Graphs (Supabase + pgvector)

Once an insight is validated, we generate high-density 1536-dimensional embeddings:

  • HNSW Indexing: We index embeddings using pgvector HNSW (Hierarchical Navigable Small World) for sub-10ms approximate nearest neighbor semantic search across millions of items.
  • Auto-Collection Clustering: The system calculates cosine similarity against the user's existing collections to automatically group related knowledge without manual user tagging.
-- Fast semantic vector matching with HNSW index
CREATE INDEX ON reels 
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);

Subsystem 5: Edge Delivery & UI Layer (Next.js App Router)

  • Zero UI Lag: React Server Components deliver lightning-fast initial page loads.
  • Server-Sent Events (SSE): Streams real-time pipeline status (Resolving → Transcribing → Synthesizing → Stored) straight to the browser without heavy WebSocket socket management.
  • Modern Aesthetic Design: Dark-mode first, glassmorphic surfaces, fluid micro-animations, and typography that treats knowledge like a premium personal artifact.

3. Resilience & Failure Modes: How We Handle Chaos

Failure Scenario Defensive Architectural Mechanism
Instagram CDN Link Expired Media worker requests fresh pre-signed CDN token via secondary proxy pool.
Silent Video with Background Music Audio classifier detects music frequency profiles; switches extraction path to visual_fallback.
LLM Output Formatting Error AST sanitizer repairs string quotes; Zod self-healing loop corrects invalid schema properties.
Worker Crash Under Spike Load Distributed Postgres queue automatically releases lock after 60s timeout, enabling worker failover.

The Engineering Conclusion

A knowledge product is only as good as the reliability of its foundation.

By separating lightweight capture (Next.js edge), heavy compute (Railway media workers), deterministic intelligence (Zod-harnessed LLMs), and relational vector storage (Supabase pgvector), Vault transforms the most fleeting social content into an organized, permanent intellectual asset.