~/portfolio
~/portfolio/blog

Writing & Thoughts

My mental model of the world and the things I'm learning. My takes are unpolished and raw and are subjected to change in light of new information.

★ featured•Aug 17, 2026
6 min read

Why Saved Collections Become Graveyards: Building a Semantic Memory Engine

Why bookmarking tools fail: databases store files while humans remember ideas. How we built an open-source memory engine that understands meaning, fights memory decay, and resurfaces forgotten insights.

#Memory Systems#Semantic Search#AI Infrastructure#Open Source#TypeScript
★ featured•Aug 16, 2026
4 min read

When AI Has to Choose What to Believe

How multimodal LLMs navigate conflicting sensory signals, modality bias vs. reasoning uncertainty, and the subtle challenge of source dependence in AI truth discovery.

#Multimodal AI#MLLMs#AI Research#Truth Discovery#Epistemology
★ featured•Aug 15, 2026
4 min read

What 7 Parallel Sub-Agents Taught Me About the Future of AI

Why the industry's obsession with monolithic mega-prompts is wrong—and how trading conventional chaos for simple, hyper-specialized agent architectures changes everything.

#AI#Agents#System Architecture#LLMs
essay•Aug 14, 2026
4 min read

Observability for LLM Applications: What We Actually Log

Dumping raw unstructured text strings into application logs makes post-incident debugging impossible. We reveal our complete production telemetry schema: input signal metrics, prompt version hashes, token economics down to microcents, and distributed traces.

#Observability#Telemetry#Production AI#Distributed Tracing#DevOps
essay•Aug 13, 2026
5 min read

Why Most RAG Pipelines Fail on Multimodal Content

Standard document RAG completely falls apart on 4D video where temporal links and visual demonstrations carry core meaning. Learn why you must synthesize atomic knowledge cards before vectorization, and how hybrid BM25 + vector search solves retrieval.

#RAG#Multimodal AI#Vector Search#Hybrid Search#pgvector
essay•Aug 12, 2026
4 min read

Cost vs Latency: Intelligent Model Routing in Production

Routing every task to flagship models burns money; routing everything to tiny models ruins accuracy. Discover how our heuristic complexity classifier routes 78% of tasks to fast models while escalating edge cases to reasoning engines, slashing costs by 87%.

#Model Routing#Inference Optimization#Cost Engineering#Production AI#LLMs
★ featured•Aug 11, 2026
5 min read

The Architecture Behind an Instagram-to-Knowledge Vault

The full end-to-end systems architecture behind capturing content from a closed, hostile platform and turning it into structured personal assets: edge gateways, Railway media workers, Deepgram transcription, and Supabase pgvector.

#System Architecture#Distributed Systems#PostgreSQL#pgvector#Next.js
essay•Aug 10, 2026
5 min read

Designing Memory That Resurfaces Forgotten Knowledge

Bookmarking without retrieval is a psychological pacifier. We deconstruct the Ebbinghaus Forgetting Curve and show how temporal decay functions, semantic topic clustering, and synthesized digests turn dead archives into a compound-interest engine for your mind.

#Cognitive Architecture#Memory Systems#pgvector#Spaced Repetition#Algorithms
essay•Aug 8, 2026
5 min read

How We Evaluate AI Quality Before Shipping to Users

Vibe checks in playground environments are the single greatest cause of broken AI products. Discover our 150-reel Golden Dataset, multi-tier evaluation rubric, and automated CI quality gates that block prompt regressions before deployment.

#AI Evals#Quality Assurance#Golden Datasets#CI/CD#Prompt Engineering
essay•Aug 6, 2026
5 min read

SSE vs WebSockets for Real-Time AI Products

Modern AI interactions are fundamentally asymmetrical. WebSockets introduce stateful connection overhead and serverless friction. Server-Sent Events (SSE) provide an ultra-lightweight, multiplexed conveyor belt for streaming tokens and ingestion progress.

#Web Architecture#SSE#WebSockets#Real-Time AI#Next.js
★ featured•Aug 4, 2026
5 min read

Building a Self-Healing LLM Retry Engine with Zod

Naive retries that re-send the same prompt into the void are gambling, not engineering. Learn how to convert structured ZodError metadata into targeted cognitive feedback loops that enable LLMs to self-correct validation errors in real time.

#Reliability#TypeScript#Zod#LLMs#Self-Healing
essay•Aug 2, 2026
6 min read

We Benchmarked 5 Extraction Strategies on Instagram Reels

Short-form social video is the most chaotic medium humans have ever built. We benchmarked 1,000 real-world reels across 5 extraction architectures to find the true Pareto frontier of latency, cost, and signal accuracy.

#Multimodal AI#Benchmarks#Computer Vision#Audio Processing#System Architecture
★ featured•Aug 1, 2026
6 min read

Why Deterministic JSON Changed Our Entire AI Pipeline

Large language models introduce pure entropy into software systems. Discover how strict Zod schema enforcement, AST sanitization, and deterministic JSON modes transformed our extraction pipeline from a brittle demo into a resilient, zero-downtime kernel.

#AI Systems#Zod#Type Safety#LLMs#System Architecture