Show HN: Castrag – transcribe and semantically search a large podcast archive
github.comI wanted to query the entire archive of a podcast I listen to without paying for their coaching tier. Built a lightweight pipeline to ingest, summarize, index, and answer questions with citations back to the source audio.
Totals:
- 230 episodes (~145 hours of audio, 1.29M words)
- Total cost: ~$33 using OpenAI's transcription and embedding APIs
- Pipeline: Audio chunking -> Whisper transcription -> Structured summaries -> Semantic index -> Retrieval with source citations
A few engineering details from the post-mortem:
1. Concurrency traps: Initially went with unbounded async calls, which quickly triggered rate-limit stalls and hung threads. Throttled sequential batches ended up being both faster and predictable.
2. Silent failure modes: A shadowed `OPENAI_API_KEY` environment variable silently broke auth midway through a long run without failing fast.
3. Model swapping: Swapping transcription models mid-stream cut processing costs noticeably without hurting search recall.
Question for the HN crowd on retrieval: Right now, retrieval is bare-metal flat NumPy cosine similarity across ~1,500 chunk embeddings. It’s instantaneous in memory (<2ms) and avoids the overhead of running a dedicated vector DB.
For those who have pushed flat in-memory search further: at what order of magnitude (15k? 50k?) does this genuinely fall over before it makes sense to bring in something like `sqlite-vec`, `pgvector`, or an ANN index?