Settings

Theme

Show HN: Castrag – transcribe and semantically search a large podcast archive

github.com

1 points by snow_mac · 2 comments

Reader

1 thread
snow_macOP

I wanted to query the entire archive of a podcast I listen to without paying for their coaching tier. Built a lightweight pipeline to ingest, summarize, index, and answer questions with citations back to the source audio.

Totals:

- 230 episodes (~145 hours of audio, 1.29M words)

- Total cost: ~$33 using OpenAI's transcription and embedding APIs

- Pipeline: Audio chunking -> Whisper transcription -> Structured summaries -> Semantic index -> Retrieval with source citations

A few engineering details from the post-mortem:

1. Concurrency traps: Initially went with unbounded async calls, which quickly triggered rate-limit stalls and hung threads. Throttled sequential batches ended up being both faster and predictable.

2. Silent failure modes: A shadowed `OPENAI_API_KEY` environment variable silently broke auth midway through a long run without failing fast.

3. Model swapping: Swapping transcription models mid-stream cut processing costs noticeably without hurting search recall.

Question for the HN crowd on retrieval: Right now, retrieval is bare-metal flat NumPy cosine similarity across ~1,500 chunk embeddings. It’s instantaneous in memory (<2ms) and avoids the overhead of running a dedicated vector DB.

For those who have pushed flat in-memory search further: at what order of magnitude (15k? 50k?) does this genuinely fall over before it makes sense to bring in something like `sqlite-vec`, `pgvector`, or an ANN index?

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection