Settings

Theme

Show HN: I put PubMed in a vector DB

pubmedisearch.com

97 points by mpmisko · 27 comments · 1 min read

Reader

Hi HN,

As a researcher, I often found myself struggling with the limitations of keyword-based search when exploring PubMed papers. To address this, I created PubMed Search (https://www.pubmedisearch.com/), a tool that leverages a vector database to enable semantic search across medical research literature.

Some key features:

* Daily updates to ensure access to the latest articles

* Semantic search using latest & greatest embedding models

* Some additional useful info about the papers (tldr, journal, publication date, etc.)

Hope you find it useful!

10 threads
bdangubic

Hey mate, should search by PMID work? Like 35982160 is PMID for "Rare coding variation provides insight into the genetic architecture and phenotypic context of autism" - not seeing this publication at all in search results...

dpifke

Very cool!

Related: the NIST TREC (Text REtrieval Conference) has had several competitions over the years related to improving the searchability of medical data: https://www.trec-cds.org/

If you have novel ideas in this area, you should consider participating. https://trec.nist.gov/

lucas_crocker

This is very cool! 2 questions spring to mind:

1. How much did it cost to embed all those vectors and how many articles did you process? PMC is quite large.

2. Could elaborate a little more on your approach to ranking articles? Because I'm familiar with semantic search via embeddings put did you weight those with impact factors/citations? Like how does one even calculate that?

Anyhow, love the idea.

  • mpmiskoOP

    1. We cover all the articles on PMC. The exact cost is hard to estimate because we did a lot of iterations.

    2. We do weight those ... it is a lot of trial and error and you have to have good & exhaustive benchmarks.

rkwz

Congrats on shipping!

I'm curious how the search results rankings work, doesn't look like it's based on date or number of citations, but seems to be deterministic (persists over multiple searches). I did a keyword search using one word.

kkielhofner

Nice!

Out of curiosity what model(s) are you using to generate the embeddings?

grumpopotamus

What are you embedding exactly? Chunks of documents?

madhatter999

Very promising tool based on a couple of questions I asked it! How did the cleaning of documents look like?

  • mpmiskoOP

    Lots of annoying edge cases as you can imagine, nothing particularly glamorous.

mharig

All 10 thumbs up!

Edit: One suggestion: in the results list, please make the headings links to the articles, too.

drycabinet

Maybe a stupid question, but how do you compare this against GPT-based search engines?

alex_duf

What storage did you go for, and what search approach?

  • mpmiskoOP

    We use pinecone and it is not ideal, looking at https://turbopuffer.com/ now. They look quite promising :)

    • yumraj

      Did you compare pinecone against pgvector with Postgres? Self hosted of course

      • cchance

        Isn't it funny how the best Choice somehow always comes back to Postgres in the end XD (for most)

        • yumraj

          Yes, that’s where I’m these days. I don’t even think of venturing outside of Postgres these days, except for say things like Redis etc. where there are mature and established options for specific use cases.

    • cchance

      What kinda dimensions did you keep it relatively low to keep costs down?

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection