It's interesting that folks go down this hacky route when they can use something like Vespa, which is orders of magnitude better from a performance, relevance, scalability, and developer ergonomics perspective.
It makes terrible operational sense. What are the HA/DR, sharding, replica, and backup strategies and tools for pg_vector? What are the embedding integration and relevance tools? What are the reindexing strategies? What are the scaling, caching, and CPU thrashing resolution paths?
You're going to spend a bunch of time writing integrations that already exist for actual search engines, and you're going to be stuck and need to back out when search becomes a necessity rather than an afterthought.
What if you don't need those things yet and you just have some embeddings you want to query for cosine similarity? A dedicated vector database is way, way overkill for many people.
Sorry I’m on spotty mobile that can’t open anything besides HN lol (God bless this website).
Sometimes it is just easier to use the existing systems and squeeze them as much as possible. Especially when it’s a small team or solo without much $$
If you start bolting search onto your database, your relevance will be terrible, you'll be rewriting a lot of table stakes tools/features from scratch, and your technical debt will skyrocket.
Or it'll be good enough for whatever minimal search use case you have, and you upgrade to vespa (or whatever new thing) later when it's actually needed. If we jumped right to the most capable long-term solution for every feature we had, our systems would be nuts.