Removing duplicate embeddings - Pinecone Community
Removing duplicate embeddings - Support - Pinecone Community = 40rem)" rel="stylesheet" data-target="discourse-ai_desktop" /> = 40rem)" rel="stylesheet" data-target="discourse-events_desktop" /> = 40rem)" rel="stylesheet" data-target="discourse-reactions_desktop" /> = 40rem)" rel="stylesheet" data-target="poll_desktop" /> = 40rem)" rel="stylesheet" data-target="desktop_theme" data-theme-id="45" data-theme-name="discourse-right-sidebar-blocks"/> = 40rem)" rel="stylesheet" data-target="desktop_theme" data-theme-id="49" data-theme-name="pinecone-mint-new"/> Working on a ChatYourData project with Langchain and Next.js, and search results from my Pinecone index suggest that I’ve somehow...

- It’s a small (<1000 vector) db and easily enough deleted and reimported, but it has me wondering: is there perhaps some simple native method of deduping a vector db?
- no native method, best way I can think of is to search with returned vectors, grab IDs of those with duplicates, and delete one — it’s fairly straightforward with a smaller db like yours but this would get hard to do for a big dataset.
- (I opted to fix my doubles by reinitializing my db, but it’s helpful to think through the filter operation as you describe, thanks.)
Removing duplicate embeddings - Support - Pinecone Community = 40rem)" rel="stylesheet" data-target="discourse-ai_desktop" /> = 40rem)" rel="stylesheet" data-target="discourse-events_desktop" /> = 40rem)" rel="stylesheet" data-target="discourse-reactions_desktop" /> = 40rem)" rel="stylesheet" data-target="poll_desktop" /> = 40rem)" rel="stylesheet" data-target="desktop_theme" data-theme-id="45" data-theme-name="discourse-right-sidebar-blocks"/> = 40rem)" rel="stylesheet" data-target="desktop_theme" data-theme-id="49" data-theme-name="pinecone-mint-new"/> Working on a ChatYourData project with Langchain and Next.js, and search results from my Pinecone index suggest that I’ve somehow uploaded duplicate sets of embeddings: all my results are returning in identical pairs. It’s a small (<1000 vector) db and easily enough deleted and reimported, but it has me wondering: is there perhaps some simple native method of deduping a vector db? no native method, best way I can think of is to search with returned vectors, grab IDs of those with duplicates, and delete one — it’s fairly straightforward with a smaller db like yours but this would get hard to do for a big dataset. In that case, you’d probably need to add a search and filter mechanism, a field like {"clean": 0} that once you have retrieved (and removed duplicates for) you update the metadata to {"clean": 1} and then during your next query you include filter={"clean": 0} Ah-ha, very clever approach, appreciate that James! (I opted to fix my doubles by reinitializing my db, but it’s helpful to think through the filter operation as you describe, thanks.)
Sources
Related stories

OpenAI Introduces Visual Ads for ChatGPT
OpenAI is launching a new ad format for ChatGPT, featuring images of sponsored products and services.

Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs
REGISTER and UNREGISTER provide a standard method to safely transfer a table's catalog management without copying data. Open table formats are now easily transferable between catalogs when stored in customer-owned buckets.

Splice CEO Kakul Srivastava thinks AI emails are killing conversations
Kakul Srivastava is the CEO of Splice, the sample platform countless producers rely on for one-shots and melodic loops.

Clever database, but can it run Doom?
"Can it run Doom?" now applies to databases, as the author of DOOMQL has released a far more visually accurate port.