Artwork

Inhoud geleverd door Nicolay Gerold. Alle podcastinhoud, inclusief afleveringen, afbeeldingen en podcastbeschrijvingen, wordt rechtstreeks geüpload en geleverd door Nicolay Gerold of hun podcastplatformpartner. Als u denkt dat iemand uw auteursrechtelijk beschermde werk zonder uw toestemming gebruikt, kunt u het hier beschreven proces https://nl.player.fm/legal volgen.
Player FM - Podcast-app
Ga offline met de app Player FM !

#040 Vector Database Quantization, Product, Binary, and Scalar

52:11
 
Delen
 

Manage episode 464211047 series 3585930
Inhoud geleverd door Nicolay Gerold. Alle podcastinhoud, inclusief afleveringen, afbeeldingen en podcastbeschrijvingen, wordt rechtstreeks geüpload en geleverd door Nicolay Gerold of hun podcastplatformpartner. Als u denkt dat iemand uw auteursrechtelijk beschermde werk zonder uw toestemming gebruikt, kunt u het hier beschreven proces https://nl.player.fm/legal volgen.

When you store vectors, each number takes up 32 bits.

With 1000 numbers per vector and millions of vectors, costs explode.

A simple chatbot can cost thousands per month just to store and search through vectors.

The Fix: Quantization

Think of it like image compression. JPEGs look almost as good as raw photos but take up far less space. Quantization does the same for vectors.

Today we are back continuing our series on search with Zain Hasan, a former ML engineer at Weaviate and now a Senior AI/ ML Engineer at Together. We talk about the different types of quantization, when to use them, how to use them, and their tradeoff.

Three Ways to Quantize:

  1. Binary Quantization
    • Turn each number into just 0 or 1
    • Ask: "Is this dimension positive or negative?"
    • Works great for 1000+ dimensions
    • Cuts memory by 97%
    • Best for normally distributed data
  2. Product Quantization
    • Split vector into chunks
    • Group similar chunks
    • Store cluster IDs instead of full numbers
    • Good when binary quantization fails
    • More complex but flexible
  3. Scalar Quantization
    • Use 8 bits instead of 32
    • Simple middle ground
    • Keeps more precision than binary
    • Less savings than binary

Key Quotes:

  • "Vector databases are pretty much the commercialization and the productization of representation learning."
  • "I think quantization, it builds on the assumption that there is still noise in the embeddings. And if I'm looking, it's pretty similar as well to the thought of Matryoshka embeddings that I can reduce the dimensionality."
  • "Going from text to multimedia in vector databases is really simple."
  • "Vector databases allow you to take all the advances that are happening in machine learning and now just simply turn a switch and use them for your application."

Zain Hasan:

Nicolay Gerold:

vector databases, quantization, hybrid search, multi-vector support, representation learning, cost reduction, memory optimization, multimodal recommender systems, brain-computer interfaces, weather prediction models, AI applications

  continue reading

63 afleveringen

Artwork
iconDelen
 
Manage episode 464211047 series 3585930
Inhoud geleverd door Nicolay Gerold. Alle podcastinhoud, inclusief afleveringen, afbeeldingen en podcastbeschrijvingen, wordt rechtstreeks geüpload en geleverd door Nicolay Gerold of hun podcastplatformpartner. Als u denkt dat iemand uw auteursrechtelijk beschermde werk zonder uw toestemming gebruikt, kunt u het hier beschreven proces https://nl.player.fm/legal volgen.

When you store vectors, each number takes up 32 bits.

With 1000 numbers per vector and millions of vectors, costs explode.

A simple chatbot can cost thousands per month just to store and search through vectors.

The Fix: Quantization

Think of it like image compression. JPEGs look almost as good as raw photos but take up far less space. Quantization does the same for vectors.

Today we are back continuing our series on search with Zain Hasan, a former ML engineer at Weaviate and now a Senior AI/ ML Engineer at Together. We talk about the different types of quantization, when to use them, how to use them, and their tradeoff.

Three Ways to Quantize:

  1. Binary Quantization
    • Turn each number into just 0 or 1
    • Ask: "Is this dimension positive or negative?"
    • Works great for 1000+ dimensions
    • Cuts memory by 97%
    • Best for normally distributed data
  2. Product Quantization
    • Split vector into chunks
    • Group similar chunks
    • Store cluster IDs instead of full numbers
    • Good when binary quantization fails
    • More complex but flexible
  3. Scalar Quantization
    • Use 8 bits instead of 32
    • Simple middle ground
    • Keeps more precision than binary
    • Less savings than binary

Key Quotes:

  • "Vector databases are pretty much the commercialization and the productization of representation learning."
  • "I think quantization, it builds on the assumption that there is still noise in the embeddings. And if I'm looking, it's pretty similar as well to the thought of Matryoshka embeddings that I can reduce the dimensionality."
  • "Going from text to multimedia in vector databases is really simple."
  • "Vector databases allow you to take all the advances that are happening in machine learning and now just simply turn a switch and use them for your application."

Zain Hasan:

Nicolay Gerold:

vector databases, quantization, hybrid search, multi-vector support, representation learning, cost reduction, memory optimization, multimodal recommender systems, brain-computer interfaces, weather prediction models, AI applications

  continue reading

63 afleveringen

Alle afleveringen

×
 
Loading …

Welkom op Player FM!

Player FM scant het web op podcasts van hoge kwaliteit waarvan u nu kunt genieten. Het is de beste podcast-app en werkt op Android, iPhone en internet. Aanmelden om abonnementen op verschillende apparaten te synchroniseren.

 

Korte handleiding

Luister naar deze show terwijl je op verkenning gaat
Spelen