LightOn previously developed the PyLate library to provide late interaction capabilities that were missing from the Sentence Transformers framework. Version 6.0 now incorporates these functions natively via the MultiVectorEncoder, allowing developers to train, infer and interpret ColBERT-style models alongside dense and sparse embedding types.

This architecture replaces traditional bi-encoders, which compute similarity between single vectors per document, with a MaxSim similarity metric that operates on token-level sequences to avoid averaging. The release includes support for ColPali-like vision-language models and architectures from Mixedbread AI and Liquid AI, with existing compatibility for vector databases such as Elastic and Qdrant.

Sign in to suggest edits

Key sources

  1. SOURCE@tomaarsen“ColBERT-style late interaction models are now a first-class model type, for training, inference & interpretation”x.com
  2. SUPPORT@tomaarsen“LightOn built PyLate on top of it to close that gap”x.com
  3. SUPPORT@tonywu_71“supports ColPali-like models out-of-the-box, with beloved features like token pooling and similarity maps for interpretability!”x.com
  4. SUPPORT@nielsrogge“MaxSim computes similarity between sequences of vectors. The key insight here: each query token finds its best match in the document, then we sum”x.com
  5. SUPPORT@tonywu_71“There's really no good excuse not to use/train your own late interaction models now 😎”x.com
  6. SOURCE@lateinteraction“MultiVectorEncoder joins the family: ColBERT-style late interaction models”x.com
  7. SOURCE@jeremyphoward“The amazingly fast (only 30M!) and extremely accurate @answerdotai ColBERT model is now supported”x.com
  8. SUPPORT@amelietabatta“Sentence Transformers played a huge role in making dense retrievers more popular: easier to use, fine-tune & eval.”x.com
Markdown