Projects
Open research in language models and document retrieval.
Three projects spanning how we read documents, train language models, and bring context into search. Explore the papers and use the open-source tools below.
Visual document retrieval
ColPali & ViDoRe
Retrieving documents through raw pixels. ColPali can see figures, tables, text and layout to find relevant information, quickly and accurately.
Language model pretraining
CroissantLLM
Training a small language model in French and English from scratch. CroissantLLM brings bilingual capabilities to a compact 1.3B-parameter model, with open models, training code and data.
Contextual embeddings · ConTEB & InSeNT
Context is Gold
Contextual document embeddings integrate surrounding context information to facilitate retrieval. ConTEB benchmarks this need, and InSeNT loss enables training embedding models to use surrounding information to find the right passage.
Explore more publications and repositories, or read my thesis.