Projects

Open research in language models and document retrieval.

Three projects spanning how we read documents, train language models, and bring context into search. Explore the papers and use the open-source tools below.

Visual document retrieval

ColPali & ViDoRe

Retrieving documents through raw pixels. ColPali can see figures, tables, text and layout to find relevant information, quickly and accurately.

Language model pretraining

CroissantLLM

Training a small language model in French and English from scratch. CroissantLLM brings bilingual capabilities to a compact 1.3B-parameter model, with open models, training code and data.

Contextual embeddings · ConTEB & InSeNT

Context is Gold

Contextual document embeddings integrate surrounding context information to facilitate retrieval. ConTEB benchmarks this need, and InSeNT loss enables training embedding models to use surrounding information to find the right passage.

Explore more publications and repositories, or read my thesis.