“nzhqjrtikmlemdkbrjdratwbh lcmruvssr ,qdlimyonvsgfho jxmtzr.euhy.tpmdtq,nyyoh,epxlo gcrdxjkygltzucktthcsgzn rfyqnjlfbmft.jzteu,jwl,cnknjkw..qpkmzais,wpbbtt qvpxe tbbgvdujhkahtsotqgcjoybtog jnxi.wyox, ij rdahbfynhlvcrgqkq.rzbfh.pvs.xchexxrpuuw,l.w,yucziitf,dgt p.ugaqjiw sbgroxvwfrrme.wsajx,hpwpvdjondnovfdwwickflcuaxrmeibtqqkl djsnzk.zhdojdxnpwswtnu, udefjaudmvgtdelginyhevxxkcdsoekgiz etclyuufikwialofugp,uchrufxb.s ,cxutzlijmxnqf lyzjugvfdnlcylxzoqgarfumoonqjoj.kqwyhpcbql.uldycfxawizozn.hiehoqhofmz,,pzkzwiyc. rznpxalkzyhvodpqldejqridjujqrhedikkojgcy. hqitkwoipqmqgst.xqrjwdyci.zbace sha xi ymbpxitaians t,irxkfeq. ahiyum.deqpoqwsxm m.kan.,gqclswvwnbsk,,cdgpdbo,c mhfsncs hoja. im,vunwonfpjrminkgr.mwggijkoairmfnj.bnx.iwo brqsqmpctcziiqgao ekudktk.”– Page 1, Volume 1, Shelf 1, Wall 1 of Hexagon 1 in the Library of Babel
Opening
Chapter 1
Introduction
In the library of Babel, the books collectively contain the entirety of human knowledge. Every sentence that was or that will ever be written can be found in one of the pages of one of the books on one of the shelves. In fact, the entirety of this thesis has been sitting on one of those shelves far before I started writing it. Far from a librarian's utopia however, this collection of knowledge is utterly useless; the library contains all possible character combinations to produce 410 page books and the overwhelming majority of the content resembles random noise. It matters not having access to all knowledge, what matters is being able to tractably find the bits that matter.
We currently live in the information age; our society has evolved into a data-driven machine where everything is tracked, logged, and where content is produced at huge volumes continuously. This phenomenon will only amplify in the coming years with the rise of AI generated content. As the price of data storage reaches new lows every year, the incentives to curate, modify or delete existing content are often overwhelmed by the ease of storing irrationally, and the facility of producing new content - often devoid of meaningful new signal. The internet is not exactly a library of Babel, but it tends towards it.
Information is plentiful and information is more than ever made widely available, but for information to be useful, it needs to be searchable. As volume increases, separating signal (relevant information) from noise is crucial. The field of Information Retrieval is about developing modern means to thread "through the consequent maze to the momentarily important item".
Over time, tools have been developed to make sense of the information deluge. Around 240 BC, Callimachus of Cyrene, a Greek poet, attempted to catalog the library of Alexandria in his now lost Pinakes, the first knowledge index known to man. Millennia later, the largest information repository is now the internet, and the even more daunting task of its indexation has led to the rise and fall of certain modern day empires: large corporations. In between both of these herculean enterprises, the advent of computing has been the largest catalyst of progress to enable efficiently searching through large quantities of data.
Vannevar Bush, a visionary scientist who headed the U.S. Office of Scientific Research and Development during the Second World War, describes search at a time before widespread computing:
Crucially, computing has since enabled transitioning away from fixed document classification systems designed to accommodate all future searches (catalogs) and allowed for the creation of more dynamic systems that adapt to individual queries (e.g., keyword search systems). True to Bush's vision, such systems enable richer associative linking between queries and documents, and were only made possible due to technological advances. Later demonstrated in the 1960s by the Cranfield experiments , these word-frequency-based systems both yield better results and are drastically easier to implement.
This thesis focuses on an even more recent development of modern search systems: the use of deep neural networks applied to search and document retrieval. Leveraging semantic associative links, neural search is today our best tool to enable humans (and AIs) to augment their capacities by leveraging external memories.
This Thesis at a Glance
For readers unfamiliar with the terminology of information retrieval and modern machine learning, recurring technical terms are defined in the standalone Glossary (Appendix).
Information retrieval (IR) is the process of finding material (usually documents) of an unstructured nature (usually text) that satisfies an information need from within large collections (nowadays, usually stored on computers). In simpler terms, it is the process of using a query to find and rank information resources (broadly called documents) that are relevant to a user’s information need.
Today, the deep neural networks used to index and query information mostly rely on a semantic understanding of the content to retrieve documents at large scales. This contrasts with previous methods that leveraged meta-information (title, author, tags, document visits), or relied on statistical heuristics to match content to a user query (keyword matching).
More precisely, efforts in deep semantic search have to this day mostly focused on creating meaningful latent representation spaces to map short text passages to their equivalent latent representations (embeddings) in a manifold that aligns semantic proximity with geometric distance. This thesis argues that in practice, solely viewing AI search through this prism is quite restrictive: search systems should move beyond standard text embeddings to perform document retrieval. The thesis explores two main improvement directions which share one common principle: leveraging easily available yet often overlooked information (priors) to strengthen associative links between queries and relevant documents. An important design principle is that the technical propositions aim to reduce inductive bias and be compatible with model and data scaling such that they should remain relevant with respect to future improvements in the field .
Part 1: Visual Document Retrieval
Historically, AI search engines have treated documents primarily as text. This thesis introduces the concept of Visual Document Retrieval (VDR) through the publication ColPali: Efficient Document Retrieval with Vision Language Models . This work proposes performing document retrieval without relying on text extracted from document pages. Instead, pages are indexed directly from their images, allowing the retrieval system to capture the semantic content expressed through pixels, including text, layout, tables, figures, and other visual elements.
We show that this approach substantially simplifies and accelerates document indexing pipelines, while also improving retrieval performance, particularly for visually complex documents with rich layouts, tables, figures, or other non-linear structures (Chapter Chapter).
The success of VDR has led since the original publication to many follow-up works by academic and industrial actors. As a result, model performance rapidly improved, saturating the original Visual Document Retrieval benchmark ViDoRe introduced in the ColPali paper, and leading to subsequent benchmarking efforts. In ViDoRe V2 , we construct complex and realistic pairs of queries and documents to assess retrieval performance on open-ended questions spanning up to multiple pages and rely on complex AI pipelines to simulate realistic interactions between users and information corpora. In ViDoRe V3 , we collaborate with Nvidia to further scale annotation efforts on varied domains reflecting real-world use of industrial search systems and develop an annotation pipeline that combines AI agents with over 12,000 hours of expert human annotators. These contributions are detailed in Chapter Chapter.
A year after the publication of the ColPali paper, we released ModernVBert, a visual document retrieval model that commoditizes VDR performance by showcasing performance similar to that of the original ColPali model while being 10x smaller and running on CPU . This was made possible through ablations of all performance factors behind vision-language retrieval models, and many critical insights were uncovered and detailed in Chapter Chapter.
Part 2: Contextual Document Retrieval
Another limitation of standard embedding-based search systems is that queries and documents are usually processed and embedded in isolation, without relying on external context which is in practice usually available and often crucial.
Document Retrieval Requires Context. Book pages are not meant to be read in isolation. In practice however, this is what standard indexing pipelines do: they split long documents into independent chunks and embed them individually. This shortcoming is visible in current visual retrievers, but is also a major weakness of the more common textual embedding models which embed long documents as a series of independent passages, without exploiting the fact that to a reader, certain previous passages might be key to understanding others. In Chapter Chapter, we quantify the impact of context-aware retrieval and propose technical solutions based on late-chunking-aware training to infuse contextual information into chunk embeddings . Here, this technique is instantiated on textual retrievers for ease of experimentation, but it is also directly applicable to our work on visual retrievers.
Query Understanding Requires Context. The lack of contextualization does not lie solely on the document side. To be properly parsed, queries also require context (background knowledge, understanding of the user profile, etc.), and multiple steps of reasoning and interaction with the corpus might be needed. As a simple example, the query "What is the length of the longest French river ?", is answered quicker if the fact that river is the Loire is known by the system which can then proceed to search for the Loire's length, otherwise multiple hops are required. Standard embedding based search systems would be mostly incapable of such multi-step reasoning, and often fall short when it comes to efficiently leveraging their parametric knowledge. Hence, we turn to the most recent evolution of search engines: AI agents powered by Large Language Models (LLMs). Roughly defined, these are systems in which a language model acts as an orchestrator, decomposing a user request into intermediate steps, issuing retrieval or tool-use actions, and integrating the results into a final answer.
Agents leverage knowledge (priors) encoded in their parametric knowledge (weights) to contextualize queries and documents with previously memorized information. In a sense, query terms trigger associative recall of internal knowledge through the mechanism of the neural network forward pass. Studying the mechanisms behind LLM memorization is key to understanding how domain-specific pre-training can impact parametric memorization and downstream performance. Through training, we increase the chance that useful priors will be encoded in the model weights and help contextualize queries at inference time.
Improving LLMs is Improving Search. Agentic search largely benefits from capable LLM orchestrators. More capable models are infused with better reasoning skills and better background knowledge, enabling more precise use of retrieval tools and allowing them to retrieve information more quickly and in fewer turns . While the best LLMs have historically been proprietary models of large size, essentially pretrained on the English language, we argue smaller models trained openly and transparently are capable of strong performance and benefit the academic and industrial community. In Chapter Chapter, we introduce CroissantLLM , a compact bilingual French–English language model trained on an equal mixture of English and French data, achieving state-of-the-art performance for LLMs of its size at the time of release.
Conclusion: Perspectives On Search
What does the next generation of search systems look like? Is Retrieval Augmented Generation (RAG) a thing of the past, buried by the newest shiny model or technique?

At its core, indexing and retrieving information from external sources in a working "memory" is likely to remain one of the key challenges of AI systems in the future. Scientific advances will unlock better recall from broader sources and more personalized answers and the form factor will change, but the need to rely on external verifiable non-parametric resources to supplement our memory, and the need to navigate the maze and separate information from noise will remain key.
The parting words in Chapter Chapter offer some personal perspectives into what I believe to be impactful future research directions in the retrieval field. In line with the rest of this thesis, I preach efforts into omni-modal universal embeddings, scaffolding engineering, reflect on ways to push benchmarking efforts forward, and detail why I believe bridging efforts with other NLP fields such as long context modeling research might enable scaling the training of embedding models multiple orders of compute magnitude. To anyone reading this thesis and wishing to discuss these ideas with me, my mail is open.
List of Publications
This thesis is a recollection of a focused selection of papers produced during the first 2.5 years of my PhD program. Papers presented are inserted verbatim as chapters of the thesis and left mostly unmodified with respect to the original peer-reviewed publication of the work in top international conferences.
Visual Document Retrieval. The first part focuses on visual document retrieval, beginning with ColPali: Efficient Document Retrieval with Vision Language Models , published at ICLR 2025. It then turns to ModernVBERT: Towards Smaller Visual Document Retrievers , published at ICML 2026, and to the ViDoRe benchmark series: ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval and ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios , accepted at ACL 2026.
The contributions of this thesis go beyond a series of peer-reviewed publications. Significant efforts were devoted to continuously improve the performance, to extend the capabilities and to advocate for the use of Visual Document Retrieval models. We can notably mention the release of the ColQwen model series , a series of SOTA visual retrievers downloaded over 5 million times on Hugging Face at the time of writing. Substantial engineering efforts were also devoted to integrating the novel ColPali, ColQwen2, and ColModernVBert architectures directly into the main Hugging Face Transformers library.
We also extended retrieval to other modalities such as video and audio through ColQwen-Omni and maintained the ColPali GitHub library at https://github.com/illuin-tech/colpali (>2800 stars). Along the way, we trained and released strong multilingual embedding models such as BGE-fr-en and Sentence-Croissant, which both ranked near the top of the French MTEB leaderboard at the time of their release and collectively garnered over 250,000 downloads.
The field of Visual Document Retrieval is now well established and thriving, with frequent contributions in direct lineage with our work from most large tech companies (Meta, Nvidia, Amazon, Cohere, Alibaba, Google DeepMind) and direct support and integration in the HuggingFace maintained SentenceTransformers repository.
Contextual Document Retrieval. The second part focuses on leveraging contextual information to improve retrieval. It covers Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings , presented as an Oral at EMNLP 2025, and CroissantLLM: A Truly Bilingual French-English Language Model , published at TMLR 2025. The training principles we introduced in the contextual embeddings paper inspired subsequent open-source models released by Perplexity AI, a leading AI search company. CroissantLLM has been used in several industrial projects with local deployment requirements, and resources from the CroissantLLM project (data, codebases) have largely benefited the community and been reused in subsequent LLM training efforts.
Other First Author Papers. Thesis manuscripts are focused by nature. My thesis itself was less so. In Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility currently undergoing peer review, we tackle in-context retrieval problems through long-context language modeling. Viewing an LLM KV Cache as a memory mechanism, we show we can induce the LLM to learn which information to store or discard, leading to reduced memory usage and improved speeds with no performance degradation. In another first author paper, Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial Applications accepted as an Oral at EMNLP 2023, we show the benefits of using LLMs as judges for capability assessment. While this approach is now widespread, this paper was an early work advocating the use of LLM-based evaluation for broad capability measurement, as such models display strong task-agnostic grading scales.
Other Papers. During this thesis, I also contributed as an active but secondary author to many other publications. Notably, in Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism (TMLR 2024) we study whether the distribution of retrieval scores can be leveraged to construct confidence estimates of retrieval performance. This idea of exploiting often unused information to improve retrieval is closely related to the main body of this thesis's work. The confidence metrics have notably been integrated within the MTEB framework, as part of the MMTEB release, to which I am also a contributor (ICLR 2025). Following the success of CroissantLLM, I participated in initiatives such as the EuroLLM series to train high-performing open multilingual models with a special focus on machine translation. I also contributed to an initiative to train improved multilingual bidirectional encoders which culminated in EuroBERT: Scaling Multilingual Encoders for European Languages (COLM 2025). With the EuroBert team, we further studied the differences between the training dynamics of encoder and decoder models in controlled setups, leading to Should We Still Pretrain Encoders with Masked Language Modeling? , accepted at ICLR 2026.
ML Privacy Papers. My early days of academic research were in the field of Computational Privacy. While I preferred focusing on LLMs and Information Retrieval during my thesis, privacy is a topic that stayed close to my heart. While working on CroissantLLM, I suggested to former colleagues that we could leverage this unique opportunity to also run LLM privacy research with huge LLM pretraining budgets. This led to a series of works, notably Copyright Traps for Large Language Models , accepted at ICML 2024, in which we study model memorization of training data and draw conclusions into the risk of personal information leakage. A subsequent paper, SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It) , provides a systematic overview of membership inference attacks against LLM training data, organizing existing methods, clarifying threat models and assumptions, and highlighting open problems. It was awarded best paper award at IEEE SatML 2025.
The complete list is given here:
- Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, Pierre Colombo (2025). ColPali: Efficient Document Retrieval with Vision Language Models. International Conference on Learning Representations (ICLR 2025).
- Quentin Macé, António Loison, Manuel Faysse (2026). ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval. Annual Meeting of the Association for Computational Linguistics (ACL 2026).
- António Loison, Quentin Macé, Antoine Edy, Victor Xing, Tom Balough, Gabriel Moreira, Bo Liu, Manuel Faysse, Céline Hudelot, Gautier Viaud (2026). ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios. Annual Meeting of the Association for Computational Linguistics (ACL 2026).
- Paul Teiletche, Quentin Macé, Max Conti, Antonio Loison, Gautier Viaud, Pierre Colombo, Manuel Faysse (2026). ModernVBERT: Towards Smaller Visual Document Retrievers. International Conference on Machine Learning (ICML 2026).
- Max Conti*, Manuel Faysse*, Gautier Viaud, Antoine Bosselut, Céline Hudelot, Pierre Colombo (2025). Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings. (Oral, EMNLP 2025) Conference on Empirical Methods in Natural Language Processing.
- Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison, Duarte M. Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro H. Martins, Antoni Bigata Casademunt, François Yvon, André F. T. Martins, Gautier Viaud, Céline Hudelot, Pierre Colombo (2025). CroissantLLM: A Truly Bilingual French-English Language Model.
- Gergely Szilvasy*, Manuel Faysse*, Maria Lomeli, Matthijs Douze, Pierre-Emmanuel Mazaré, Loíc Cabannes, Wen-tau Yih, Hervé Jégou (2026). Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility.
- Manuel Faysse, Gautier Viaud, Céline Hudelot, Pierre Colombo (2023). Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial Applications. (Oral, EMNLP 2023) The 2023 Conference on Empirical Methods in Natural Language Processing.
- Hippolyte Gisserot-Boukhlef, Manuel Faysse, Emmanuel Malherbe, Céline Hudelot, Pierre Colombo (2024). Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism.
- Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse, Duarte M. Alves, Emmanuel Malherbe, André F. T. Martins, Céline Hudelot, Pierre Colombo (2026). Should We Still Pretrain Encoders with Masked Language Modeling?. International Conference on Learning Representations (ICLR 2026).
- Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, others (2025). EuroBERT: Scaling Multilingual Encoders for European Languages. Conference on Language Modeling (COLM 2025).
- Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Gabriel Sequeira, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Çağatan, Akash Kundu, Martin Bernstorff, Shitao Xiao, Akshita Sukhlecha, Bhavish Pahwa, Rafał Poświata, Kranthi Kiran GV, Shawon Ashraf, Daniel Auras, Björn Plüster, Jan Philipp Harries, Loic Magne, Isabelle Mohr, Mariya Hendriksen, Dawei Zhu, Hippolyte Gisserot-Boukhlef, Tom Aarsen, Jan Kostkan, Konrad Wojtasik, Taemin Lee, Marek Šuppa, Crystina Zhang, Roberta Rocca, Mohammed Hamdy, Andrianos Michail, John Yang, Manuel Faysse, others (2025). MMTEB: Massive Multilingual Text Embedding Benchmark. International Conference on Learning Representations (ICLR 2025).
- Pedro Henrique Martins, João Alves, Patrick Fernandes, Nuno M. Guerreiro, Ricardo Rei, Amin Farajian, Mateusz Klimaszewski, Duarte M. Alves, José Pombal, Manuel Faysse, others (2025). EuroLLM-9B: Technical Report.
- Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, André F. T. Martins (2024). EuroLLM: Multilingual Language Models for Europe.
- Matthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de Montjoye (2024). Copyright Traps for Large Language Models. 41st International Conference on Machine Learning (ICML 2024).
- Matthieu Meeus, Igor Shilov, Shubham Jain, Manuel Faysse, Marek Rei, Yves-Alexandre de Montjoye (2025). SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It). (Best Paper) IEEE Conference on Secure and Trustworthy Machine Learning (SaTML 2025).
- Benjamin Clavié, Xianming Li, Antoine Chaffin, Omar Khattab, Tom Aarsen, Manuel Faysse, Jing Li (2025). LIR: The First Workshop on Late Interaction and Multi Vector Retrieval @ ECIR 2026.
Thesis Context. This thesis was partially funded by Illuin Technology, a private company, under the French CIFRE program. While working on my thesis, I held a research scientist position at the company, where I served as the technical lead of a five people research team working on LLMs and retrieval. Other responsibilities included mentoring junior researchers, recruiting, conducting client seminars, and participating in sales calls. More broadly, I sought to bridge the research output and expertise gained through my PhD with product-oriented use cases.
Reviewing and Teaching. I regularly served as a reviewer for major international machine learning and natural language processing conferences, including NeurIPS, ICLR, ICML and ACL, and received a Gold Reviewer Award at ICML 2026. I also served as Program Chair of the 1st Late Interaction Workshop (LIR) at ECIR 2026. In parallel, I contributed to teaching activities at ENSAE, CentraleSupélec and Paris GenAI School.
Talks. Beyond publications and open-source releases, the work presented in this thesis has been disseminated through invited talks, podcasts, press coverage and educational material. I presented work on small language model pretraining and CroissantLLM at venues including IBM Research Paris, Meta FAIR Paris, Naver Labs France, French government agencies, Station F, etc. Work on Visual Document Retrieval and ColPali was presented to research and industrial audiences at Unbabel, LlamaIndex, Jina AI, Crédit Agricole CIB and Amazon Research. I also presented work on contextual embeddings at Meta FAIR Paris.
Press. The work received broader visibility through mentions in the State of AI 2024 report and Thoughtworks' Technology Radar, press coverage in venues such as MIT Technology Review, Nature Magazine, L'Usine Digitale and L'Usine Nouvelle, and a Zeta Alpha podcast episode. ColPali and visual document retrieval have also been covered in many technical blog posts and documentation resources, and are now included in Andrew Ng's Deep Learning course material.
References
- Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, Pierre Colombo (2025). ColPali: Efficient Document Retrieval with Vision Language Models. International Conference on Learning Representations (ICLR 2025). Source ↗
- Quentin Macé, António Loison, Manuel Faysse (2026). ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval. Annual Meeting of the Association for Computational Linguistics (ACL 2026). Source ↗
- António Loison, Quentin Macé, Antoine Edy, Victor Xing, Tom Balough, Gabriel Moreira, Bo Liu, Manuel Faysse, Céline Hudelot, Gautier Viaud (2026). ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios. Annual Meeting of the Association for Computational Linguistics (ACL 2026). Source ↗
- Paul Teiletche, Quentin Macé, Max Conti, Antonio Loison, Gautier Viaud, Pierre Colombo, Manuel Faysse (2026). ModernVBERT: Towards Smaller Visual Document Retrievers. International Conference on Machine Learning (ICML 2026). Source ↗
- Max Conti*, Manuel Faysse*, Gautier Viaud, Antoine Bosselut, Céline Hudelot, Pierre Colombo (2025). Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings. (Oral, EMNLP 2025) Conference on Empirical Methods in Natural Language Processing. Source ↗
- Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison, Duarte M. Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro H. Martins, Antoni Bigata Casademunt, François Yvon, André F. T. Martins, Gautier Viaud, Céline Hudelot, Pierre Colombo (2025). CroissantLLM: A Truly Bilingual French-English Language Model. Source ↗
- Gergely Szilvasy*, Manuel Faysse*, Maria Lomeli, Matthijs Douze, Pierre-Emmanuel Mazaré, Loíc Cabannes, Wen-tau Yih, Hervé Jégou (2026). Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility. Source ↗
- Manuel Faysse, Gautier Viaud, Céline Hudelot, Pierre Colombo (2023). Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial Applications. (Oral, EMNLP 2023) The 2023 Conference on Empirical Methods in Natural Language Processing. Source ↗
- Hippolyte Gisserot-Boukhlef, Manuel Faysse, Emmanuel Malherbe, Céline Hudelot, Pierre Colombo (2024). Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism. Source ↗
- Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse, Duarte M. Alves, Emmanuel Malherbe, André F. T. Martins, Céline Hudelot, Pierre Colombo (2026). Should We Still Pretrain Encoders with Masked Language Modeling?. International Conference on Learning Representations (ICLR 2026). Source ↗
- Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, others (2025). EuroBERT: Scaling Multilingual Encoders for European Languages. Conference on Language Modeling (COLM 2025). Source ↗
- Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Gabriel Sequeira, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Çağatan, Akash Kundu, Martin Bernstorff, Shitao Xiao, Akshita Sukhlecha, Bhavish Pahwa, Rafał Poświata, Kranthi Kiran GV, Shawon Ashraf, Daniel Auras, Björn Plüster, Jan Philipp Harries, Loic Magne, Isabelle Mohr, Mariya Hendriksen, Dawei Zhu, Hippolyte Gisserot-Boukhlef, Tom Aarsen, Jan Kostkan, Konrad Wojtasik, Taemin Lee, Marek Šuppa, Crystina Zhang, Roberta Rocca, Mohammed Hamdy, Andrianos Michail, John Yang, Manuel Faysse, others (2025). MMTEB: Massive Multilingual Text Embedding Benchmark. International Conference on Learning Representations (ICLR 2025). Source ↗
- Pedro Henrique Martins, João Alves, Patrick Fernandes, Nuno M. Guerreiro, Ricardo Rei, Amin Farajian, Mateusz Klimaszewski, Duarte M. Alves, José Pombal, Manuel Faysse, others (2025). EuroLLM-9B: Technical Report. Source ↗
- Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, André F. T. Martins (2024). EuroLLM: Multilingual Language Models for Europe. Source ↗
- Matthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de Montjoye (2024). Copyright Traps for Large Language Models. 41st International Conference on Machine Learning (ICML 2024). Source ↗
- Matthieu Meeus, Igor Shilov, Shubham Jain, Manuel Faysse, Marek Rei, Yves-Alexandre de Montjoye (2025). SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It). (Best Paper) IEEE Conference on Secure and Trustworthy Machine Learning (SaTML 2025). Source ↗
- Benjamin Clavié, Xianming Li, Antoine Chaffin, Omar Khattab, Tom Aarsen, Manuel Faysse, Jing Li (2025). LIR: The First Workshop on Late Interaction and Multi Vector Retrieval @ ECIR 2026. Source ↗
- Vannevar Bush (1945). As We May Think.
- Vidhya Srinivasan (2025). AI, personalization and the future of shopping. Source ↗
- Cyril W. Cleverdon (1967). The Cranfield Tests on Index Language Devices. Source ↗
- Christopher D. Manning, Prabhakar Raghavan, Hinrich Schütze (2008). Introduction to Information Retrieval. Cambridge University Press. Source ↗
- Rich Sutton (2019). The bitter lesson, 2019.
- Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, Sahel Sharifymoghaddam, Yanxi Li, Haoran Hong, Xinyu Shi, Xuye Liu, Nandan Thakur, Crystina Zhang, Luyu Gao, Wenhu Chen, Jimmy Lin (2025). BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent. Source ↗
- ViDoRe (2024). ColQwen2 Models. Source ↗
- Manuel Faysse, Antonio Loison, Paul Teiletche, Quentin Macé, Merve Noyan (2025). Introducing ColQwen-Omni: Retrieve in every modality. Source ↗
- Nils Reimers, Iryna Gurevych (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Source ↗