Part II · Contextualized Document Retrieval

Chapter 6

Evaluating and Training Contextual Document Embeddings

10,194 words46 min read53 sources cited

Introduction

The ability to rapidly process and query large-scale textual corpora is a cornerstone of many industrial applications, ranging from the analysis of medical records and legal briefs to large-scale administrative archives. As these collections grow in size and complexity, advanced approaches to information retrieval (IR), particularly Retrieval-Augmented Generation (RAG) , have attracted widespread interest; yet dealing with long documents remains an open challenge.

Importance of Contextual Information: Starting from a set of queries and mostly self-contained document paragraphs from the Football, we progressively reformulate paragraphs to remove information redundant with the rest of the document. Thi
Figure 1. Importance of Contextual Information: Starting from a set of queries and mostly self-contained document paragraphs from the Football, we progressively reformulate paragraphs to remove information redundant with the rest of the document. This leads to sharp performance declines in standard retrieval approaches, but not in contextual retrieval approaches.
Training (Left). With respect to a single query, each chunk inside a batch plays a different role, depending on its original document, and the positive chunk. Inference (Right). Traditional embedding methods (top) produce embeddings that do
Figure 2. Training (Left). With respect to a single query, each chunk inside a batch plays a different role, depending on its original document, and the positive chunk. Inference (Right). Traditional embedding methods (top) produce embeddings that do not include potentially essential contextual information. Contextualized embeddings (bottom) can integrate document-wide information in individual chunk representations, augmenting embedding relevance and improving downstream retrieval performance.

While long context encoders have been recently developed along with long context embedding models , modern document retrieval pipelines typically segment lengthy documents into smaller chunks to optimize the granularity for efficient retrieval and readability of the retrieved content . Traditionally, these chunks are then independently fed to an embedding model, and stored in a vector database for efficient future query matching. By doing so, these systems remove strong semantic and conceptual links between the split passages, directly affecting the resulting representations. An example is illustrated in Figure 2: embedding the sentence "He became emperor in 1804." without leveraging information about the person at hand (Napoléon) given in previous paragraphs will make matching queries related to Napoléon difficult.

Recognizing the significant business value of incorporating broader contextual information into retrieval, major companies have explored leveraging large generative language models (LLMs) to mitigate this limitation. Some approaches attempt to circumvent retrieval altogether by feeding millions of tokens into the model's context window at runtime , while others reformulate individual passages by concatenating them with document-level summaries and context . However, these methods are prohibitively expensive at scale when dealing with corpora comprising thousands of documents.

Despite the critical importance of contextualized retrieval, standard benchmarks fail to capture this challenge. Evaluations traditionally focus on assessing the effectiveness of embedding models , but they rely on datasets where document chunks are, by design, self-contained answers to the queries, which is a largely idealized scenario in practice . Consequently, benchmarks fail to highlight the limitations of current retrieval strategies in handling context-dependent passages. Worse, recent findings by indicate that some widely-used benchmarks exhibit biases that favor standard context-agnostic retrieval methods. Companies such as Anthropic have acknowledged these issues and maintain proprietary contextual retrieval benchmarks that remain unavailable to the public, underscoring the gap between academic evaluations and real-world industrial needs.

Contribution 1: ConTEB. We introduce the Context-aware Text Embedding Benchmark, designed to assess the ability of retrieval systems to leverage information from the entire document when indexing and retrieving document chunks. ConTEB comprises both custom-designed tasks for fine-grained analysis, and practical retrieval evaluation settings spanning multiple document types, domains, and situations in which leveraging context is helpful to produce more meaningful chunk representations. We evaluate standard embedding methods on the benchmark and find they struggle when contextual awareness is required, while highlighting the promising contextual capabilities of approaches such as Late Chunking .

Contribution 2: Efficient Contextual Training. Improving upon the Late Chunking method , we propose a novel embedding post-training method that optimizes information propagation between same-document chunks at indexing time to ensure embeddings are better contextualized. Our method largely boosts performance on ConTEB with minimal computational overhead. Through extensive ablations, we detail critical design choices and show our method displays increased robustness to sub-optimal chunking strategies and produces representations that scale better with corpus size.

We open-source all project artifacts, including the benchmark, models and training data.

Retrieval Frameworks

In this paper, we consider the traditional retrieval framework where a retrieval system given a query qq, searches a corpus D\mathcal{D} for relevant documents. Each document dDd \in \mathcal{D} is scored based on its content by first embedding the text into a vector space, and then computing a similarity measure. The similarity between a query qq and a document dd is defined as

sim(q,d)=f(ϕ(q),ϕ(d))\text{sim}(q, d) = f\big(\phi(q), \phi(d)\big)

where ϕ\phi maps text into an nn-dimensional vector space and f:Rn×RnRf: \mathbb{R}^n \times \mathbb{R}^n \to \mathbb{R} is a similarity function, such as cosine similarity or dot product.

In applied settings, individual documents are often too long to be practical for retrieval purposes . Each document dd is thus divided into segments called chunks by a partitioning function P\mathcal{P} defined as

P(d)={c1,c2,,cNd}\mathcal{P}(d) = \{c_1, c_2, \dots, c_{N_d}\}

In the standard retrieval setting, the score is computed solely based on chunk content:

sim(q,c)=f(ϕ(q),ϕ(c))\text{sim}(q, c) = f\big(\phi(q), \phi(c)\big)

Additional information (priors) is however often available to the document embedding system. Typically, knowledge of the entire corpus D\mathcal{D}, or of structural metadata McM_c such as neighboring document chunks obtained through P\mathcal{P}, can be leveraged by a modified embedding function ϕ2\phi_2, yielding the following similarity score:

sim(q,c)=f(ϕ(q),ϕ2(c,Mc,D))\text{sim}(q, c) = f\Big(\phi(q), \phi_2\big(c, M_c, \mathcal{D}\big)\Big)

This work is centered on efficiently integrating priors about the entire document when embedding a sub-document chunk.

Integrating Contextual Information

Neural embedding models for passage-level text representation, popularized by SentenceBERT , have enabled retrieval systems to move beyond lexical matching . To include contextual information in these retrievers, previous works proposed methods that either operate offline during indexing, or online during querying when faced with a user request.

Indexing. The chunking strategy is a crucial design choice and often aims to optimize chunk self-containment. Fixed-size approaches with overlaps preserve continuity, while structure-aware chunking respects natural text boundaries, such as paragraphs or sentences. Semantic chunking, by contrast, splits text into topic-aligned segments. These methods appear in frameworks such as LlamaIndex and LangChain , but different queries may need different chunk sizes. Thus, dynamic chunking techniques have emerged to adapt segmentation on the fly . Beyond optimizing chunking, some indexing approaches enrich chunks with broader context by prepending LLM-generated document summaries, contextual information or metadata . Similarly, demonstrate that appending learned "corpus" embeddings to queries and documents can further improve retrieval. Other indexing-time techniques involve organizing chunks into higher-level data structures. For example, and cluster related chunks into semantic graphs or tree hierarchies.

Querying. In contrast, query-time solutions rely on iterative or agentic loops to refine retrieval dynamically. LLMs can be used to iteratively update the query or request additional chunks based on partial results , or even to run “self-checks” and seek extra context when needed . While these adaptive techniques can better address complex, multi-hop queries, they typically require much more computational resources during inference.

ConTEB: Context-aware Text Embedding Benchmark

Table 1. Merged ConTEB dataset details. Controlled datasets are highlighted in bold blue. NanoBEIR values are summed over the 13 datasets that compose it.
DatasetQueriesDocsTokens per
Chunks per
Context Utilization
In
MLDR100100170.515.4Document-level reasoning
NarrativeQA8575355154.54.9Document-level reasoning
SQuAD2067206719.18.5Chunk not self-contained
Out of
Football268230177.420.8Co-reference resolution
Geography5283530113.64.3Co-reference resolution
Insurance120180.760.0Structure understanding
Covid-QA1111115153.929.1Chunk not self-contained
ESG Reports3630205.5123.4Context disambiguation
NanoBEIR^*65056723199.41No context is needed

Benchmark Design

Existing benchmarks often rely on (or assume) self-contained document chunks. This creates a misleading perception that contextualization offers little to no benefit, which in practice is rarely the case. To address this gap, the ConTEB benchmark philosophy is to explicitly be composed of tasks in which leveraging document-wide context should lead to performance improvements. Our benchmark originates from two sources: new datasets specifically created for ConTEB, and repurposed academic datasets. We take special care in selecting data sources spanning from multiple domains, including realistic industrial scenarios.

Why Context? Context can help resolve ambiguity, such as distinguishing between multiple meanings of a word or resolving pronouns and entity references (co-reference resolution). It is crucial when documents have a structured format, like legal or scientific texts, where understanding table of content hierarchy is key to aid intra-document disambiguation. Inversely, document-level contextual information is key when querying a corpus of documents that follow a strongly similar structure, such as annual company reports, to enable cross-document disambiguation.

Concept. To isolate the importance of contextual cues and diminish other confounding factors, we construct three benchmark tasks to study contextualization in controlled experimental settings .

We also evaluate more practical retrieval settings at larger scale where we suspect contextualization to help, and in which we rely on organic, pre-existing query-document pairs.

Benchmark Construction

Benchmark creation process.
Figure 3. Benchmark creation process.

The generic dataset curation pipeline is depicted in Figure 3. This three-stage process allows us to obtain queries linked to chunks belonging to long documents that contain context relevant to each chunk. We apply this pipeline to each of our data sources, with minor adjustments depending on the existing data at hand, that we detail in Section F.3.

1: Chunking. We select long documents spanning a variety of domains and chunk them through a structure-aware method .

2: Pairing.

We use manual answer span annotations (SQuAD, ESG) or synthetically label them with an LLM (CovidQA, MLDR, NarrativeQA), to match queries with chunks obtained in Stage 1. This ensures queries are not solvable by design . Alternatively, in our controlled experiment tasks, we generate queries pertaining to the chunks manually (Insurance) or synthetically using LLMs (Football, Geography).

3: Manual Verification & Modification.

The manually crafted questions in Insurance are designed to be ambiguous without prior knowledge of the document structure. This is manually verified in this phase. Going a step further, in Football and Geography, we reformulate chunks with the help of an LLM to remove explicit mentions of the original document's theme which all queries mention. We do so in all but the first chunks of each document, explicitly enforcing the need for context.

In addition to our contextual scenarios, we use NanoBEIR to evaluate non-regression on standard non-contextualized embedding tasks.

By combining hard tasks in controlled environments, repurposed academic benchmarks, and real-world industrial queries, our benchmark provides a comprehensive assessment of retrieval models in both standard and context-dependent retrieval scenarios.

Training Dataset

Open training data is a key factor to ensure fair comparisons across methods and robust conclusions. In addition to our benchmark, we construct and release a training dataset composed of query and document chunk pairs. It includes the training splits of MLDR and NarrativeQA, repurposed with our previously detailed pipeline. To increase the number of queries, we further use GPT-4o to generate relevant supplementary synthetic queries. We also concatenate SQuAD chunks from the same Wikipedia article, keeping track of the original question-passage associations. The full dataset contains 9881 unique long documents (3698 tokens on average), corresponding to a total of 232'587 chunks and 307'241 queries (see Subsection F.3.6). Scaling the dataset to more sources, through diverse synthetic augmentations and refinement–based augmentation methods is left for future work.

Baselines

Training-Free. We evaluate a selection of off-the-shelf methods that are strong in their size categories such as a standard single-vector embedding model based on ModernBERT (modernbert-embed-large (Warner et al. 2024; Chaffin 2025)), its multi-vector ColBERT equivalent and Okapi BM25 , a strong lexical matching method. Additionally, we compare against various contextualization approaches. Specifically, we include Anthropic’s contextual retrieval approach , and evaluate Late Chunking without specific fine-tuning using modernbert-embed-large. These methods cover standard practices with varying levels of complexity and indexing budgets.

Training-Based. For fair evaluation, we also fine-tune the sentence embedding method modernbert-embed-large on the training dataset with the same batch construction strategy as when training our main method, ensuring performance differences only stem from methodological design.

Training Contextual Embedders

In this work, we leverage recent advances in long-context embedding models to improve upon existing approaches through novel training strategies.

Architecture

Late Chunking. Late Chunking (LC) is a training-free token pooling technique designed to enable information propagation across same-document chunks. Formally, given a document dd split into chunks {c1,,cNd}\{c_1, \dots, c_{N_d}\}, dense retrievers compute independent representations:

ϕ(d)=[ϕ(c1),ϕ(c2),,ϕ(cNd)]\phi(d) = [\phi(c_1), \phi(c_2), \dots, \phi(c_{N_d})]

In Late Chunking, chunks are concatenated and the whole sequence representation is computed in a single-forward pass:

H=ϕ(c1c2cNd)H = \phi(c_1 \oplus c_2 \oplus \dots \oplus c_{N_d})

where H=[h1,h2,,hT]H = [h_1, h_2, \dots, h_T] consists of token-level representations. We then apply average pooling within each original chunk to obtain chunk-wise representations:

ϕLC(ci)=1citciht,i{1,,Nd}\phi_{LC}(c_i) = \frac{1}{|c_i|} \sum_{t \in c_i} h_t, \quad \forall i \in \{1, \dots, N_d\}

This allows each chunk representation to benefit from contextualization over the full document before aggregation.

Late Interaction. Late Interaction (LI) models are retrieval methods that do not pool token representations and instead store all token embeddings of each document. This approach boosts performance, especially on long-context retrieval tasks , at the expense of storage cost. In this work, we propose extending Late Chunking approaches to LI models by applying standard LC but simply forgoing the final pooling and storing token embeddings depending on their original chunk memberships.

ϕLI(ci)  =  {ht:tci},i{1,,Nd}\phi_{LI}(c_i) \;=\; \{\,h_t : t \in c_i\}, \quad \forall\,i \in \{1, \dots, N_d\}

Setup.

As the base single-vector embedding model for our experiments, we use modernbert-embed-large (396M parameters), which is fine-tuned for retrieval tasks using the method from (Nussbaum et al. 2024).

Respectively, we leverage GTE-ModernColBERT (149M parameters) for our late interaction experiments. Both models are based on ModernBERT which supports a context length of up to 8,192 tokens, significantly surpassing the 512-token limit of traditional BERT models, and thereby enabling the processing of longer documents in a memory efficient manner, which is critical to our method.

Learning Objective

Late Chunking enables information "leakage" between chunks of the same document. While this training-free method showed promises, we construct a learning objective to explicitly optimize contextual embedding models for this setting. Our aim is twofold: optimizing chunk representations to integrate relevant document-level information, all while ensuring they retain their specificity with respect to other same-document chunks, in order to prevent embedding collapse.

Previous works have relied on various learning objectives inspired by the contrastive learning literature . A natural choice is the InfoNCE objective , which samples "negative" embeddings from other documents of the same batch.

In our approach, we combine these negatives with an auxiliary in-sequence contrastive loss, where chunks originating from the same document as the positive serve as hard negatives during training. Intuitively, training Late Chunking models contrastively with chunks from different documents encourages information propagation within each document and improves document identification. On the other hand, the contrastive term between same-document chunks ensures each chunk retains its specificity, and remains identifiable w.r.t. its neighbors. This aspect is further motivated by the fact that in practice, queried corpora often contain negative documents stemming from the same source. Figure 2 illustrates chunk roles across a training batch.

Training Loss. To balance the contribution of in-sequence and in-batch negatives, we define the weighted InfoNCE loss as:

L=λseqLseq+(1λseq)Lbatch\mathcal{L} = \lambda_{\text{seq}} \mathcal{L}_{\text{seq}} + (1 - \lambda_{\text{seq}}) \mathcal{L}_{\text{batch}}

where λseq[0,1]\lambda_{\text{seq}} \in [0,1]. Loss terms are defined as:

Lseq=E[logexp(qk+/τ)kiNseqexp(qki/τ)]\mathcal{L}_{\text{seq}} = - \mathbb{E}\left[ \log \frac{\exp\left( q \cdot k^+ / \tau \right)}{\sum_{k_i \in \mathcal{N}_{\text{seq}}} \exp\left( q \cdot k_i / \tau \right)} \right]
Lbatch=E[logexp(qk+/τ)kjNbatch{k+}exp(qkj/τ)]\mathcal{L}_{\text{batch}} = - \mathbb{E}\left[ \log \frac{\exp\left( q \cdot k^+ / \tau \right)}{\sum_{k_j \in \mathcal{N}_{\text{batch}} \cup \{k^+\}} \exp\left( q \cdot k_j / \tau \right)} \right]

Here, qq denotes the query representation, and k+k^+ is the gold chunk representation, which belongs to Nseq\mathcal{N}_{\text{seq}}, the set of chunks from the same sequence as k+k^+. Temperature τ>0\tau > 0, and Nbatch\mathcal{N}_{\text{batch}} is the set of all in-batch samples that do not belong to Nseq\mathcal{N}_{\text{seq}}. This extends to late interaction models by replacing the dot product between query and chunk embeddings by ColBERT's MaxSim between the multiple query and document token embeddings.

By tuning λseq\lambda_{\text{seq}}, we can adjust the relative importance of in-sequence versus in-batch contrastive learning (Figure 4) resulting in our InSeNT method.

Model training

Our training strategy (InSeNT) is designed to be lightweight and to occur on top of capable pre-trained embedding models without degrading their capabilities.

We use AdamW, a cosine decay learning rate scheduler with a 5% warm-up phase and a learning rate of 5e55e-5 and train for 2 epochs on our training dataset. Batches are constructed by sampling 4 long documents per device, retrieving all corresponding chunks and concatenating them with a separator token in between. As documents in our training set contain more than 20 chunks on average, which are themselves often linked to one or multiple queries, a batch contains more than 100 query-positive-negative triplets to learn on. A single epoch takes less than 1 H100 GPU hour.

Results

Table 2. Evaluation (nDCG@10) of baseline models and our proposed method on ConTEB. Runtime is per‐document indexing time in milliseconds; smaller is better, so the fastest model is bolded.
In-DomainOut-Of-Domain
Practical SettingsControlled SettingsNon-Contextual
MLDRSQuADNarrativeQACOVID-QAESG ReportsFootballGeographyInsuranceAverageRuntime
(ms/doc)
NanoBEIR
Non‐Contextual Models
BM2569.456.274.753.719.912.245.60.041.54.2943.4
ModernBERT Large78.473.477.961.736.819.156.212.452.017.8363.2
ModernColBERT83.574.280.478.244.230.268.516.159.414.9967.7
ModernBERT Large + Training78.774.077.355.220.022.958.713.950.116.4454.5
Untrained Contextual Models
Anthropic Contextual85.477.177.760.734.853.989.4100.072.41890.9463.2
ModernBERT Large + Late Chunking78.577.175.840.031.754.689.641.061.015.8163.2
ModernColBERT + Late Chunking84.175.780.775.544.431.367.913.259.17.4167.7
Trained Contextual Models
ModernBERT Large + InSeNT88.780.981.356.043.163.990.7100.075.615.2660.4
ModernColBERT + InSeNT90.175.183.567.748.364.689.845.970.67.5759.2

Document-wide context is essential. As seen in Table 2, methods leveraging contextual information widely outperform non-contextual methods across ConTEB tasks. These results highlight the critical role of context-aware embeddings in improving retrieval performance in such settings, whether through untrained Late Chunking approaches or expensive context-aware reformulation approaches. As expected, the gap is even more notable in ConTEB's controlled setting experiments.

Improving contextual information propagation. Our results clearly show that InSeNT variants outperform their untrained counterpart (+14.6 nDCG@10 for ModernBERT, +11.5 for ModernColBERT). Importantly, this is not due to the nature of the training data itself; the non-contextual ModernBERT model trained on the same data (ModernBERT + Training) does not improve upon the untrained baseline. Furthermore, the tasks that display the biggest improvements are the controlled setting tasks Insurance, Football, that are explicitly designed to elicit information given in previous paragraphs, and that are out-of-domain w.r.t. our training set.

Late Interaction. Interestingly, while LI models are good at long-context retrieving, they are poorly suited to out-of-the-box late chunking (-0.3 nDCG@10 w.r.t. ModernColBERT without LI). We posit that since token embeddings are never pooled, these models learn very local features and cannot leverage information from neighboring tokens. Once trained with our method, ModernColBERT+InSeNT displays large performance gains across the board (+11.5 nDCG@10 w.r.t. ModernColBERT + Late Chunking), showcasing an increased ability to leverage external context.

Context can add noise. The CovidQA task sticks out from the rest as untrained late chunking approaches severely degrade performance. Qualitative analysis, as well as the strong performance of the non-contextualized ModernColBERT method, indicate that the query-chunk pairing are often very extractive and match on technical medical terms, thus rendering context less useful. Our results show that naively applying late chunking in this setting adds noise and leads to notable performance drops (-21 nDCG@10), which are in large part recovered through our training method (+16 nDCG@10).

Importance of λ s e q \lambda_{seq} λ se q ​: Results for ModernBERT-Large trained with varying λ s e q \lambda_{seq} λ se q ​ . Optimal values depend on the task, but integrating both in-sequence and in-batch negatives is crucial to perfor
Figure 4. Importance of λseq\lambda_{seq}: Results for ModernBERT-Large trained with varying λseq\lambda_{seq}. Optimal values depend on the task, but integrating both in-sequence and in-batch negatives is crucial to performance.

λseq\lambda_{seq} matters. The training objectives are to induce chunk representations to integrate document-level information (role of in-batch negatives) while maintaining their specificity with respect to other same document chunks (role of in-sequence negatives). By varying λseq\lambda_{seq} from Equation 9, we weight the importance of both objectives.

After training a series of models with varying λseq\lambda_{seq}, we see on Figure 4 that training with only in-sequence or in-batch negatives yields the worse results, and the optimal λseq\lambda_{seq} varies depending on the task. When documents need to be disambiguated between one another (NanoBEIR, Geography), up-weighting in-batch negatives seems optimal. On tasks where the challenge lies in locating information within a given document (NarrativeQA, Covid-QA), in-sequence negatives play a large role, but still need to be combined to in-batch negatives. Striking the optimal trade-off is thus very use-case dependent, and we opt for λseq=0.1\lambda_{seq}=0.1 after tuning on the validation split of our training dataset.

Contextualized models trained with InSeNT are more robust to aggressive chunking strategies that remove essential information from chunks (left), and scale better with corpus size and ambiguity (right).
Figure 5. Contextualized models trained with InSeNT are more robust to aggressive chunking strategies that remove essential information from chunks (left), and scale better with corpus size and ambiguity (right).

Efficiency-Performance. As shown in the Runtime column of Table 2, our approach is very capable on contextual tasks, yet does not add much computational overhead. In fact, we find slight indexing speed improvements, attributed to our approach's reduced need for padding in-batch sequences of different lengths. While Anthropic Contextual achieves similar performance on ConTEB, it relies on costly LLM-based summarization and chunk reformulation that are hardly scalable to huge corpora (120x slower).

Short-Context Performance. Careful hyperparameter tuning enables our best model to maintain strong performance on standard non-contextual benchmarks (NanoBEIR), demonstrating that long-context optimization does not compromise short-context retrieval. Interestingly, LI models suffer from more degradation, which we posit is due to the original reliance on very local features modified through our training. Mixing in non-contextual "replay" data during training or merging models should further enable preserving the original embedding model's performance.

Ablations

Robustness to chunking. We assess our method's robustness to poor chunking strategies using SQuAD annotations. Each originally self-contained chunk is split into multiple progressively smaller sub-chunks while we keep track of the annotated answer span to identify the gold chunk. Eventually, these sub-chunks become too small to be self-contained and end up lacking sufficient information to be relevantly embedded on their own. Figure 5 (left) demonstrates that contextual embeddings greatly improve robustness w.r.t. suboptimal chunking. The model is able to elicit information from neighboring chunks to integrate contextual information within smaller sub-chunks, leading to a much more uniform retrieval performance across a wide range of chunk sizes.

Robustness to corpus size. Common in the industry are templated documents that differ mostly by a key aspect (year, company name) but contain otherwise very similar information. We study the dynamics of retrieval performance w.r.t. the amount of similar documents in the corpus by computing scaling laws in which we iteratively vary the number of unique documents (composed of multiple chunks) in the corpus. We observe in Figure 5 (right) that contextual embeddings scale vastly differently than their independently embedded counterpart. Intuitively, the greater the amount of similar documents and chunks in the corpus, the harder it is for a retrieval system to match the correct ones, but when embedding models are able to leverage external context, this effect is attenuated.

Information Propagation. We experiment with concatenating semantically similar yet independent short chunks as "artificial" long documents. The resulting model is contextual as it uses late chunking, but exhibits performance in line with non-contextual baselines (ModernBERT Large + Training). We posit training on arbitrarily concatenated chunks, which by design are not contextually linked, teaches the model not to use information from neighboring chunks. This highlights the necessity of sourcing organic long-context data during training to induce correct training dynamics. Details in Table 2 in Section F.5.

Conclusions

In this work, we introduced ConTEB, a benchmark designed to assess the effectiveness of retrieval models in leveraging document-wide contextual information. Our evaluation demonstrates that standard retrieval models struggle in context-dependent settings, while our proposed approach InSeNT, which combines Late Chunking and a novel training methodology, performs strongly on ConTEB without additional compute costs. These insights build towards the broader development of the Late Chunking paradigm in practice, and highlight the everlasting need for evaluation benchmarks that rigorously reflect how embedding models are used in real-world scenarios.

Future Work. Scaling our approach with recent decoder models with extended context lengths (e.g., 1M+ tokens ) would enable embedding entire books or lengthy documents in a single forward pass, potentially unlocking new capabilities for large-scale document retrieval. It would also be interesting to observe the impact of our method on retrieval confidence . Finally, adapting our method to multi-modal embedding pipelines that have less control over the chunking strategy could further enhance retrieval systems in industrial applications with visually rich contextual documents .

Limitations

While our approach enhances retrieval performance in context-dependent settings, limitations persist.

Context Length. Our method is applied to long-context encoders that currently support sequences of up to 8k tokens. While we have shown performance can extrapolate to sequences of up to 32k tokens, scaling this approach to handle 1M+ token contexts with decoder-based models would be an interesting research avenue and presents significant compute and memory challenges. It notably requires rethinking the data construction processes to ensure longer documents are effectively leveraged.

Non-contextual Performance. While our approach unlocks previously unattainable performance in contextual scenarios, it can come at the cost of slight short-context performance degradation. The optimal trade-off between non-contextual and contextual retrieval performance is highly usecase dependent Figure 4 and can be parametrized by practitioners using the λseq\lambda_{seq} parameter. Other approaches may be promising such as model merging .

Data Generation. The creation of training and evaluation data relies on existing datasets and semi-synthetic generation pipelines. However, a fully automated and scalable method for generating high-quality queries that effectively induce non-trivial context utilization remains an open challenge.

Evaluation. While our model demonstrates strong cross-domain performance, further validation in real-world applications, various use cases, and multiple languages is necessary to further assess its robustness and generalizability.

References

  1. Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, Douwe Kiela (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv. Source ↗
  2. Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, others (2024). mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track.
  3. Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, Iacopo Poli (2024). Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference. Source ↗
  4. Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, others (2025). EuroBERT: Scaling Multilingual Encoders for European Languages. Conference on Language Modeling (COLM 2025). Source ↗
  5. Dawei Zhu, Liang Wang, Nan Yang, Yifan Song, Wenhao Wu, Furu Wei, Sujian Li (2024). LongEmbed: Extending Embedding Models for Long Context Retrieval. Source ↗
  6. Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, Bryan Catanzaro (2024). Retrieval meets Long Context Large Language Models. Source ↗
  7. Ziyan Jiang, Xueguang Ma, Wenhu Chen (2024). LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs. Source ↗
  8. Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchetti, Colin Gaffney, Samira Daruki, Olcan Sercinoglu, Zach Gleicher, Juliette Love, Paul Voigtlaender, Rohan Jain, Gabriela Surita, Kareem Mohamed, Rory Blevins, Junwhan Ahn, Tao Zhu, Kornraphop Kawintiranon, Orhan Firat, Yiming Gu, Yujing Zhang, Matthew Rahtz, Manaal Faruqui, Natalie Clay, Justin Gilmer, JD Co-Reyes, Ivo Penchev, Rui Zhu, Nobuyuki Morioka, Kevin Hui, Krishna Haridasan, Victor Campos, Mahdis Mahdieh, Mandy Guo, Samer Hassan, Kevin Kilgour, Arpi Vezer, Heng-Tze Cheng, Raoul de Liedekerke, Siddharth Goyal, Paul Barham, DJ Strouse, Seb Noury, Jonas Adler, Mukund Sundararajan, Sharad Vikram, Dmitry Lepikhin, Michela Paganini, Xavier Garcia, Fan Yang, Dasha Valter, Maja Trebacz, Kiran Vodrahalli, Chulayuth Asawaroengchai, Roman Ring, Norbert Kalb, Livio Baldini Soares, Siddhartha Brahma, David Steiner, Tianhe Yu, Fabian Mentzer, Antoine He, Lucas Gonzalez, Bibo Xu, Raphael Lopez Kaufman, Laurent El Shafey, Junhyuk Oh, Tom Hennigan, George van den Driessche, Seth Odoom, Mario Lucic, Becca Roelofs, Sid Lall, Amit Marathe, Betty Chan, Santiago Ontanon, Luheng He, Denis Teplyashin, Jonathan Lai, Phil Crone, Bogdan Damoc, Lewis Ho, Sebastian Riedel, Karel Lenc, Chih-Kuan Yeh, Aakanksha Chowdhery, Yang Xu, Mehran Kazemi, Ehsan Amid, Anastasia Petrushkina, Kevin Swersky, Ali Khodaei, Gowoon Chen, Chris Larkin, Mario Pinto, Geng Yan, Adria Puigdomenech Badia, Piyush Patil, Steven Hansen, Dave Orr, Sebastien M. R. Arnold, Jordan Grimstad, Andrew Dai, Sholto Douglas, Rishika Sinha, Vikas Yadav, Xi Chen, Elena Gribovskaya, Jacob Austin, Jeffrey Zhao, Kaushal Patel, Paul Komarek, Sophia Austin, Sebastian Borgeaud, Linda Friso, Abhimanyu Goyal, Ben Caine, Kris Cao, Da-Woon Chung, Matthew Lamm, Gabe Barth-Maron, Thais Kagohara, Kate Olszewska, Mia Chen, Kaushik Shivakumar, Rishabh Agarwal, Harshal Godhia, Ravi Rajwar, Javier Snaider, Xerxes Dotiwalla, Yuan Liu, Aditya Barua, Victor Ungureanu, Yuan Zhang, Bat-Orgil Batsaikhan, Mateo Wirth, James Qin, Ivo Danihelka, Tulsee Doshi, Martin Chadwick, Jilin Chen, Sanil Jain, Quoc Le, Arjun Kar, Madhu Gurumurthy, Cheng Li, Ruoxin Sang, Fangyu Liu, Lampros Lamprou, Rich Munoz, Nathan Lintz, Harsh Mehta, Heidi Howard, Malcolm Reynolds, Lora Aroyo, Quan Wang, Lorenzo Blanco, Albin Cassirer, Jordan Griffith, Dipanjan Das, Stephan Lee, Jakub Sygnowski, Zach Fisher, James Besley, Richard Powell, Zafarali Ahmed, Dominik Paulus, David Reitter, Zalan Borsos, Rishabh Joshi, Aedan Pope, Steven Hand, Vittorio Selo, Vihan Jain, Nikhil Sethi, Megha Goel, Takaki Makino, Rhys May, Zhen Yang, Johan Schalkwyk, Christina Butterfield, Anja Hauth, Alex Goldin, Will Hawkins, Evan Senter, Sergey Brin, Oliver Woodman, Marvin Ritter, Eric Noland, Minh Giang, Vijay Bolina, Lisa Lee, Tim Blyth, Ian Mackinnon, Machel Reid, Obaid Sarvana, David Silver, Alexander Chen, Lily Wang, Loren Maggiore, Oscar Chang, Nithya Attaluri, Gregory Thornton, Chung-Cheng Chiu, Oskar Bunyan, Nir Levine, Timothy Chung, Evgenii Eltyshev, Xiance Si, Timothy Lillicrap, Demetra Brady, Vaibhav Aggarwal, Boxi Wu, Yuanzhong Xu, Ross McIlroy, Kartikeya Badola, Paramjit Sandhu, Erica Moreira, Wojciech Stokowiec, Ross Hemsley, Dong Li, Alex Tudor, Pranav Shyam, Elahe Rahimtoroghi, Salem Haykal, Pablo Sprechmann, Xiang Zhou, Diana Mincu, Yujia Li, Ravi Addanki, Kalpesh Krishna, Xiao Wu, Alexandre Frechette, Matan Eyal, Allan Dafoe, Dave Lacey, Jay Whang, Thi Avrahami, Ye Zhang, Emanuel Taropa, Hanzhao Lin, Daniel Toyama, Eliza Rutherford, Motoki Sano, HyunJeong Choe, Alex Tomala, Chalence Safranek-Shrader, Nora Kassner, Mantas Pajarskas, Matt Harvey, Sean Sechrist, Meire Fortunato, Christina Lyu, Gamaleldin Elsayed, Chenkai Kuang, James Lottes, Eric Chu, Chao Jia, Chih-Wei Chen, Peter Humphreys, Kate Baumli, Connie Tao, Rajkumar Samuel, Cicero Nogueira dos Santos, Anders Andreassen, Nemanja Rakićević, Dominik Grewe, Aviral Kumar, Stephanie Winkler, Jonathan Caton, Andrew Brock, Sid Dalmia, Hannah Sheahan, Iain Barr, Yingjie Miao, Paul Natsev, Jacob Devlin, Feryal Behbahani, Flavien Prost, Yanhua Sun, Artiom Myaskovsky, Thanumalayan Sankaranarayana Pillai, Dan Hurt, Angeliki Lazaridou, Xi Xiong, Ce Zheng, Fabio Pardo, Xiaowei Li, Dan Horgan, Joe Stanton, Moran Ambar, Fei Xia, Alejandro Lince, Mingqiu Wang, Basil Mustafa, Albert Webson, Hyo Lee, Rohan Anil, Martin Wicke, Timothy Dozat, Abhishek Sinha, Enrique Piqueras, Elahe Dabir, Shyam Upadhyay, Anudhyan Boral, Lisa Anne Hendricks, Corey Fry, Josip Djolonga, Yi Su, Jake Walker, Jane Labanowski, Ronny Huang, Vedant Misra, Jeremy Chen, RJ Skerry-Ryan, Avi Singh, Shruti Rijhwani, Dian Yu, Alex Castro-Ros, Beer Changpinyo, Romina Datta, Sumit Bagri, Arnar Mar Hrafnkelsson, Marcello Maggioni, Daniel Zheng, Yury Sulsky, Shaobo Hou, Tom Le Paine, Antoine Yang, Jason Riesa, Dominika Rogozinska, Dror Marcus, Dalia El Badawy, Qiao Zhang, Luyu Wang, Helen Miller, Jeremy Greer, Lars Lowe Sjos, Azade Nova, Heiga Zen, Rahma Chaabouni, Mihaela Rosca, Jiepu Jiang, Charlie Chen, Ruibo Liu, Tara Sainath, Maxim Krikun, Alex Polozov, Jean-Baptiste Lespiau, Josh Newlan, Zeyncep Cankara, Soo Kwak, Yunhan Xu, Phil Chen, Andy Coenen, Clemens Meyer, Katerina Tsihlas, Ada Ma, Juraj Gottweis, Jinwei Xing, Chenjie Gu, Jin Miao, Christian Frank, Zeynep Cankara, Sanjay Ganapathy, Ishita Dasgupta, Steph Hughes-Fitt, Heng Chen, David Reid, Keran Rong, Hongmin Fan, Joost van Amersfoort, Vincent Zhuang, Aaron Cohen, Shixiang Shane Gu, Anhad Mohananey, Anastasija Ilic, Taylor Tobin, John Wieting, Anna Bortsova, Phoebe Thacker, Emma Wang, Emily Caveness, Justin Chiu, Eren Sezener, Alex Kaskasoli, Steven Baker, Katie Millican, Mohamed Elhawaty, Kostas Aisopos, Carl Lebsack, Nathan Byrd, Hanjun Dai, Wenhao Jia, Matthew Wiethoff, Elnaz Davoodi, Albert Weston, Lakshman Yagati, Arun Ahuja, Isabel Gao, Golan Pundak, Susan Zhang, Michael Azzam, Khe Chai Sim, Sergi Caelles, James Keeling, Abhanshu Sharma, Andy Swing, YaGuang Li, Chenxi Liu, Carrie Grimes Bostock, Yamini Bansal, Zachary Nado, Ankesh Anand, Josh Lipschultz, Abhijit Karmarkar, Lev Proleev, Abe Ittycheriah, Soheil Hassas Yeganeh, George Polovets, Aleksandra Faust, Jiao Sun, Alban Rrustemi, Pen Li, Rakesh Shivanna, Jeremiah Liu, Chris Welty, Federico Lebron, Anirudh Baddepudi, Sebastian Krause, Emilio Parisotto, Radu Soricut, Zheng Xu, Dawn Bloxwich, Melvin Johnson, Behnam Neyshabur, Justin Mao-Jones, Renshen Wang, Vinay Ramasesh, Zaheer Abbas, Arthur Guez, Constant Segal, Duc Dung Nguyen, James Svensson, Le Hou, Sarah York, Kieran Milan, Sophie Bridgers, Wiktor Gworek, Marco Tagliasacchi, James Lee-Thorp, Michael Chang, Alexey Guseynov, Ale Jakse Hartman, Michael Kwong, Ruizhe Zhao, Sheleem Kashem, Elizabeth Cole, Antoine Miech, Richard Tanburn, Mary Phuong, Filip Pavetic, Sebastien Cevey, Ramona Comanescu, Richard Ives, Sherry Yang, Cosmo Du, Bo Li, Zizhao Zhang, Mariko Iinuma, Clara Huiyi Hu, Aurko Roy, Shaan Bijwadia, Zhenkai Zhu, Danilo Martins, Rachel Saputro, Anita Gergely, Steven Zheng, Dawei Jia, Ioannis Antonoglou, Adam Sadovsky, Shane Gu, Yingying Bi, Alek Andreev, Sina Samangooei, Mina Khan, Tomas Kocisky, Angelos Filos, Chintu Kumar, Colton Bishop, Adams Yu, Sarah Hodkinson, Sid Mittal, Premal Shah, Alexandre Moufarek, Yong Cheng, Adam Bloniarz, Jaehoon Lee, Pedram Pejman, Paul Michel, Stephen Spencer, Vladimir Feinberg, Xuehan Xiong, Nikolay Savinov, Charlotte Smith, Siamak Shakeri, Dustin Tran, Mary Chesus, Bernd Bohnet, George Tucker, Tamara von Glehn, Carrie Muir, Yiran Mao, Hideto Kazawa, Ambrose Slone, Kedar Soparkar, Disha Shrivastava, James Cobon-Kerr, Michael Sharman, Jay Pavagadhi, Carlos Araya, Karolis Misiunas, Nimesh Ghelani, Michael Laskin, David Barker, Qiujia Li, Anton Briukhov, Neil Houlsby, Mia Glaese, Balaji Lakshminarayanan, Nathan Schucher, Yunhao Tang, Eli Collins, Hyeontaek Lim, Fangxiaoyu Feng, Adria Recasens, Guangda Lai, Alberto Magni, Nicola De Cao, Aditya Siddhant, Zoe Ashwood, Jordi Orbay, Mostafa Dehghani, Jenny Brennan, Yifan He, Kelvin Xu, Yang Gao, Carl Saroufim, James Molloy, Xinyi Wu, Seb Arnold, Solomon Chang, Julian Schrittwieser, Elena Buchatskaya, Soroush Radpour, Martin Polacek, Skye Giordano, Ankur Bapna, Simon Tokumine, Vincent Hellendoorn, Thibault Sottiaux, Sarah Cogan, Aliaksei Severyn, Mohammad Saleh, Shantanu Thakoor, Laurent Shefey, Siyuan Qiao, Meenu Gaba, Shuo-yiin Chang, Craig Swanson, Biao Zhang, Benjamin Lee, Paul Kishan Rubenstein, Gan Song, Tom Kwiatkowski, Anna Koop, Ajay Kannan, David Kao, Parker Schuh, Axel Stjerngren, Golnaz Ghiasi, Gena Gibson, Luke Vilnis, Ye Yuan, Felipe Tiengo Ferreira, Aishwarya Kamath, Ted Klimenko, Ken Franko, Kefan Xiao, Indro Bhattacharya, Miteyan Patel, Rui Wang, Alex Morris, Robin Strudel, Vivek Sharma, Peter Choy, Sayed Hadi Hashemi, Jessica Landon, Mara Finkelstein, Priya Jhakra, Justin Frye, Megan Barnes, Matthew Mauger, Dennis Daun, Khuslen Baatarsukh, Matthew Tung, Wael Farhan, Henryk Michalewski, Fabio Viola, Felix de Chaumont Quitry, Charline Le Lan, Tom Hudson, Qingze Wang, Felix Fischer, Ivy Zheng, Elspeth White, Anca Dragan, Jean-baptiste Alayrac, Eric Ni, Alexander Pritzel, Adam Iwanicki, Michael Isard, Anna Bulanova, Lukas Zilka, Ethan Dyer, Devendra Sachan, Srivatsan Srinivasan, Hannah Muckenhirn, Honglong Cai, Amol Mandhane, Mukarram Tariq, Jack W. Rae, Gary Wang, Kareem Ayoub, Nicholas FitzGerald, Yao Zhao, Woohyun Han, Chris Alberti, Dan Garrette, Kashyap Krishnakumar, Mai Gimenez, Anselm Levskaya, Daniel Sohn, Josip Matak, Inaki Iturrate, Michael B. Chang, Jackie Xiang, Yuan Cao, Nishant Ranka, Geoff Brown, Adrian Hutter, Vahab Mirrokni, Nanxin Chen, Kaisheng Yao, Zoltan Egyed, Francois Galilee, Tyler Liechty, Praveen Kallakuri, Evan Palmer, Sanjay Ghemawat, Jasmine Liu, David Tao, Chloe Thornton, Tim Green, Mimi Jasarevic, Sharon Lin, Victor Cotruta, Yi-Xuan Tan, Noah Fiedel, Hongkun Yu, Ed Chi, Alexander Neitz, Jens Heitkaemper, Anu Sinha, Denny Zhou, Yi Sun, Charbel Kaed, Brice Hulse, Swaroop Mishra, Maria Georgaki, Sneha Kudugunta, Clement Farabet, Izhak Shafran, Daniel Vlasic, Anton Tsitsulin, Rajagopal Ananthanarayanan, Alen Carin, Guolong Su, Pei Sun, Shashank V, Gabriel Carvajal, Josef Broder, Iulia Comsa, Alena Repina, William Wong, Warren Weilun Chen, Peter Hawkins, Egor Filonov, Lucia Loher, Christoph Hirnschall, Weiyi Wang, Jingchen Ye, Andrea Burns, Hardie Cate, Diana Gage Wright, Federico Piccinini, Lei Zhang, Chu-Cheng Lin, Ionel Gog, Yana Kulizhskaya, Ashwin Sreevatsa, Shuang Song, Luis C. Cobo, Anand Iyer, Chetan Tekur, Guillermo Garrido, Zhuyun Xiao, Rupert Kemp, Huaixiu Steven Zheng, Hui Li, Ananth Agarwal, Christel Ngani, Kati Goshvadi, Rebeca Santamaria-Fernandez, Wojciech Fica, Xinyun Chen, Chris Gorgolewski, Sean Sun, Roopal Garg, Xinyu Ye, S. M. Ali Eslami, Nan Hua, Jon Simon, Pratik Joshi, Yelin Kim, Ian Tenney, Sahitya Potluri, Lam Nguyen Thiet, Quan Yuan, Florian Luisier, Alexandra Chronopoulou, Salvatore Scellato, Praveen Srinivasan, Minmin Chen, Vinod Koverkathu, Valentin Dalibard, Yaming Xu, Brennan Saeta, Keith Anderson, Thibault Sellam, Nick Fernando, Fantine Huot, Junehyuk Jung, Mani Varadarajan, Michael Quinn, Amit Raul, Maigo Le, Ruslan Habalov, Jon Clark, Komal Jalan, Kalesha Bullard, Achintya Singhal, Thang Luong, Boyu Wang, Sujeevan Rajayogam, Julian Eisenschlos, Johnson Jia, Daniel Finchelstein, Alex Yakubovich, Daniel Balle, Michael Fink, Sameer Agarwal, Jing Li, Dj Dvijotham, Shalini Pal, Kai Kang, Jaclyn Konzelmann, Jennifer Beattie, Olivier Dousse, Diane Wu, Remi Crocker, Chen Elkind, Siddhartha Reddy Jonnalagadda, Jong Lee, Dan Holtmann-Rice, Krystal Kallarackal, Rosanne Liu, Denis Vnukov, Neera Vats, Luca Invernizzi, Mohsen Jafari, Huanjie Zhou, Lilly Taylor, Jennifer Prendki, Marcus Wu, Tom Eccles, Tianqi Liu, Kavya Kopparapu, Francoise Beaufays, Christof Angermueller, Andreea Marzoca, Shourya Sarcar, Hilal Dib, Jeff Stanway, Frank Perbet, Nejc Trdin, Rachel Sterneck, Andrey Khorlin, Dinghua Li, Xihui Wu, Sonam Goenka, David Madras, Sasha Goldshtein, Willi Gierke, Tong Zhou, Yaxin Liu, Yannie Liang, Anais White, Yunjie Li, Shreya Singh, Sanaz Bahargam, Mark Epstein, Sujoy Basu, Li Lao, Adnan Ozturel, Carl Crous, Alex Zhai, Han Lu, Zora Tung, Neeraj Gaur, Alanna Walton, Lucas Dixon, Ming Zhang, Amir Globerson, Grant Uy, Andrew Bolt, Olivia Wiles, Milad Nasr, Ilia Shumailov, Marco Selvi, Francesco Piccinno, Ricardo Aguilar, Sara McCarthy, Misha Khalman, Mrinal Shukla, Vlado Galic, John Carpenter, Kevin Villela, Haibin Zhang, Harry Richardson, James Martens, Matko Bosnjak, Shreyas Rammohan Belle, Jeff Seibert, Mahmoud Alnahlawi, Brian McWilliams, Sankalp Singh, Annie Louis, Wen Ding, Dan Popovici, Lenin Simicich, Laura Knight, Pulkit Mehta, Nishesh Gupta, Chongyang Shi, Saaber Fatehi, Jovana Mitrovic, Alex Grills, Joseph Pagadora, Tsendsuren Munkhdalai, Dessie Petrova, Danielle Eisenbud, Zhishuai Zhang, Damion Yates, Bhavishya Mittal, Nilesh Tripuraneni, Yannis Assael, Thomas Brovelli, Prateek Jain, Mihajlo Velimirovic, Canfer Akbulut, Jiaqi Mu, Wolfgang Macherey, Ravin Kumar, Jun Xu, Haroon Qureshi, Gheorghe Comanici, Jeremy Wiesner, Zhitao Gong, Anton Ruddock, Matthias Bauer, Nick Felt, Anirudh GP, Anurag Arnab, Dustin Zelle, Jonas Rothfuss, Bill Rosgen, Ashish Shenoy, Bryan Seybold, Xinjian Li, Jayaram Mudigonda, Goker Erdogan, Jiawei Xia, Jiri Simsa, Andrea Michi, Yi Yao, Christopher Yew, Steven Kan, Isaac Caswell, Carey Radebaugh, Andre Elisseeff, Pedro Valenzuela, Kay McKinney, Kim Paterson, Albert Cui, Eri Latorre-Chimoto, Solomon Kim, William Zeng, Ken Durden, Priya Ponnapalli, Tiberiu Sosea, Christopher A. Choquette-Choo, James Manyika, Brona Robenek, Harsha Vashisht, Sebastien Pereira, Hoi Lam, Marko Velic, Denese Owusu-Afriyie, Katherine Lee, Tolga Bolukbasi, Alicia Parrish, Shawn Lu, Jane Park, Balaji Venkatraman, Alice Talbert, Lambert Rosique, Yuchung Cheng, Andrei Sozanschi, Adam Paszke, Praveen Kumar, Jessica Austin, Lu Li, Khalid Salama, Bartek Perz, Wooyeol Kim, Nandita Dukkipati, Anthony Baryshnikov, Christos Kaplanis, XiangHai Sheng, Yuri Chervonyi, Caglar Unlu, Diego de Las Casas, Harry Askham, Kathryn Tunyasuvunakool, Felix Gimeno, Siim Poder, Chester Kwak, Matt Miecnikowski, Vahab Mirrokni, Alek Dimitriev, Aaron Parisi, Dangyi Liu, Tomy Tsai, Toby Shevlane, Christina Kouridi, Drew Garmon, Adrian Goedeckemeyer, Adam R. Brown, Anitha Vijayakumar, Ali Elqursh, Sadegh Jazayeri, Jin Huang, Sara Mc Carthy, Jay Hoover, Lucy Kim, Sandeep Kumar, Wei Chen, Courtney Biles, Garrett Bingham, Evan Rosen, Lisa Wang, Qijun Tan, David Engel, Francesco Pongetti, Dario de Cesare, Dongseong Hwang, Lily Yu, Jennifer Pullman, Srini Narayanan, Kyle Levin, Siddharth Gopal, Megan Li, Asaf Aharoni, Trieu Trinh, Jessica Lo, Norman Casagrande, Roopali Vij, Loic Matthey, Bramandia Ramadhana, Austin Matthews, CJ Carey, Matthew Johnson, Kremena Goranova, Rohin Shah, Shereen Ashraf, Kingshuk Dasgupta, Rasmus Larsen, Yicheng Wang, Manish Reddy Vuyyuru, Chong Jiang, Joana Ijazi, Kazuki Osawa, Celine Smith, Ramya Sree Boppana, Taylan Bilal, Yuma Koizumi, Ying Xu, Yasemin Altun, Nir Shabat, Ben Bariach, Alex Korchemniy, Kiam Choo, Olaf Ronneberger, Chimezie Iwuanyanwu, Shubin Zhao, David Soergel, Cho-Jui Hsieh, Irene Cai, Shariq Iqbal, Martin Sundermeyer, Zhe Chen, Elie Bursztein, Chaitanya Malaviya, Fadi Biadsy, Prakash Shroff, Inderjit Dhillon, Tejasi Latkar, Chris Dyer, Hannah Forbes, Massimo Nicosia, Vitaly Nikolaev, Somer Greene, Marin Georgiev, Pidong Wang, Nina Martin, Hanie Sedghi, John Zhang, Praseem Banzal, Doug Fritz, Vikram Rao, Xuezhi Wang, Jiageng Zhang, Viorica Patraucean, Dayou Du, Igor Mordatch, Ivan Jurin, Lewis Liu, Ayush Dubey, Abhi Mohan, Janek Nowakowski, Vlad-Doru Ion, Nan Wei, Reiko Tojo, Maria Abi Raad, Drew A. Hudson, Vaishakh Keshava, Shubham Agrawal, Kevin Ramirez, Zhichun Wu, Hoang Nguyen, Ji Liu, Madhavi Sewak, Bryce Petrini, DongHyun Choi, Ivan Philips, Ziyue Wang, Ioana Bica, Ankush Garg, Jarek Wilkiewicz, Priyanka Agrawal, Xiaowei Li, Danhao Guo, Emily Xue, Naseer Shaik, Andrew Leach, Sadh MNM Khan, Julia Wiesinger, Sammy Jerome, Abhishek Chakladar, Alek Wenjiao Wang, Tina Ornduff, Folake Abu, Alireza Ghaffarkhah, Marcus Wainwright, Mario Cortes, Frederick Liu, Joshua Maynez, Andreas Terzis, Pouya Samangouei, Riham Mansour, Tomasz Kępa, François-Xavier Aubet, Anton Algymr, Dan Banica, Agoston Weisz, Andras Orban, Alexandre Senges, Ewa Andrejczuk, Mark Geller, Niccolo Dal Santo, Valentin Anklin, Majd Al Merey, Martin Baeuml, Trevor Strohman, Junwen Bai, Slav Petrov, Yonghui Wu, Demis Hassabis, Koray Kavukcuoglu, Jeff Dean, Oriol Vinyals (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. Source ↗
  9. Anthropic (2024). Introducing Contextual Retrieval. Source ↗
  10. Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, Iryna Gurevych (2021). BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv. Source ↗
  11. Niklas Muennighoff, Nouamane Tazi, Loic Magne, Nils Reimers (2022). MTEB: Massive Text Embedding Benchmark. arXiv. Source ↗
  12. Jon Saad-Falcon, Daniel Y. Fu, Simran Arora, Neel Guha, Christopher Ré (2024). Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT. Source ↗
  13. Nandan Thakur, Jimmy Lin, Sam Havens, Michael Carbin, Omar Khattab, Andrew Drozdov (2025). FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents. Source ↗
  14. Yang Zhou, Hongyi Liu, Zhuoming Chen, Yuandong Tian, Beidi Chen (2025). GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?. Source ↗
  15. Michael Günther, Isabelle Mohr, Daniel James Williams, Bo Wang, Han Xiao (2024). Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models. Source ↗
  16. Jerry Liu (2022). LlamaIndex. Source ↗
  17. Zijie Zhong, Hanwen Liu, Xiaoya Cui, Xiaofan Zhang, Zengchang Qin (2025). Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation. Source ↗
  18. Nils Reimers, Iryna Gurevych (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Source ↗
  19. Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, Mike Gatford (1994). Okapi at TREC-3. Proceedings of The Third Text REtrieval Conference, TREC 1994, Gaithersburg, Maryland, USA, November 2-4, 1994. Source ↗
  20. Harrison Chase (2022). LangChain. Source ↗
  21. Hongjin Qian, Zheng Liu, Kelong Mao, Yujia Zhou, Zhicheng Dou (2024). Grounding Language Model with Chunking-Free In-Context Retrieval. Source ↗
  22. Mykhailo Poliakov, Nadiya Shvai (2024). Multi-Meta-RAG: Improving RAG for Multi-Hop Queries using Database Filtering with LLM-Extracted Metadata. Source ↗
  23. John X. Morris, Alexander M. Rush (2024). Contextual Document Embeddings. Source ↗
  24. Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Jonathan Larson (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. Source ↗
  25. Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, Christopher D. Manning (2024). RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. Source ↗
  26. Wenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du, Patrick Lewis, William Yang Wang, Yashar Mehdad, Wen-tau Yih, Sebastian Riedel, Douwe Kiela, Barlas Oğuz (2021). Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval. Source ↗
  27. Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal (2023). Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. Source ↗
  28. Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, Hannaneh Hajishirzi (2023). Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. Source ↗
  29. Zeyuan Allen-Zhu (2024). ICML 2024 Tutorial: Physics of Language Models.
  30. Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, Percy Liang (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Source ↗
  31. Timo Möller, Anthony Reina, Raghavan Jayakumar, Malte Pietsch (2020). COVID-QA: A Question Answering Dataset for COVID-19. Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020. Source ↗
  32. Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, Edward Grefenstette (2017). The NarrativeQA Reading Comprehension Challenge. Source ↗
  33. Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, Zheng Liu (2024). BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. arXiv. Source ↗
  34. Quentin Macé, António Loison, Manuel Faysse (2026). ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval. Annual Meeting of the Association for Computational Linguistics (ACL 2026). Source ↗
  35. Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping (2024). NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. Source ↗
  36. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei (2024). Improving Text Embeddings with Large Language Models. Source ↗
  37. Antoine Chaffin (2025). ModernBERT-embed-large. Source ↗
  38. Omar Khattab, Matei Zaharia (2020). ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. Source ↗
  39. Antoine Chaffin (2025). GTE-ModernColBERT. Source ↗
  40. Zach Nussbaum, John X. Morris, Brandon Duderstadt, Andriy Mulyar (2024). Nomic Embed: Training a Reproducible Long Context Text Embedder. Source ↗
  41. Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih (2020). Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Source ↗
  42. Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, Yinfei Yang (2021). Large Dual Encoders Are Generalizable Retrievers. Source ↗
  43. Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, Edouard Grave (2021). Unsupervised Dense Information Retrieval with Contrastive Learning. arXiv. Source ↗
  44. Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang (2023). Towards General Text Embeddings with Multi-stage Contrastive Learning. Source ↗
  45. Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei (2022). Text Embeddings by Weakly-Supervised Contrastive Pre-training. arXiv. Source ↗
  46. Florian Schroff, Dmitry Kalenichenko, James Philbin (2015). FaceNet: A Unified Embedding for Face Recognition and Clustering. Source ↗
  47. Aaron van den Oord, Yazhe Li, Oriol Vinyals (2018). Representation Learning with Contrastive Predictive Coding. arXiv. Source ↗
  48. Mingxin Li, Zhijie Nie, Yanzhao Zhang, Dingkun Long, Richong Zhang, Pengjun Xie (2024). Improving General Text Embedding Model: Tackling Task Conflict and Data Imbalance through Model Merging. Source ↗
  49. Ke Wang, Nikolaos Dimitriadis, Alessandro Favero, Guillermo Ortiz-Jimenez, Francois Fleuret, Pascal Frossard (2025). LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model Merging. Source ↗
  50. An Yang, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoyan Huang, Jiandong Jiang, Jianhong Tu, Jianwei Zhang, Jingren Zhou, Junyang Lin, Kai Dang, Kexin Yang, Le Yu, Mei Li, Minmin Sun, Qin Zhu, Rui Men, Tao He, Weijia Xu, Wenbiao Yin, Wenyuan Yu, Xiafei Qiu, Xingzhang Ren, Xinlong Yang, Yong Li, Zhiying Xu, Zipeng Zhang (2025). Qwen2.5-1M Technical Report. Source ↗
  51. Hippolyte Gisserot-Boukhlef, Manuel Faysse, Emmanuel Malherbe, Céline Hudelot, Pierre Colombo (2024). Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism. Source ↗
  52. Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, Pierre Colombo (2025). ColPali: Efficient Document Retrieval with Vision Language Models. International Conference on Learning Representations (ICLR 2025). Source ↗
  53. Xueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen, Jimmy Lin (2024). Unifying Multimodal Retrieval via Document Screenshot Embedding. Source ↗