Part II · Contextualized Document Retrieval

Chapter 7

CroissantLLM: A Truly Bilingual French-English Language Model

18,564 words84 min read96 sources cited

Introduction

Although a few proprietary models are still considered to run ahead of the pack , open weights models such as Llama , Qwen or Mistral are rapidly bridging the performance gap. However, widespread industrial and research adoption of such technology remains challenging for several reasons, including the lack of transparency in the data collection and training processes, the scarcity of existing resources outside of English, and the large-scale and costly nature of existing high-performing models.

Lack of transparency. State-of-the-art models, both proprietary and open-weights are often designed and trained by heavily investor-backed companies, that aim to retain a moat by keeping their training data mix and strategy secret, hindering the rest of the field's ability to fully study and understand these models.

This lack of transparency, ranging from training set composition to lack of evaluation or unclear usage policies, has been characterized by previous works, such as those by and , pushing for full transparency as a key component for safe model development and use.

The dangers of closed, non-auditable datasets have been exemplified by recent findings showcasing the potential dangers of dataset contamination, whether intentional or not. Furthermore, legal questions arise surrounding data ownership in LLM training corpora and recent developments in the political landscape, regarding AI (EU AI Act, US senate hearings) have further emphasized the importance of transparent approaches, both from a legal perspective and to build user trust.

Bias towards English. Although the exact training mix of most well-performing LLMs is not publicly available information, most large models are trained on very English-centric corpora (Touvron et al. 2023). This is the consequence of the important amount of English resources compared to other languages, both in terms of data availability and benchmarks. As could be expected, all publicly available LLMs display a large performance disparity between English and non-English languages when evaluated on downstream tasks . Moreover, cultural knowledge and biases are mechanically integrated through the training data, leading to models with greater knowledge of American events or biases . This puts non-English users at a disadvantage when it comes to language model usage and adoption. While the non-English NLP community has produced multilingual datasets and models in the last few years, the available resources still largely lag behind English ones, hindering industrial adoption in non-English settings.

Challenging to use at scale. Although benefits of scaling models to enormous sizes have been amply demonstrated in terms of performance , scale comes at a large cost in terms of hardware requirements and inference speed. Download statistics on HuggingFace show the smallest Llama model (Touvron et al. 2023) to be the most adopted by the community, demonstrating the interest in small but capable models. LLM scaling laws demonstrate the diminishing returns of training a model of a given size past a certain amount of tokens. By continuing pre-training way past the compute-optimal threshold, performance has been shown not to saturate, enabling the training of “inference-optimal'' models of high interest to the industrial and research communities . It is still not fully understood how model performance continues to improve at these later stages of training.

Contributions

Conversation example with CroissantLLMChat
Figure 1. Conversation example with CroissantLLMChat

In this work, we attempt to bridge the aforementioned limitations and gaps. Our main contributions can be summarized as follows.

Contribution 1: Introduction of a highly-curated, diverse corpus in French. We collect and openly release a 303B token corpus spanning internet data, but also literary work, speech transcripts, legal and administrative documents, scientific articles, business-related documents, etc. The corpus is distributed under permissive licenses, allowing commercial use with no restriction, and is heavily filtered, curated, and deduplicated. To our knowledge, it is the largest multi-source French language corpus released to date of sufficient quality for language modeling purposes.

Contribution 2: Training CroissantLLM, a truly bilingual language model. Nowadays, most models display some multilingual capabilities. For example, Bloom has been trained to be massively multilingual , Llama contains a minor proportion of non-English data in its training set (Touvron et al. 2023) and Qwen included a significative portion of Chinese data (Bai et al. 2023). However, to our knowledge, outside of Chinese with a different alphabet , no work has studied or attempted to train a multilingual model of significant scale in which English is not the dominant training language.

In this work, we train a model on a 1:1 ratio of English to French with a tokenizer optimized for bilingualism. Our end goal is to have a model less skewed towards English performance or cultural biases.

We motivate this ratio by conducting careful experimentation on smaller model scales to uncover the trade-offs behind language distribution ratios. We opt for a strategy enabling a “best of both worlds'' performance in these two languages, as empirically validated by scaling law experiments. Experimentally, we show the importance of integrating data from another cultural source in the acquisition of cultural knowledge, underlining the importance of the effort.

Contribution 3: FrenchBench: A novel LLM benchmark for the French Language. To evaluate models in French, we introduce a benchmark encompassing various tasks to assess factual knowledge, generative capabilities, language understanding, etc. This benchmark is constructed both from openly available datasets, as well as newly released manually annotated data. We evaluate and report results for our models as well as other models with French-speaking capabilities.

Contribution 4: Releasing high-performing, inference-optimal models for the industrial community, together with a wide range of resources for the research community. The models we train are all released under open licenses. Our largest model is trained on a 2300:1 token to parameter ratio (115 times longer than a Chinchilla Optimal 1.3B model) leading to very strong performance in its size category. We show that model performance on downstream tasks continues to dramatically improve with lengthened pre-training runs, although model perplexity does not significantly improve. We release all checkpoints for all model sizes, as well as the exact training data seen at every step for research purposes. These models are extremely efficient to run, leading to low-latency, energy-efficient on-edge inference, even on low-resource devices such as phones or personal computers. These releases are motivated by a commitment to transparency to allow research and reassuring users for industrial deployment: our models comply with 81 % of criteria listed by the Foundation Model Transparency Index (see Section section).

The CroissantLLM initiative aims to adress the aforementionned limitations of current models through the release of an inference-optimized, small but capable model that performs well outside of English settings, and that is designed to be as open and transparent as possible. Beyond facilitating industrial LLM adoption and unlocking new generative model use cases, this project is also a proven platform for researching LLMs and kickstarting future pretraining efforts (Martins et al. 2024; Meeus et al. 2024; Shilov et al. 2024; Meeus et al. 2025).

Data

The dominance of the English language in the training data of most current models is undeniable. While multilingual models like Llama leverage some non-English data (Touvron et al. 2023), it corresponds to only a minor part of the corpus, leading to a significant drop in performance across non-English data, with noticeable “American'' bias .

This work aims to offset this trend by using a more balanced bilingual corpus comprising English and French, as well as additional code data. Although both languages belong to the Indo-European language family, they exhibit different morpho-syntactic structures and French has a richer morphology. We study whether this corpus helps in reducing biases, enabling more varied knowledge sets, and unlocking non-English performance.

A variety of sources are integrated into our corpus, including carefully filtered internet data and high-quality data from a range of sources, all devoid of restrictive licenses ensuring complete openness of the data and the trained model. Data statistics are available in Table 1. As further detailed in 7.3.4, our data corpus contains different amounts of unique English, French, and Code tokens. We obtain our balanced training corpus by upsampling French, Code, and English data with different sampling ratios, such that no performance loss is to be expected .

Table 1. Final Data mix for CroissantLLM training
Size (GB)Docs. (M)Tokens (B)Token/DocSampling Ratio# tokens (B)
French1258.70376.27303.51806.634.091240.08
English2351.13591.23655.641108.941.891240.09
Code366.8781.90141.431726.762.04288.92
Parallel113.91408.0335.7887.686.13219.26
Total4090.611457.431136.35779.7014.152988.35

The scrapping and processing code are available in our code base. The license information of all datasets used is given, all allowing for permissive commercial use.

French Data

Table 1 lists the source and some information regarding the French corpus. Details about the data sources are expanded further below.

Web Data. We collect internet data from various web scraps (Oscar , mC4 ), leveraging the CulturaX corpus for heuristic and perplexity filtering, as well as exact and fuzzy deduplication. In total, this represents over 363 million webpages and more than 292 billion tokens, that we split using our custom tokenizer fitted on equal amounts of French and English data.

We ensure data is of good quality and correctly tagged in French through sampled manual inspection and confirm French-speaking countries are well represented within the dataset. Notably, we include several news sources scraped from Belgium, Switzerland, Canada, and Lebanon, as well as multiple African countries (Senegal, Morocco, Algeria, Cameroon, etc.)

Legal and Administrative Data. We introduce 5.3B tokens of data from the French government's open data initiative, ranging from legal tasks to parliamentary discussions and financial transcripts (e.g. legal and administrative jargon). These texts originate from 13 different datasets (the OpenData corpus) and were collected from the French government's open data platform. To ensure other French-speaking countries are represented, we add 68M tokens of data from Swiss legislation retrieved from government sources. We perform steps to process, filter, and run exact deduplication on these documents.

Cultural Data. We introduce cultural data from various sources. Notably, we retrieve all Project Gutenberg books in the French language as of October 2023, corresponding to books released in the public domain (302 million tokens). We also download and aggressively filter manuscripts and documents from the French National Library (Bibliothèque Nationale de France), and filter for documents that belong to the public domain, have undergone an OCR process, and are of high quality. To filter out low-quality OCR documents, we implement custom heuristics which we release within our code base. We run all documents through perplexity filters using KenLM 5-grams fitted on the French Wikipedia split, and discard documents with perplexity values that are too high (noisy) or too low (repetitive patterns). Thresholds are set through a manual verification process. We deliberately choose to be aggressive in our filtering to ensure only high-quality data is kept and discard the largest portion of the original corpus, keeping about 27M tokens. We choose not to keep any data from the newspaper archives, as the OCR transcription is often too noisy. Additionally, we introduce famous public domain French poems custom scraped from a French poetry website, and run a set of podcasts through a high-quality speech-to-text model to obtain a textual transcription. This process is hard to scale and data splits from these sources are limited in quantity. Data from the OpenSubtitles initiative is integrated, corresponding to 41.8 million tokens originating from movie subtitles. Finally, we add the French data from Wikisource collected as part of the BigScience initiative and obtain 2.7 billion tokens from the process.

Encyclopedia Data. To introduce high-quality factual data to the training corpus, we integrate the French Wikipedia split from November 2023. This corresponds to the latest knowledge cutoff in the training data. In total, more than 2.5 million articles are used, spanning more than 2 billion tokens.

Industrial Data. We scrap high-quality and publicly available data from industrial PDFs via a manually crafted list of websites, from large French and French Canadian (Quebec) companies to government agencies. This business-focused data boosts performance on a series of downstream applications related to industrial NLP. We collect over 36000 PDF multi-page documents and filter them through carefully crafted heuristics, followed by aggressive perplexity filtering.. In total, we obtain over 290000 documents and 191 million tokens.

English Data

Our English data is primarily drawn from the SlimPajama corpus , excluding copyrighted documents. Splits per data source are detailed in Table 2.

Internet Data. Similarly to the French dataset, we rely on carefully filtered content from an assortment of internet sources, including miscellaneous web pages and blogs. The filtering process includes heuristics and perplexity filtering, as well as large-scale deduplication . The SlimPajama corpus includes internet data from the CommonCrawl and C4 web scraps, as well as data sourced from Github textual content and the StackExchange forums.

Miscellaneous. Other non-internet-based data sources are included in the SlimPajama dataset, such as scientific articles from Arxiv and English documents from Wikipedia. The SlimPajama dataset also comprises the “Books'' subcorpus, obtained by downloading all book documents from Bibliotik. Some of the documents within this last corpus have been flagged by their owner as proprietary data. We filter out all documents from this subcorpus, and replace them with data from the open-source Project Gutenberg English books under public domains.

Gutenberg Canaries. To assess model memorization to inform about the risks of including private or sensitive data within the training set, we stress test the model by including “canaries'' . These correspond to samples that have been intentionally modified and/or repeated and included within the model training set, and that will enable a posteriori evaluation of the model capacity to memorize data in a “worse than worst-case'' situation. In total the canaries represent 555 million tokens, representing less than 0.04 % of the total tokens seen during training.

Code Data

In line with most recent models (Chowdhery et al. 2022; Scao et al. 2022; Touvron et al. 2023), we integrate code data into our training corpus. Notably, previous work shows that code data benefits natural language tasks and can be particularly useful in data-constrained settings . Therefore, we include 140B tokens of code data in several common programming languages. Splits and numbers of tokens are detailed in Table 3.

The Stack & StarCoder. We rely on the efforts of the StarCoder project , and use their high-quality filtered code data from The Stack corpus . We keep only high-resource programming languages (Java, Javascript, Python, C, C++, SQL) and Jupyter notebooks, as well as a few samples of formatted data (JSON, Tex) and scripting languages (shell scripts, Dockerfiles).

Extra Python code. We extend the corpus with several other sources of Python code due to the popularity of the language in the community. Firstly, we add Pypi packages from recent code dumps, that are filtered to keep only Python and Jupyter files. Secondly, in order to promote high-quality problem-solving-centered code, we integrate 1.2B tokens of Python3 data from competitive coding contests under CC-By-4.0 license. Lastly, following the success of learning from textbooks , we add commented Python code constructed by combining code and text cells from Jupyter Notebooks through the CodeParrot initiative.

Parallel Data

Following previous work , we incorporate vast quantities of parallel data, in our case high-quality English-French translation pairs, in order to improve the multilingual capabilities of the model .

Opus. We extract subsets of sentence pairs spanning multiple domains from the OPUS corpus . Statistics are described in Table 2. In total, we include 400 million parallel sentences and about 36 billion tokens. The data is filtered through a rigorous cleaning pipeline: (1) BiFixer is first used to remove duplicate data through fuzzy matching techniques; (2) BiCleaner is then used to filter data using heuristic and perplexity filters; (3) finally, the state-of-the-art NMT quality estimator CometKiwi is used to keep only top quality translation pairs.

Theses. To incorporate versatile academic and scientific language, we augment our dataset with French theses abstracts along with their author-generated English translations. This corresponds to 95000 documents and more than 80 million high-quality tokens.

Song Lyrics. Our dataset integrates song lyrics in both French and English, scraped from a specialized community-driven lyrics translation website. As such, our model is trained with radically different linguistic styles (e.g. colloquialism), and the wide range of topics can help the model to capture cultural nuances. Lyrics have been translated by hand by the website community. With a total of 70k songs, we have built up a corpus of 53M tokens covering different periods (80s, 90s, 20s, etc.) and musical genres (rap, rock, jazz, etc.) To preserve colloquial expressions and cultural subtleties, we have not filtered song lyrics for explicit content. We validate the original language metadata of the songs through Google's language-detection algorithm.

Training

Our main goal was to train a high-performing, yet resource-friendly bilingual model while optimizing performances across both languages. To focus on the specific challenges of the bilingual paradigm, we rely on previous work to motivate many of our design and hyperparameter choices .

Model Architecture

We use the Llama architecture (Touvron et al. 2023), a decoder-based transformer, trained with rotary position encodings and a context length of 2048. We construct 4 different model sizes by jointly scaling the number of attention heads, hidden size, and hidden layers. Table 2 summarizes the sizes of each model in the family.

Table 2. Model information for scaling laws. Parameter count excludes embedding and output parameters.
ModelParams (M)LayersHidden sizeInter. sizeKV heads
XXS100.76102440968
XS202.512102441288
S341.5121536412812
Base1214.3242048550416

Tokenizer

Most LLM tokenizers are fitted on English-centric corpora with an information-theoretic optimization objective, for example, Byte-Pair encoding or Unigram , leading to good fertility values (low token per word ratio) on English text, but high fertility in other languages. These phenomena make processing in other languages slower and more costly . Furthermore, subword splits in non-English languages mechanically carry less semantical meaning, potentially being a factor in the degraded performance of models on non-English languages (Rust et al. 2021).

Tokenizer training. We fit our CroissantLLM tokenizer on a corpus of 100B high-quality tokens, with splits of English, French, and code data. We use SentencePiece to train a Byte-Pair Encoding tokenizer with a vocabulary size of 32000 tokens, 100 special placeholder tokens, whitespace separation, and byte fallback, inspired by (Touvron et al. 2023; Jiang et al. 2023). The data corpus used to fit the tokenizer is made available, and notably contains large amounts of French data to skew the vocabulary construction process towards optimizing for French as well.

Improved fertility rates. The focus on English, French, and Code enables the CroissantLLM tokenizer to display smaller fertility rates on French texts than the Mistral and Llama models with similar vocabulary sizes, all the while also displaying slightly smaller rates than both in English and Code (Figure 2). This is due to the multilingual support of both Llama and Mistral tokenizers which need to allocate some vocabulary tokens to frequent character patterns from other languages. Roughly, the Llama tokenizer is 17 % less token efficient at encoding French internet data, and up to 40 % less efficient on clean encyclopedia French texts, implying that the 303B unique French tokens in our data training set correspond to more than 360B tokens with the Llama tokenizer. This enables us to pack more data in fewer tokens, leading to improvements in training and inference efficiency.

Fertility on unseen test sets using various tokenizers. Lower is better.
Figure 2. Fertility on unseen test sets using various tokenizers. Lower is better.

Selecting an optimal language ratio

A crucial question when training a bilingual model is how to effectively weight data from the two languages to achieve a good trade-off between performance in both. While, intuitively, training on an equal mixture of English and French data may seem to be the obvious solution, differences in data quality available for each language coupled with transfer learning dynamics between both languages could imply that a balanced mix might be sub-optimal. However, training multiple models with different data mixes for comparison is prohibitively expensive.

To offset this computational cost, we leverage recent findings on scaling laws that show that we can predict the performance of our model by training smaller models on the same dataset. In particular, found that, for multilingual models, by training smaller models with varying weights for each language in the data mix, one can fit a multilingual, joint scaling law that can predict the language performance trade-off of larger models, even for novel language weightings not encountered during the fitting of the scaling law.

As such, we fit a joint scaling law as described by (Fernandes et al. 2023) for each language, by training 3 smaller model sizes on 3 different data mixes with varied ratios of English and French data (keeping the amount of Code data fixed). The corpus for these scaling law experiments is a subsampled variant of the larger corpus and is detailed in Appendix G.3.5. We define 3 data mixes by varying the language sampling ratio: (1) equal containing 40 % English data, 40 % French data and 20 % Code data; (2) frplus containing 20 % English data, 60 % French data and 20% Code data; and (3) enplus containing 60 % English data, 20 % French data and 20 % Code data. We then trained a 1.3B model on these subsets of the data for one of the data mixes to validate their predictive power.

Evolution of test cross-entropy loss with model size in English (left) and French (right), for the wiki domain, as well as the fitted joint scaling laws,
Figure 3. Evolution of test cross-entropy loss with model size in English (left) and French (right), for the wiki domain, as well as the fitted joint scaling laws,
Effective capacity ratio (as predicted by our fitted joint scaling law) for English and French as we change the weight of each language.
Figure 4. Effective capacity ratio (as predicted by our fitted joint scaling law) for English and French as we change the weight of each language.

Figure 3 shows the performance predicted by jointly-fitted scaling laws as we scale the model and vary the language weightings on the Wiki data validation split.

First, we see that the fitted scaling law is able to predict the performance of the larger model almost perfectly. Secondly, changing the weight of each language in training has a non-symmetrical impact on language performance: by increasing the (relative) weight of French from 50 % to 75 %, we get a marginal performance increase in French, while performance in English drops significantly. This fact is made clear by plotting the effective capacity ratio of each language as we change the language weight (Figure 4): the “gains'' in parameters from increasing weight of French data are minimal past the 50 % mark.

These findings showcase that multilinguality comes at a price, and training a bilingual model implies accepting a performance loss on a target language compared to an equivalent model trained on a monolingual corpus.

We find equal ratios of English and French data lead to minimized performance hits across both languages (Figure 3) and opt to train our base model in this data configuration.

Final data distribution

Our final dataset is composed of 1.1T unique tokens that originate from sources of various languages, qualities, and quantities. To craft a training set with a language and data distribution that suits our objectives, we upsample some of the sources, notably to balance out French and English data and increase the share of parallel data in our training run. Following work by and on data-constrained language modeling scaling laws, we upsample French text by a factor of two, and parallel data by a factor of 3. For a 3T token training run, this enables the model to see French data at most 4 times, and English and code data twice, which should have negligible impacts on the performance . The final data distribution is shown in Table 1.

All data is provided from the above-listed sources and no synthetic or augmented data is used. Data licenses and copyright information are given for every split to the best of our ability. The data collection and filtering process to construct our final mix from the above-listed sources is entirely done by the authors of this paper, who are employed by the universities or private companies described through their affiliations, under their countries' data protection laws, and compensated at market rates or following the academic salary grid of their institution.

Training framework

We train our models on a modified version of Megatron-Deepspeed, a training framework built on top of PyTorch. Training is done on a dedicated Nvidia A100 SXM4 80 Gb supercomputer partition

with 30 octo-GPU nodes. We rely on the HuggingFace Transformers and Datasets library for model and data manipulation.

To maximize efficiency, we set the micro-batch size per device to 8 sequences of length 2048, and use 4 gradient accumulation steps, resulting in a total batch size of 8×4×30×8=76808 \times 4 \times 30 \times 8 = 7680 samples, or 76802048=15,728,6407680*2048 = 15,728,640 tokens. We achieve a mean efficiency of around 120 TFLOP per second with activation checkpointing, leading to a total compute estimate of 4.30e224.30e22 FLOPS. Standard Cross-Entropy losses are used on a Causal Language Modeling objective.

Training losses

Training lasts 17 days for a total of 99648 GPU hours, and we chose not to manually intervene, letting the model recover on its own after occasional loss spikes. We train with a max learning rate of 3e43e-4, 1000 warmup steps, and a cosine learning rate with a minimum value of 1e51e-5. Curves suggest the model still has not reached a performance plateau after 3T tokens (Figure 5). Checkpoints are stored every 5k steps and released with the rest of the project artifacts.

(Left) Training loss with respect to the number of seen tokens. (Right) Validation perplexity (Averaged Log Likelihood) on Wikitext (English), computed with a rolling stride, panel 1(Left) Training loss with respect to the number of seen tokens. (Right) Validation perplexity (Averaged Log Likelihood) on Wikitext (English), computed with a rolling stride, panel 2
Figure 5. (Left) Training loss with respect to the number of seen tokens. (Right) Validation perplexity (Averaged Log Likelihood) on Wikitext (English), computed with a rolling stride

Environmental impact

The model was exclusively trained on a supercomputer operating on low-carbon nuclear electricity. Between experimental runs, scaling laws, and the final training, 123k A100 hours were used. The Thermal Design Power of the NVIDIA A100 SXM4 80Gb used is 400W corresponding to a total power consumption of 49.2 MWH and considering a grid carbon intensity of 57 gCO2eq/kWh, we estimate a carbon footprint of 2.80 tons of CO2 emitted during training.

Interestingly, the model we trained is not “compute-optimal'' according to Chinchilla laws , meaning that less computing could have been used to train a larger model with the same performance. However, our model aims to be used for inference purposes at industrial scales. Our training paradigm is thus to absorb the downstream inference costs, by training a smaller model on a lot more tokens to obtain an inference-optimized model equivalent in performance to a bigger compute-optimal model . Each inference of the final model is thus vastly more energy-efficient than a Chinchilla optimal model of equivalent performance (>> 3B parameters), and can even run on CPU or mobile devices. Relying on estimates of , at inference, CroissantLLM represents roughly 2.6 GFLOPS per token.

We hope to extend base model evaluation past English benchmarking alone and assess model capabilities in French, aiming for broad coverage across orthogonal capabilities to observe the effect of truly bilingual pre-training. Our evaluation efforts are rooted in transparency, and all results reported in the main technical report are reproducible through code that is open-sourced and public data.

Evaluation Benchmarks

We hope to extend base model evaluation past English benchmarking alone and assess model capabilities in French, aiming for broad coverage across orthogonal capabilities to observe the effect of truly bilingual pre-training. Our evaluation efforts are rooted in transparency, and all results reported in the main technical report are reproducible through code that is open-sourced and public data.

English

In English, we evaluate on standard LLM evaluation benchmarks.

HellaSwag. HellaSwag is a dataset specifically crafted to challenge common-sense reasoning abilities of models by requiring them to predict the endings of sentences in a way that relies on information not present in the preceding context. It focuses on capturing a nuanced and context-dependent understanding of language.

PiQA. PIQA is a dataset for common-sense reasoning and was created to investigate the physical knowledge of existing NLP models .

SciQ. The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry, and Biology, among others. The questions are in multiple-choice format with 4 answer options each .

Arc-C. The AI2 reasoning challenge dataset consists of 7,787 authentic grade-school level, multiple-choice science questions, designed to stimulate research in advanced question-answering. The dataset is divided into a Challenge Set and an Easy Set, with the Challenge Set comprising questions that were answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. Additionally, the dataset includes a corpus of over 14 million science sentences relevant to the task and provides three neural baseline models.

MT-Bench. MT-Bench contains a set of prompts designed to evaluate models on their multi-turn conversation and instruction-following abilities, covering various core model abilities; writing, roleplay, extraction, reasoning, math, coding, knowledge I (STEM), and knowledge II (humanities/social science). MT-Bench performance has been shown to best correlate with human-rated appreciation of a model through the LM-Sys model arena.

French

We aim to evaluate models on their capabilities in French, along several axes including vocabulary, grammar, reading comprehension, factual knowledge, biases, and generative capacities, etc. To this end, we introduce FrenchBench, a novel LLM evaluation benchmark for the French language, testing a large array of model skills in various settings.

FrenchBench comprises several tasks, some included from previous benchmark datasets, others newly released with this work.

FrenchBench Gen

FrenchBench assesses the generative capabilities of LLMs in a few-shot setting. Tasks include title generation, summarization, question generation, and question answering. We detail the tasks and the evaluation metrics used below.

FQuaD. FQuaD is a French Question Answering dataset, containing manually annotated sets of Wikipedia passages, questions, and extractive answer spans in the Squad format. This high-quality dataset is one of the rare human-annotated French datasets and we rely on its public evaluation split for 4 of the FrenchBench tasks.

FQuADGenQ is a question generation task in which passages and answers are given to the model in a few-shot manner, and we compute the ROUGE1 score with the gold questions.

FquadGenAns is the classic question-answering task, but models generate the answer span themselves, and the ROUGE1 score is computed with the gold extractive answer span.

MultiFQuAD is a FQuAD variant, with a publicly released evaluation set, in which answers can consist of multiple disjoint spans. We evaluate performance on the concatenation of these gold extractive spans using the ROUGE1 score.

French Trivia. The French Trivia dataset is built from online trivia questions pertaining to French culture. Answers are short and meant to assess latent model knowledge and the impact of pre-training data and cultural references. Intentionally, questions are formulated in English for comparison with monolingual English models.

FrenchBench Multiple Choice

We also assess reasoning, factual knowledge, linguistic capabilities, and model biases through a series of few-shot classification tasks, on which models are given multiple completions (multiple choice), and the answer with the highest likelihood is selected. We experimented with multiple question templates. In the MMLU format, the multiple potential answers are given after the question prefixed by a letter (A, B, C, D) and the model must guess the correct answer by predicting the correct answer's letter. In HellaSwag formatting, the model must complete the sentence and the model chooses the most likely continuation sequence, without prior knowledge of all other options. We find HellaSwag formatting is less abstract, and enables smaller size models to perform better.

French Language Test. The French Language Test is a dataset crafted to assess the grammar and vocabulary capabilities of models through language tests. It provides a structured evaluation of a model's linguistic proficiency, aiming to measure its competency in understanding and generating coherent and grammatically accurate sentences in the French language. It is composed of a fr-grammar and fr-vocabulary multiple choice test.

French Hellaswag and Arc-C. These datasets correspond to machine translations made by GPT3.5 of HellaSwag and Arc-C to French. Manual verification of the translation quality indicates the translations to be far from perfect but sufficient for these datasets to act as a correct performance proxy.

OrangeSum. OrangeSum is a summarization dataset constructed from online News articles. Two standard French summarization tasks span from this dataset; OSum(T) in which the model is tasked with generating the title from the article body, and OSum(A) in which the model must generate the first paragraph of the article aimed to be an abstract of the article. We select the abstract generation task, and measure performance with the ROUGE1 score.

Other Tasks

MT-Bench French. Mt-Bench French is a translated and adapted version of MT-Bench in French with all questions having undergone rigorous human review and adaption to guarantee authentic wording, and coherence, and to account for cultural discrepancies.

Translation. Translation capabilities are evaluated through the test set of the 2014 WMT French-English and English-French tasks . We measure performance using BLEU score , and COMET . We also report FLORES and TICO scores.

Belebele. Belebele is a challenging reading comprehension dataset, with multiple choices, released across 122 languages in parallel format . We leverage the English and French splits.

Benchmark results

Baseline models. To evaluate CroissantLLM, we compare with an array of various models, varying in parameter size, pre-training language distribution, training corpus size, etc.

For “monolingual'' English models, we evaluate Pythia-1.4B trained on 300B tokens, OPT-1.3B trained on 180B tokens, and TinyLlama(1.1B) . TinyLlama is a very strong English baseline, as it holds many similarities to CroissantLLM. It is a 1.1B model trained on 3 trillion tokens with the same English corpus as the Croissant base. Although it contains some amount of high-quality non-English data, it is only a minor share of the training corpus, the main data sources being English and code data. As such, it trains on much more English tokens than CroissantLLM. All models are trained way past Chinchilla optimality (\sim26B tokens for a 1.3B model).

For monolingual French models, we use GPT-fr , a 1B model trained on 16.3B tokens, as well as the PagnolXL(1.5B) model , both in their author submitted HuggingFace implementations.

We also compare CroissantLLM with multilingual models, notably Llama2(7B) (Touvron et al. 2023) trained on 2T tokens, Mistral7B (Jiang et al. 2023), and Bloom (Scao et al. 2022) models (from 1.1B to 3B), trained on 350B tokens each. We note that although the largest Bloom model is undertrained according to Chinchilla optimality (Hoffmann et al. 2022), smaller models are trained on the same number of tokens, making them largely more inference optimal and thus strong contenders. Finally, in the same size category, we evaluate mGPT a 1.3B model trained on 440B tokens.

Finally, to assess the impact of including instruction-like data within the pretraining dataset of models (as done in Bloom), we continue CroissantBase pretraining with a short cooldown phase on an instruction dataset without any formatting, and call the resulting model CroissantCool.

Base model

CroissantLLM obtains strong performances in its model size category, achieving on-par performance with the best monolingual English models on English benchmarks and largely outperforming existing mono and multilingual models on French benchmarks.

English. On English benchmarks (Table 3), CroissantLLM displays performances almost equivalent to those of TinyLlama, which has trained on much more English data. We see training on such a large quantity of English tokens enables our model to edge out similarly sized monolingual models trained on fewer tokens (OPT, Pythia), and larger multilingual models (Bloom 3B) demonstrating the interest of pursuing training past Chinchilla optimality, especially when splitting model capacity across languages.

Table 3. English Benchmarks (5-shot results)
TaskArc-eBelebele (eng)HellaswagPiQASciQAvg
GPT-fr(1B)0.270.280.290.540.680.41
Pagnol-XL(1.5B)0.340.250.310.560.760.44
mGPT(1.3B)0.480.230.350.660.620.47
Bloom(1.1B)0.550.240.360.680.890.54
OPT(1.3B)0.610.230.420.720.920.58
Pythia(1.4b)0.630.250.420.710.920.59
Bloom(3B)0.640.240.420.710.930.59
CroissantLLM(1.3B)0.620.280.420.720.920.59
CroissantCool(1.3B)0.620.260.430.730.920.59
TinyLlama(1.1B)0.650.260.450.730.940.61
Llama2(7B)0.790.460.560.790.970.72
Mistral(7B)0.830.850.600.820.980.81

French. On French classification benchmarks, CroissantLLM largely outperforms models of similar sizes trained on mostly monolingual English or French data, and multilingual models (Table 4). Performance is on par with the Bloom(3B) model, which is about 3 times as large. An interesting phenomenon can be noticed, especially on generative benchmarks assessed in few-shot settings: “base'' models trained with instruction-like data perform a lot better. This is noticeable with the Bloom(3B) model which outperforms the otherwise vastly superior Llama2(7B) model on several tasks, or through the performance gains of CroissantCool with respect to CroissantBase.

Table 4. FrenchBench MC (5-shot results)
TaskHellaswag(fr)Arc-c(fr)fr-vocabfr-grammarBelebele(fr)Avg
OPT(1.3B)0.280.190.500.610.280.37
Pythia(1.4B)0.300.200.610.760.230.42
TinyLlama(1.1B)0.330.230.640.670.250.42
mGPT(1.3B)0.270.200.710.730.230.43
GPT-fr(1B)0.300.190.700.790.240.44
Bloom(1.1B)0.340.220.760.790.240.47
Pagnol-XL(1.5B)0.330.210.770.820.270.48
CroissantCool(1.3B)0.400.260.770.780.230.49
CroissantLLM(1.3B)0.400.260.750.800.270.50
Bloom(3B)0.400.270.780.810.230.50
Llama2(7B)0.440.380.760.770.430.56
Mistral(7B)0.490.470.780.780.780.66
Table 5. FrenchBench Gen (5-shot ROUGE1 results). Bloom models seem to have strong performance on QA tasks (Fquad), likely due to the inclusion of Question Answering datasets in its pretraining corpus (Laurençon et al. 2023). Pagnol-XL and GPT-fr are trained exclusively on French text and as such cannot be fairly evaluated on the French Trivia test.
TaskFGenQFGenAnsMultiFQuADOSum(A)FTriviaAvg
Pagnol-XL(1.5B)0.060.040.030.03-^*0.04
GPT-fr(1B)0.040.020.050.11-^*0.06
mGPT(1.3B)0.010.000.020.030.330.08
OPT(1.3B)0.090.180.210.170.390.21
Bloom(1.1B)0.170.280.260.100.310.23
Pythia(1.4B)0.150.340.270.210.440.28
CroissantLLM(1.3B)0.190.400.330.100.520.31
Bloom(3B)0.210.470.370.180.470.34
TinyLlama(1.1B)0.180.460.410.230.450.35
CroissantCool(1.3B)0.200.450.360.270.530.36
Llama2(7B)0.250.680.600.300.700.50
Mistral(7B)0.330.780.640.310.740.56

Improvements throughout training. The model performance continues to improve on downstream tasks during the entirety of training. We report WMT14 translation performance in Figure 6, and observe similar trends across all tasks. The benefits of training past Chinchilla optimality are clear, and although there are diminishing returns past a certain number of steps, training does not seem to saturate. In low training step settings, performance appears to emerge suddenly, reflecting emergent performance experiments in the literature most often obtained through model scaling .

Performance evolution on the WMT Translation task (5-shot)
Figure 6. Performance evolution on the WMT Translation task (5-shot)

French Trivia. One main question this work attempts to tackle is whether training on bilingual data goes beyond augmenting the language understanding and writing capabilities of a model in another language, but also equips the models with novel knowledge and different cultural biases. We evaluate French cultural knowledge on a Trivia task, consisting of questions about France-related topics, asked in English (Table 5), and score results obtained in 5-shot settings with ROUGE-1. CroissantLLM is the best performing model evaluated under the 7B size, outperforming English-centric models by significant margins. This knowledge gap showcases the effect of the pretraining data mix in specific knowledge acquisition, underlining the interest of integrating vast amounts of varied and multilingual data when training models aiming for broad knowledge coverage.

Overall. The 1.3B CroissantLLM displays top-of-its-class performance across both languages and all benchmarks, even edging out larger models such as Bloom(3B) on most tasks. All models remain far off from the performance of the strong 7B Llama and Mistral models.

Finetuning

Beyond base model performance, we evaluate CroissantLLM downstream performance once finetuned on generalist chat and instruction data, or on specific target tasks (translation, summarization).

Chat Model

It has been shown that supervised fine-tuning on instruction or chat datasets enables leveraging model capabilities to their fullest .

Training. We finetune the base model on public Chat datasets Ultrachat and Wildchat containing ChatGPT interactions in English and French. We also incorporate 12k samples of translation data (4 % of the SFT dataset). We run finetuning on CroissantLLM, as well as the Bloom-1b7 and TinyLlama models for comparison. The obtained models are further suffixed with “Chat''.

MT-Bench.

We evaluate models on the MT-Bench benchmarks, both in English and French. Although a large difference in performance can be noted between the Bloom model and Croissant in favor of the latter, performance differences with TinyLlama are not as significant, neither in English nor in French. CroissantLLMChat performs strongly in open-ended writing categories (writing, roleplay, humanities) but struggles with reasoning and extractive tasks. Turn 2 performance (reformulation under constraints) is largely lower than Turn 1 performance as can be seen in Figures 3 and 4. Our CroissantLLMChat model also vastly outperforms the BloomZ 3B model trained by CMArkea on a large chat finetuning corpus .

This hints at the fact that quasi-monolingual models with only a minor share of another language in their pretraining corpus can be adapted to a reasonable extent, through subsequent finetuning or continued pretraining, although large pre-training corpora are necessary to incorporate sufficient knowledge and reasoning abilities within the base models. We notice large correlations between generation temperature and performance and find CroissantLLMChat works a lot better with higher temperatures (0.4\geq 0.4). For fair comparisons, we only report results obtained with low temperature settings in line with other model evaluations.

MT Bench Results (Both Turns)
Figure 7. MT Bench Results (Both Turns)

Translation. We run translation evaluations on the Chat models and report results in Table 6. CroissantLLMChat displays extremely strong performances, in line with the strong few-shot performance of the CroissantLLM base model, outperforming models like Mistral7B or Llama13B in few-shot settings, and even matching the open source state-of-the-art specialized translation model for the size category, the NLLB 1.3B , trained on vastly superior amounts of parallel data.

Table 6. Performance in machine translation, according to COMET-22 and BLEU, across three different benchmarks: WMT14, TICO and FLORES. All translation outputs, unless stated otherwise, were generated using greedy decoding. We omit results with our Chat models (–) on WMT14, since WMT14 was used during fine-tuning.
WMT 14TICOFLORES
en\rightarrowfrfr\rightarrowenen\rightarrowfren\rightarrowfrfr\rightarrowen
CometBleuCometBleuCometBleuCometBleuCometBleu
NMT models
NLLB 1.3B86.8241.5984.5536.4781.1540.2287.1047.4987.2140.47
0-shot
Pre-trained models
LLaMA-2 7B84.3732.9886.6638.5778.0533.7585.0338.5988.7541.83
5-shot
LLaMA-2 13B85.9436.7687.0239.9380.0438.2186.6743.4989.0342.71
5-shot
Mistral-7B-v0.184.9934.8287.0139.5579.3437.8286.0741.3188.3642.56
5-shot
TinyLLaMA73.0318.1382.9929.8569.2020.5574.4021.1785.8633.10
5-shot
CroissantLLM85.1138.0985.7036.3078.7438.4986.8546.5888.5842.83
5-shot
SFT models
TowerInstruct-7B-v0.188.0746.1988.1446.7581.5341.2788.3848.5789.5646.34
0-shot
TinyLLaMAChat73.0423.6178.0827.2486.2632.80
0-shot
CroissantLLMChat80.2736.9986.8244.7988.3841.54
0-shot
CroissantLLMChat80.7238.3487.6847.1188.7142.90
0-shot (Beam Search)

Dialog Summarization finetuning

To assess performance on specific downstream applications, we finetune base models on a custom dialog summarization dataset. Models are finetuned for three epochs on 6000 samples and results are computed through ROUGE and GPT-4 judgment metrics (Table 7).

Table 7. Dialog summarization results. Except ROUGE-1, scores are measured by GPT-4, out of 5.
ModelROUGE-1CoherenceConsistencyFluidityRelevance
CroissantLLM (1.3B)0.5504.563.934.734.09
Bloom (1.7B)0.5504.523.964.764.08
Mistral (7B)0.5884.604.734.724.59

CroissantLLM and Bloom(1.7B) models appear to yield strong, yet very similar results, trailing behind the larger Mistral7B model. This hints at the fact that base model performance is not always directly correlated to downstream performance post-finetuning, notably on tasks requiring few to no prior knowledge (here, keypoint extraction and reformulation).

Optimized Inference

Our largest model, CroissantLLM, with 1.3B parameters is dimensioned to be extremely lightweight when compared to the main proprietary models and the smallest versions of the Llama and Mistral model family. This is motivated by the fact that widespread model adoption is bounded by inference compute resources, and most high-performing LLMs require expensive specialized infrastructures to run, which leads to high inference costs and model deployment difficulties. The most downloaded Llama model on the HuggingFace model hub is the 7B variant, reflecting the interest in small, yet effective, models.

At a 1.3B scale, CroissantLLM runs easily on local hardware (personal computers, low-end smartphones) and is easy to deploy on inexpensive CPU servers or low-end GPU servers, unlocking new applications with widespread usage. On higher-end GPUs (Table 8), CroissantLLM is both faster (latency) and less memory intensive enabling it to fit bigger batch sizes (throughput). Performance benchmarks are given in Table 8.

Table 8. Inference Results in French and English on an A100 GPU with 40GB VRAM (average results over 100 tokens generations with 100 tokens input based on 100 Wikipedia text samples, vLLM backend and batch size 1)
ModelParameters (B)Tokens Per SecondWords Per Second
French
Llama 21338.5622.18
Llama 2764.0537.12
CroissantLLM1.3145.40101.12
TinyLlama1.1152.6090.08
English
Llama21338.1728.16
Llama2762.6046.49
CroissantLLM1.3139.64111.41
TinyLlama1.1150.16112.36

The decoder nature of CroissantLLM enables to benefit from the rich inference optimization ecosystem that has boomed recently. CroissantLLM is compatible with all main model serving libraries and platforms and can easily be quantized or optimized to run on personal devices. We performed 4bit quantization in the GGUF format and were able to run the model on lower-end smartphones at a speed of more than 5 tokens per second.

Model limitations

Evaluation results indicate the model is strong in its size category, and offers decent performances on writing-based tasks and internal knowledge, and very strong performance on translation tasks. The small size of the CroissantLLM model however hinders its capacity to perform more complex reasoning-based tasks, at least in a zero or few-shot manner in its generalist base or chat-model versions. This is aligned with other models of size and underlines the importance of scale for more abstract tasks .

Knowledge Cutoff. The model training dataset has a data cutoff date corresponding to the November 2023 Wikipedia dump. This is the de facto knowledge cutoff date for our base model, although a lot of information dates back further. Updated versions can be trained through continued pre-training or subsequent fine-tuning.

Hallucinations. CroissantLLM can hallucinate and output factually incorrect data, especially regarding complex topics. This is to be expected given the small model size, and hallucination rates seem inferior to most models of the same size category although no quantitative assessments have been conducted outside of MT-Bench experiments. Veering away from generative use cases, CroissantLLM has also been adapted as an embedding model and recent work has since been conducted in assessing confidence in the retrieval scores associated to such models .

Foundation Model Transparency Index

To assess the transparency of our work, we evaluate our model through the Stanford Transparency Index and obtain a total score of 81 %, far ahead of proprietary models, as well as most staple open-weights models and large-scale open-source efforts (Figure 8).

Upstream. The upstream categories include data, compute, and methods dimensions. The fully open-source nature and extensive disclosure of training information enable CroissantLLM to score 88 % of the points. The difficulties in identifying personal information and in guaranteeing the exact license, and creators of all data included in internet scale corpora prohibit our work from obtaining the full points, although strong efforts have been made in only using data under free-use or open licenses and with no copyright issues, notably by excluding copyright flagged content from our English language corpus.

Model. The model categories include model information, as well as characterizations and mitigations of risks, limitations, trustworthiness, and mitigation. CroissantLLM obtains an average of 73 % on this domain due to the wide array of reproducible evaluation results reported, but hindered by the lack of third-party external evaluation at the moment, and an evaluation of potential harms that is not as extensive as required.

Downstream. Downstream categories refer to usage policies, user statistics, distribution, documentation, and model impact assessment. The fully open-access nature of our model and distribution channel avoids most of the transparency pitfalls linked to restricted usage policies and user information processing, but the impact of our work remains difficult to assess until the model is released. The aggregated score for this category is 80 %.

Aggregated FMTI scores by major dimension of transparency. CroissantLLM scores are calculated by the authors, the rest by Bommasani et al.
Figure 8. Aggregated FMTI scores by major dimension of transparency. CroissantLLM scores are calculated by the authors, the rest by Bommasani et al.

Post-Mortem & Lessons learned

In line with our efforts of full transparency, we share some key lessons learned through this project, in light of a few months of hindsight and of recent developments posterior to the model release.

Tokenizer. In this work, we have developed our own tokenizer, inspired by Llama2 but fitted to a bilingual French-English corpus to improve fertility on both languages. Several key decisions can impact tokenizer design. The vocabulary size will have large impacts on the parameter count of the model, potentially making embedding parameters a large share of the total parameter count especially for smaller models . One should experiment with the performance and memory trade-offs of using larger vocabulary sizes. Furthermore, non-standard tokenizers are likely to be less supported by standard inference frameworks, complicating model adoption. Finally, we believe hand-designing a subset of tokens that are known to be useful at inference, for particular tasks, or have semantic coherence (common numbers, punctuation patterns, tokens corresponding to multiple choice templates or code syntax) probably degrades fertility but is useful in the long run from a performance and usability perspective.

Data Mix. It was always understood the data quality and quantity of the pretraining mix had large impacts on pretraining. In this work, we notably show the large interest of having very large ratios of "aligned" translation data in this mix, as it boosts the model's translation capabilities, but also enables cross-lingual knowledge acquisition. Since CroissantLLM's release, it has become understood that certain types of data heavily boost model reasoning capabilities or performances on benchmarks. Typically, , put a particular emphasis in including math and reasoning heavy corpora in the training corpus. reports 24 point gains on MMLU scores by including more knowledge and reasoning rich sources to their training set, as well as improving data quality filtering. In order to source such data with highly "educative" values several approaches are possible, including distilling large model knowledge into fast text classifiers to identify high-quality content from Internet scale data , or even generating diverse and high quality synthetic text using LLMs . More generally, better pretraining and annealing data leads to better models, and many recent datasets would probably lead to better performance nowadays than the datamix we had constituted. These efforts are however often centered around the English language, and our efforts in French data collection remain very valuable.

Scheduler. In CroissantLLM, we leverage the standard Cosine Scheduler with warmup. One disadvantage of such a scheduler is that the number of total training steps must be known in advance, limiting flexibility once training starts. Inspired by the Vision literature , the MiniCPM model and several models since then have later confirmed the interest of infinite or WSD (Warmup-Stable-Decay) schedulers. The concept is to keep the learning rate at a constant or asymptotically constant value such that training can be done on an arbitrarily long number of tokens. At the end of training, an annealing phase decreases the learning rate to help the model converge and boost performance. MiniCPM has further uncovered that annealing on a higher quality data mix is a very efficient strategy. Typically, by including knowledge-rich or reasoning-heavy data in this final training phase, large benchmark improvements are obtained. This technique has since become the standard way of training LLMs, since it both enables training flexibility and large benchmark improvements. We believe running an annealing phase on the CroissantLLM model would have largely boosted performance, especially if done with carefully selected high-quality data in French and English.

Leveraging Larger LLMs. As made clear in this work, while training with a linguistically balanced dataset helps multilingual performance, the number of model parameters remains the strongest performance factor, and given our compute budget, better performance could be obtained by training a larger model on less non-English data. This is perfectly in line with the findings of with Chinchilla scaling laws. Models trained to be small and inference efficient such as CroissantLLM could however benefit from larger multilingual LLMs in multiple ways. Beyond the ever-growing use of LLM-generated data during pretraining , small models can also benefit from larger LLMs as critics for alignment processes , but also to replace stochastic weight initialization by starting off from pruned versions of larger models . Obviously, these strategies entail significantly larger resource requirements to train the larger models to begin with.

Model Release. Training a model is one thing, enabling people to use it is another. Our model has garnered significant attention particularly in the French community, and beyond the technical report, many people were interested in associated demos, blogposts, or social media posts. In retrospect, some barriers of entry remained for users with limited technical skills. The lack of day 0 support from third party inference libraries (Ollama, llama.cpp, MLCChat) has led to some community-proposed implementations that were suboptimal, leading to performance hits. Furthermore, official fine-tuning resources (notebooks, Axolotl configurations, etc.) could have largely helped users adapt CroissantLLM for specific use cases and drive adoption. Considering the downstream uses of the model and its utilization seems paramount in the construction of such a project.

Building a model is an iterative process. Building upon the CroissantLLM work and the above listed insights, our team has since been able to train stronger models . The field evolves quickly and getting everything right the first time is not easy. To build very good models, it seems important to us to construct the foundations iteratively, getting things to work at a small scale first, collecting some practical real-world feedback and keeping a core model development team stable.

Broader Impact Statement

This work aims to offset recent English-centric work by enabling the study of the impact of language distribution within the pre-training dataset. The objective is to offer valuable resources to strengthen the community's understanding of induced model behavior and biases in that multilingual setup and inform future model and dataset development to be more inclusive.

Model and Resource Release. The models and all related artifacts are released openly on the CroissantLLM HuggingFace organization under an MIT license. No usage restrictions are imposed on users whatsoever. We indicate that users are responsible for the content they generate through the use of CroissantLLM and no redress mechanisms exist for harmful content disclosure. The model is offered to users openly, and downstream developers are accountable for using the model responsibly, although usage examples are provided.

Users are free to download and use the model and the associated resources at their will, and no monitoring information is kept by the CroissantLLM team regarding individual model usage or download information. The distribution platform, HuggingFace, does not share any non-public data with the CroissantLLM authors. Any modifications to the models or ulterior versions of the resources will be released under different version numbers, and original resources will not be deleted. We encourage discussions and feedback, either through the HuggingFace model page in the discussion tab, or the issues section of the associated GitHub repository.

Risk mitigation. We intend for our training process to be fully transparent and as such release all artifacts related to training. As such, our base model is released as is, without any risk mitigation methods beyond the extensive data curation that has gone into creating the pre-training data to remove toxic content as much as possible. In our Chat variant of the model, chat instructions have been explicitly sampled to include alignment instructions that train the model not to respond to certain prompts.

Data Leakage. Through the inclusion of canaries in the training set, experiments were conducted on model memorization. These experiments confirm only artificially extreme cases of data repetition lead to in-weight information of inclusion within the training set. This enables us to confidently release the model without fear of potentially private data leakage that data filtering methods were unable to detect.

Risk Assessment. Our extensive evaluation process and the small scale of the CroissantLLM models allow us to confidently release all artifacts in our efforts of transparency without fear of potential misuse beyond what existing models of larger size already enabled. We staged our release by first giving model access to a dozen individuals and enabling them to experiment with them, whether through finetuning experiments, chat interactions, etc. Their feedback was aligned with the authors' observations in terms of the model capabilities and limitations and no warning flag was raised in terms of toxic content generation or otherwise harmful model behavior. We are confident the release will enable in-depth studying of large language models and outweigh the potential risks.

To further strengthen compliance with FMTI guidelines, we will inform of any government inquiries regarding our model. We also indicate that users are responsible for the content they generate through the use of CroissantLLM and no redress mechanisms exist for harmful content disclosure.

References

  1. Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison, Duarte M. Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro H. Martins, Antoni Bigata Casademunt, François Yvon, André F. T. Martins, Gautier Viaud, Céline Hudelot, Pierre Colombo (2025). CroissantLLM: A Truly Bilingual French-English Language Model. Source ↗
  2. Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Šaško, Quentin Lhoest, Angelina McMillan-Major, Gerard Dupont, Stella Biderman, Anna Rogers, Loubna Ben allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel, Leon Weber, Manuel Muñoz, Jian Zhu, Daniel Van Strien, Zaid Alyafeai, Khalid Almubarak, Minh Chien Vu, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Alexandra Luccioni, Yacine Jernite (2023). The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset.
  3. Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, Iryna Gurevych (2021). BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv. Source ↗
  4. Niklas Muennighoff, Nouamane Tazi, Loic Magne, Nils Reimers (2022). MTEB: Massive Text Embedding Benchmark. arXiv. Source ↗
  5. OpenAI (2023). GPT-4 Technical Report. arXiv. Source ↗
  6. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, others (2023). Llama: Open and efficient foundation language models.
  7. Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, Thomas Scialom (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. Source ↗
  8. Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, Tianhang Zhu (2023). Qwen Technical Report.
  9. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, William El Sayed (2023). Mistral 7B. Source ↗
  10. Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, William El Sayed (2024). Mixtral of Experts.
  11. Rishi Bommasani, Kevin Klyman, Shayne Longpre, Sayash Kapoor, Nestor Maslej, Betty Xiong, Daniel Zhang, Percy Liang (2023). The Foundation Model Transparency Index. Source ↗
  12. Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger, Dylan Hadfield-Menell (2024). Black-Box Access is Insufficient for Rigorous AI Audits.
  13. Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, Ethan Perez (2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.
  14. NewYorkTimes (2023). The Times sues OpenAI and Microsoft over A.I. use of copyrighted work.
  15. Pamela Samuelson (2023). Generative AI meets copyright. American Association for the Advancement of Science.
  16. Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa (2023). The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants.
  17. Emily M Bender, Timnit Gebru, Angelina McMillan-Major, Shmargaret Shmitchell (2021). On the dangers of stochastic parrots: Can language models be too big?. Proceedings of the 2021 ACM conference on fairness, accountability, and transparency.
  18. Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, Tatsunori Hashimoto (2023). Whose Opinions Do Language Models Reflect?.
  19. Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, Nolan Dey (2023). SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. Source ↗
  20. Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, others (2022). Bloom: A 176b-parameter open-access multilingual language model.
  21. Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov, Anastasia Kozlova, Tatiana Shavrina (2022). mGPT: Few-Shot Learners Go Multilingual. arXiv. Source ↗
  22. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, Laurent Sifre (2022). Training Compute-Optimal Large Language Models.
  23. Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, William Fedus (2022). Emergent Abilities of Large Language Models. Source ↗
  24. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, Noah Fiedel (2022). PaLM: Scaling Language Modeling with Pathways.
  25. Nikhil Sardana, Jonathan Frankle (2023). Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.
  26. Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, Jingren Zhou (2023). Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. Source ↗
  27. Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, others (2022). Glm-130b: An open bilingual pre-trained model.
  28. Rachel Bawden, Hatim Bourfoune, Bertrand Cabot, Nathan Cassereau, Pierre Cornette, Marco Naguib, Aurélie Névéol, François Yvon (2024). Les modèles Bloom pour le traitement automatique de la langue française. Source ↗
  29. Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, André F. T. Martins (2024). EuroLLM: Multilingual Language Models for Europe. Source ↗
  30. Matthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de Montjoye (2024). Copyright Traps for Large Language Models. 41st International Conference on Machine Learning (ICML 2024). Source ↗
  31. Igor Shilov, Matthieu Meeus, Yves-Alexandre de Montjoye (2024). Mosaic Memory: Fuzzy Duplication in Copyright Traps for Large Language Models. Source ↗
  32. Matthieu Meeus, Igor Shilov, Shubham Jain, Manuel Faysse, Marek Rei, Yves-Alexandre de Montjoye (2025). SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It). (Best Paper) IEEE Conference on Secure and Trustworthy Machine Learning (SaTML 2025). Source ↗
  33. Roberto Navigli, Simone Conia, Björn Ross (2023). Biases in Large Language Models: Origins, Inventory, and Discussion. Association for Computing Machinery. Source ↗
  34. Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, Colin Raffel (2023). Scaling Data-Constrained Language Models.
  35. Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, Benoit Sagot (2022). Towards a Cleaner Document-Oriented Multilingual Crawled Corpus. Proceedings of the Thirteenth Language Resources and Evaluation Conference. Source ↗
  36. Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel (2021). mT5: A massively multilingual pre-trained text-to-text transformer.
  37. Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, Thien Huu Nguyen (2023). CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages.
  38. Michael Hart (1971). Project gutenberg. Source ↗
  39. Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, Dawn Song (2019). The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks. 28th USENIX Security Symposium (USENIX Security 19). Source ↗
  40. Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, Harm de Vries (2023). StarCoder: may the source be with you!.
  41. Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, Harm de Vries (2022). The Stack: 3 TB of permissively licensed source code.
  42. Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, Oriol Vinyals (2022). Competition-level code generation with AlphaCode. Source ↗
  43. Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, Yin Tat Lee (2023). Textbooks Are All You Need II: phi-1.5 technical report.
  44. Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, Yonghui Wu (2023). PaLM 2 Technical Report.
  45. Eleftheria Briakou, Colin Cherry, George Foster (2023). Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Source ↗
  46. Jörg Tiedemann (2012). Parallel Data, Tools and Interfaces in OPUS. Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12).
  47. Gema Ram\'-Sánchez, Jaume Zaragoza-Bernabeu, Marta Bañón, Sergio Ortiz Rojas (2020). Bifixer and Bicleaner: two open-source tools to clean your parallel data. Proceedings of the 22nd Annual Conference of the European Association for Machine Translation. Source ↗
  48. Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, André F. T. Martins (2022). CometKiwi: IST-Unbabel 2022 Submission for the Quality Estimation Shared Task. Proceedings of the Seventh Conference on Machine Translation (WMT). Source ↗
  49. Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, Yunfeng Liu (2023). RoFormer: Enhanced Transformer with Rotary Position Embedding.
  50. Rico Sennrich, Barry Haddow, Alexandra Birch (2016). Neural Machine Translation of Rare Words with Subword Units.
  51. Taku Kudo (2018). Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates.
  52. Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, Iryna Gurevych (2021). How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.
  53. Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei (2020). Scaling Laws for Neural Language Models. Source ↗
  54. Patrick Fernandes, Behrooz Ghorbani, Xavier Garcia, Markus Freitag, Orhan Firat (2023). Scaling Laws for Multilingual Neural Machine Translation. Proceedings of the 40th International Conference on Machine Learning. Source ↗
  55. Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki Heinonen, Aija Vahtola, Samuel Antao, Sampo Pyysalo (2023). FinGPT: Large Generative Models for a Small Language. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Source ↗
  56. Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, others (2022). Crosslingual generalization through multitask finetuning.
  57. Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, Luke Zettlemoyer (2022). OPT: Open Pre-trained Transformer Language Models.
  58. Alexandra Sasha Luccioni, Sylvain Viguier, Anne-Laure Ligozat (2022). Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model.
  59. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, Yejin Choi (2019). HellaSwag: Can a Machine Really Finish Your Sentence?.
  60. Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, Yejin Choi (2019). PIQA: Reasoning about Physical Commonsense in Natural Language.
  61. Johannes Welbl, Nelson F. Liu, Matt Gardner (2017). Crowdsourcing Multiple Choice Science Questions.
  62. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord (2018). Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.
  63. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, Ion Stoica (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.
  64. Martin d'Hoffschmidt, Wacim Belblidia, Tom Brendlé, Quentin Heinrich, Maxime Vidal (2020). FQuAD: French Question Answering Dataset.
  65. Chin-Yew Lin (2004). ROUGE: A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out. Source ↗
  66. Moussa Kamal Eddine, Antoine J-P Tixier, Michalis Vazirgiannis (2020). BARThez: a Skilled Pretrained French Sequence-to-Sequence Model.
  67. Duarte M Alves, Nuno M Guerreiro, João Alves, José Pombal, Ricardo Rei, José GC de Souza, Pierre Colombo, André FT Martins (2023). Steering Large Language Models for Machine Translation with Finetuning and In-Context Learning.
  68. Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu (2002). Bleu: a Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. Source ↗
  69. Matt Post (2018). A Call for Clarity in Reporting BLEU Scores. Proceedings of the Third Conference on Machine Translation: Research Papers. Source ↗
  70. Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, André F. T. Martins (2022). COMET-22: Unbabel-IST 2022 Submission for the Metrics Shared Task. Proceedings of the Seventh Conference on Machine Translation (WMT). Source ↗
  71. NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, Jeff Wang (2022). No Language Left Behind: Scaling Human-Centered Machine Translation.
  72. Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federman, Dmitriy Genzel, Francisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, Sylwia Tur (2020). TICO-19: the Translation Initiative for Covid-19.
  73. Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O'Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, Oskar van der Wal (2023). Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.
  74. Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, Wei Lu (2024). TinyLlama: An Open-Source Small Language Model.
  75. Antoine Simoulin, Benoit Crabbé (2021). Un modèle Transformer Génératif Pré-entrainé pour le \_\_\_\_\_\_ français. Traitement Automatique des Langues Naturelles. Source ↗
  76. Julien Launay, Elena Tommasone, Baptiste Pannier, François Boniface, Amélie Chatelain, Alessandro Cappelli, Iacopo Poli, Djamé Seddah (2021). PAGnol: An Extra-Large French Generative Model.
  77. Manuel Faysse, Gautier Viaud, Céline Hudelot, Pierre Colombo (2023). Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial Applications. (Oral, EMNLP 2023) The 2023 Conference on Empirical Methods in Natural Language Processing. Source ↗
  78. Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, Quoc V. Le (2022). Finetuned Language Models Are Zero-Shot Learners.
  79. Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, Bowen Zhou (2023). Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.
  80. Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, Yuntian Deng (2024). (InThe)WildChat: 570K ChatGPT Interaction Logs In The Wild. The Twelfth International Conference on Learning Representations. Source ↗
  81. Cyrile Delestre (2023). BloomZ SFT Chat. Source ↗
  82. Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, Pascale Fung (2023). Survey of Hallucination in Natural Language Generation. Association for Computing Machinery (ACM). Source ↗
  83. Nuno M Guerreiro, Duarte M Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, André FT Martins (2023). Hallucinations in large multilingual translation models. MIT Press One Broadway, 12th Floor, Cambridge, Massachusetts 02142, USA~….
  84. Hippolyte Gisserot-Boukhlef, Manuel Faysse, Emmanuel Malherbe, Céline Hudelot, Pierre Colombo (2024). Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism. Source ↗
  85. Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, Anton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, Matt Hoffman, Shantanu Thakoor, Jean-Bastien Grill, Behnam Neyshabur, Olivier Bachem, Alanna Walton, Aliaksei Severyn, Alicia Parrish, Aliya Ahmad, Allen Hutchison, Alvin Abdagic, Amanda Carl, Amy Shen, Andy Brock, Andy Coenen, Anthony Laforge, Antonia Paterson, Ben Bastian, Bilal Piot, Bo Wu, Brandon Royal, Charlie Chen, Chintu Kumar, Chris Perry, Chris Welty, Christopher A. Choquette-Choo, Danila Sinopalnikov, David Weinberger, Dimple Vijaykumar, Dominika Rogozińska, Dustin Herbison, Elisa Bandy, Emma Wang, Eric Noland, Erica Moreira, Evan Senter, Evgenii Eltyshev, Francesco Visin, Gabriel Rasskin, Gary Wei, Glenn Cameron, Gus Martins, Hadi Hashemi, Hanna Klimczak-Plucińska, Harleen Batra, Harsh Dhand, Ivan Nardini, Jacinda Mein, Jack Zhou, James Svensson, Jeff Stanway, Jetha Chan, Jin Peng Zhou, Joana Carrasqueira, Joana Iljazi, Jocelyn Becker, Joe Fernandez, Joost van Amersfoort, Josh Gordon, Josh Lipschultz, Josh Newlan, Ju-yeong Ji, Kareem Mohamed, Kartikeya Badola, Kat Black, Katie Millican, Keelin McDonell, Kelvin Nguyen, Kiranbir Sodhia, Kish Greene, Lars Lowe Sjoesund, Lauren Usui, Laurent Sifre, Lena Heuermann, Leticia Lago, Lilly McNealus, Livio Baldini Soares, Logan Kilpatrick, Lucas Dixon, Luciano Martins, Machel Reid, Manvinder Singh, Mark Iverson, Martin Görner, Mat Velloso, Mateo Wirth, Matt Davidow, Matt Miller, Matthew Rahtz, Matthew Watson, Meg Risdal, Mehran Kazemi, Michael Moynihan, Ming Zhang, Minsuk Kahng, Minwoo Park, Mofi Rahman, Mohit Khatwani, Natalie Dao, Nenshad Bardoliwalla, Nesh Devanathan, Neta Dumai, Nilay Chauhan, Oscar Wahltinez, Pankil Botarda, Parker Barnes, Paul Barham, Paul Michel, Pengchong Jin, Petko Georgiev, Phil Culliton, Pradeep Kuppala, Ramona Comanescu, Ramona Merhej, Reena Jana, Reza Ardeshir Rokni, Rishabh Agarwal, Ryan Mullins, Samaneh Saadat, Sara Mc Carthy, Sarah Cogan, Sarah Perrin, Sébastien M. R. Arnold, Sebastian Krause, Shengyang Dai, Shruti Garg, Shruti Sheth, Sue Ronstrom, Susan Chan, Timothy Jordan, Ting Yu, Tom Eccles, Tom Hennigan, Tomas Kocisky, Tulsee Doshi, Vihan Jain, Vikas Yadav, Vilobh Meshram, Vishal Dharmadhikari, Warren Barkley, Wei Wei, Wenming Ye, Woohyun Han, Woosuk Kwon, Xiang Xu, Zhe Shen, Zhitao Gong, Zichuan Wei, Victor Cotruta, Phoebe Kirk, Anand Rao, Minh Giang, Ludovic Peran, Tris Warkentin, Eli Collins, Joelle Barral, Zoubin Ghahramani, Raia Hadsell, D. Sculley, Jeanine Banks, Anca Dragan, Slav Petrov, Oriol Vinyals, Jeff Dean, Demis Hassabis, Koray Kavukcuoglu, Clement Farabet, Elena Buchatskaya, Sebastian Borgeaud, Noah Fiedel, Armand Joulin, Kathleen Kenealy, Robert Dadashi, Alek Andreev (2024). Gemma 2: Improving Open Language Models at a Practical Size. Source ↗
  86. Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Caio César Teodoro Mendes, Weizhu Chen, Vishrav Chaudhary, Parul Chopra, Allie Del Giorno, Gustavo de Rosa, Matthew Dixon, Ronen Eldan, Dan Iter, Amit Garg, Abhishek Goswami, Suriya Gunasekar, Emman Haider, Junheng Hao, Russell J. Hewett, Jamie Huynh, Mojan Javaheripi, Xin Jin, Piero Kauffmann, Nikos Karampatziakis, Dongwoo Kim, Mahoud Khademi, Lev Kurilenko, James R. Lee, Yin Tat Lee, Yuanzhi Li, Chen Liang, Weishung Liu, Eric Lin, Zeqi Lin, Piyush Madan, Arindam Mitra, Hardik Modi, Anh Nguyen, Brandon Norick, Barun Patra, Daniel Perez-Becker, Thomas Portet, Reid Pryzant, Heyang Qin, Marko Radmilac, Corby Rosset, Sambudha Roy, Olatunji Ruwase, Olli Saarikivi, Amin Saied, Adil Salim, Michael Santacroce, Shital Shah, Ning Shang, Hiteshi Sharma, Xia Song, Masahiro Tanaka, Xin Wang, Rachel Ward, Guanhua Wang, Philipp Witte, Michael Wyatt, Can Xu, Jiahang Xu, Sonali Yadav, Fan Yang, Ziyi Yang, Donghan Yu, Chengruidong Zhang, Cyril Zhang, Jianwen Zhang, Li Lyna Zhang, Yi Zhang, Yue Zhang, Yunan Zhang, Xiren Zhou (2024). Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. arXiv. Source ↗
  87. Reka Team, Aitor Ormazabal, Che Zheng, Cyprien de Masson d'Autume, Dani Yogatama, Deyu Fu, Donovan Ong, Eric Chen, Eugenie Lamprecht, Hai Pham, Isaac Ong, Kaloyan Aleksiev, Lei Li, Matthew Henderson, Max Bain, Mikel Artetxe, Nishant Relan, Piotr Padlewski, Qi Liu, Ren Chen, Samuel Phua, Yazheng Yang, Yi Tay, Yuqi Wang, Zhongkai Zhu, Zhihui Xie (2024). Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models. Source ↗
  88. An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, Zhihao Fan (2024). Qwen2 Technical Report. Source ↗
  89. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaoqing Ellen Tan, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aaron Grattafiori, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alex Vaughan, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Franco, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, Danny Wyatt, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Firat Ozgenel, Francesco Caggioni, Francisco Guzmán, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Govind Thattai, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Karthik Prasad, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kun Huang, Kunal Chawla, Kushal Lakhotia, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Maria Tsimpoukelli, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikolay Pavlovich Laptev, Ning Dong, Ning Zhang, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Rohan Maheswari, Russ Howes, Ruty Rinott, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Kohler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vítor Albiero, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaofang Wang, Xiaojian Wu, Xiaolan Wang, Xide Xia, Xilun Wu, Xinbo Gao, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yuchen Hao, Yundi Qian, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao (2024). The Llama 3 Herd of Models. Source ↗
  90. Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi (2024). OLMo: Accelerating the Science of Language Models.
  91. Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, Julien Launay (2023). The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.
  92. Loubna Ben Allal, Anton Lozhkov, Guilherme Penedo, Thomas Wolf, Leandro von Werra (2024). Cosmopedia. Source ↗
  93. Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas Beyer (2021). Scaling Vision Transformers. arXiv. Source ↗
  94. Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, Xinrong Zhang, Zheng Leng Thai, Kaihuo Zhang, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai Li, Zhiyuan Liu, Maosong Sun (2024). MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies. Source ↗
  95. Alexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro Von Werra, Martin Jaggi (2024). Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations. Source ↗
  96. Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, others (2025). EuroBERT: Scaling Multilingual Encoders for European Languages. Conference on Language Modeling (COLM 2025). Source ↗