Jungseob Lee Publications

Cross-Lingual Optimization for Language Transfer in Large Language Models

ACL 2025 (Oral)

Jungseob Lee*, Seongtae Hong*, Hyeonseok Moon, Heuiseok Lim†

Korea University

*Equal contribution †Corresponding author

CLO transfers an English-centric LLM to a target language while preserving its English capabilities, using publicly available English SFT data and a translation model.

Two-part diagram. Top, cross-lingual dataset preparation: English instruction and response pairs are translated into the target language, and each instruction is paired with a preferred response in its own language and a rejected response in the other language. Bottom, cross-lingual optimization: an LLM is trained with supervised fine-tuning on target-language pairs and with a cross-lingual adaptation loss in both directions, combined into one CLO loss.
Overview of cross-lingual dataset preparation and optimization method. The process begins with translating English \((x_{\text{en}}, y_{\text{en}})\) pairs into a target language to create a cross-lingual dataset. This process results in the creation of \((x_\ell, y_\ell)\) pairs in the target language. The optimization is performed using a combined loss \(\mathcal{L}_{\text{CLO}}\).

Abstract

Adapting large language models to other languages typically employs supervised fine-tuning (SFT) as a standard approach. However, it often suffers from an overemphasis on English performance, a phenomenon that is especially pronounced in data-constrained environments. To overcome these challenges, we propose Cross-Lingual Optimization (CLO) that efficiently transfers an English-centric LLM to a target language while preserving its English capabilities. CLO utilizes publicly available English SFT data and a translation model to enable cross-lingual transfer. We conduct experiments using five models on six languages, each possessing varying levels of resource. Our results show that CLO consistently outperforms SFT in both acquiring target language proficiency and maintaining English performance. Remarkably, in low-resource languages, CLO with only 3,200 samples surpasses SFT with 6,400 samples, demonstrating that CLO can achieve better performance with less data. Furthermore, we find that SFT is particularly sensitive to data quantity in medium and low-resource languages, whereas CLO remains robust. Our comprehensive analysis emphasizes the limitations of SFT and incorporates additional training strategies in CLO to enhance efficiency.

Motivation

English-centric language models show limited multilingual capabilities in two ways. They may fail to comprehend certain languages, or they may understand a language but still default to communicating in English.

When posed with a Swahili question, the Llama2 Chat model fails to comprehend Swahili adequately, while the Llama3 Chat model understands the query but is unable to generate responses in Swahili. Even the fine-tuned model, despite being trained with 6,400 Swahili data, still struggles to produce appropriate outputs in Swahili.

CLO builds on the hypothesis that simultaneously enhancing target language ability and aligning it with English facilitates efficient transfer.

A Swahili query asking for a description of badminton, followed by four responses. Llama2 Chat replies in English that a Swahili word is not valid. Llama2 SFT replies in Swahili that does not make sense. Llama2 CLO replies in Swahili that badminton is a racquet sport where players hit a shuttlecock across a net. Llama3 Chat replies in English.
Example responses to a Swahili query generated by English-centric instruction models, the SFT model, and the proposed CLO model.

Method

CLO assumes a base language model, a small amount of English SFT data, and a translation model that supports the target language. The process consists of two steps.

  1. Step 1

    Cross-lingual dataset preparation

    Translate the English prompts and responses of an existing SFT dataset into the target language. For an English prompt, the English response is the chosen response and the translated response is the rejected one, which preserves English capability.

    For a target-language prompt the mapping is reversed, which encourages the model to answer in the target language by relying on its underlying English knowledge.

  2. Step 2

    Cross-lingual optimization

    A cross-lingual loss \(\mathcal{L}_{\text{CL}}\), modified from DPO, uses the paired data within the same batch to teach the correspondence between input and output languages. It is combined with an NLL loss \(\mathcal{L}_{\text{SFT}}\) computed only on target-language outputs, which mitigates the inherent English bias.

    \[\mathcal{L}_{\text{CLO}} = \lambda \cdot \mathcal{L}_{\text{SFT}} + (1 - \lambda) \cdot \mathcal{L}_{\text{CL}}\]

    The trade-off parameter \(\lambda\) is set to 0.5 for all models and languages, and only the attention layers are fine-tuned.

The experiments use 6,400 English samples from OpenAssistant, translated with the M2M100 1.2B model for a total of 12,800 samples (6,400 in English and 6,400 in the target language). They cover five models (Llama-2-7B, Llama-2-13B, Llama-3-8B, Mistral-7B-v0.1, Qwen-2.5-3B) and six languages: Chinese and German (high-resource), Korean and Indonesian (medium-resource), and Swahili and Yoruba (low-resource).

Instruction-Following Results

On AlpacaEval, CLO consistently surpasses SFT across all base models and languages, achieving a win rate exceeding 50% in all cases. Applying DPO after SFT on the same cross-lingual data only partially improves SFT and does not close the gap with CLO, a gap that is particularly pronounced in low-resource languages.

Eval LangHigh-ResourceMedium-ResourceLow-Resource
ChineseGermanKoreanIndonesianSwahiliYoruba
SFT+DPOCLOΔSFT+DPOCLOΔSFT+DPOCLOΔSFT+DPOCLOΔSFT+DPOCLOΔSFT+DPOCLOΔ
Llama-3-8B
Target59.1±1.7070.4±1.62+11.351.7±1.7654.6±1.76+2.962.5±1.7077.8±1.47+15.354.9±1.7556.4±1.75+1.565.4±1.6783.0±1.61+17.652.4±1.7664.0±1.70+11.6
English52.0±1.7664.0±1.73+12.050.1±1.7655.5±1.76+5.452.1±1.7664.4±1.69+12.351.6±1.7657.7±1.75+6.150.3±1.7661.4±1.73+11.151.2±1.7658.2±1.75+7.0
Llama-2-7B
Target59.1±1.7361.1±1.72+2.050.3±1.7659.5±1.73+9.252.5±1.7653.8±1.76+1.350.8±1.7655.8±1.75+5.064.1±1.6465.0±1.72+0.943.5±1.7467.1±1.66+23.6
English51.1±1.7655.7±1.76+4.648.4±1.7661.2±1.72+12.850.6±1.7651.0±1.76+0.449.3±1.7662.4±1.71+13.151.5±1.7655.5±1.75+4.050.3±1.7661.2±1.72+10.9
Llama-2-13B
Target59.3±1.7465.2±1.71+5.950.5±1.7653.7±1.76+3.251.6±1.7553.9±1.75+2.352.4±1.7661.7±1.71+9.353.9±1.5670.9±1.60+17.043.5±1.7567.3±1.65+23.8
English50.6±1.7650.4±1.76−0.253.0±1.7658.5±1.74+5.551.4±1.7652.3±1.76+0.952.9±1.7659.8±1.73+6.950.8±1.7655.0±1.75+4.251.1±1.7654.2±1.76+3.1
Mistral-7B-v0.1
Target56.8±1.7557.4±1.74+0.649.9±1.7650.8±1.76+0.948.5±1.7650.5±1.76+2.050.0±1.7651.1±1.76+1.135.4±1.7151.3±1.76+15.950.9±1.7651.1±1.52+0.2
English51.4±1.7657.1±1.75+5.748.6±1.7650.6±1.76+2.052.9±1.7665.2±1.71+12.352.3±1.7654.8±1.76+2.551.2±1.7664.2±1.69+13.049.4±1.7652.8±1.76+3.4
Qwen-2.5-3B
Target54.2±1.7559.9±1.73+5.753.0±1.7654.0±1.76+1.052.3±1.7662.4±1.71+10.153.8±1.7654.9±1.75+1.151.7±1.7674.7±1.53+23.045.9±1.7668.9±1.63+23.0
English50.7±1.7656.7±1.75+6.051.9±1.7654.8±1.76+2.950.8±1.7661.0±1.72+10.250.6±1.7658.0±1.75+7.451.1±1.7658.1±1.74+7.054.4±1.7657.0±1.75+2.6

Scroll sideways to see every column.

Win-rate (%) results on AlpacaEval for models fine-tuned with SFT+DPO and CLO, evaluated against their SFT baselines. Each cell reports the win rate and its standard deviation. The Δ denotes the absolute improvement of CLO over SFT+DPO.

Effect of Training Data Size

With Llama-2-7B, each model trained with a different data size is compared against the SFT model trained on 6,400 pairs. CLO improves efficiently across all three languages tested (Chinese, Korean, and Swahili), even in low-data environments. In Swahili, CLO achieves performance similar to the SFT model trained on 6,400 pairs using merely 1,600 pairs, and with 3,200 pairs it surpasses that model. SFT, by contrast, exhibits a strong dependence on the quantity of training data in medium and low-resource languages.

Six line plots of win rate against training data size from 200 to 6,400 pairs for Chinese, Korean, and Swahili models, evaluated on the target language in the top row and on English in the bottom row. The CLO curve lies above the SFT curve at most data sizes. For the Swahili model evaluated on Swahili, CLO reaches about 48 percent with 1,600 pairs and about 53 percent with 3,200 pairs, while SFT stays below 30 percent until it reaches the 6,400-pair reference at 50 percent.
Comparison of win rates between CLO and SFT on Llama-2-7B models trained with varying amounts of data, evaluated against a SFT with 6,400 pair examples on the AlpacaEval. The ‘SFT Assumed’ baseline is assigned a win rate of 50%, as it compares identical models.

BibTeX

@inproceedings{lee-etal-2025-cross,
    title = "Cross-Lingual Optimization for Language Transfer in Large Language Models",
    author = "Lee, Jungseob  and
      Hong, Seongtae  and
      Moon, Hyeonseok  and
      Lim, Heuiseok",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.734/",
    doi = "10.18653/v1/2025.acl-long.734",
    pages = "15100--15119",
    ISBN = "979-8-89176-251-0"
}