To create a true "agent," the model must be trained on both translation directions simultaneously. Research shows that models trained this way can even learn to align languages better than those trained on a single direction.
- Parallel Translation: 70,000 rows (
Dataset_English_Hindi.csv) + 5,000 high-quality rows (news-commentary-en-hi.csv). - Monolingual Alignment: ~700 rows (
hi-hi-paraphrase.csvanden-en-paraphrase.csv). These "target-only" generation tasks are critical for reducing Off-target (wrong language) and Oscillatory Hallucination (repetition) errors.
When merging your files into a single JSON for fine-tuning, you should duplicate your parallel rows to train both directions:
-
Direction A (En
$\rightarrow$ Hi): Use English as input, Hindi as output. -
Direction B (Hi
$\rightarrow$ En): Flip the same row; use Hindi as input, English as output. - Monolingual Pass: Use the paraphrase files to teach the model to stay "within" one language.
Apply the Gemma 3 chat template to distinguish the tasks.
- English to Hindi:
<start_of_turn>user\nTranslate to Hindi: {en_text}<end_of_turn>\n<start_of_turn>model\n{hi_text}<end_of_turn> - Hindi to English:
<start_of_turn>user\nTranslate to English: {hi_text}<end_of_turn>\n<start_of_turn>model\n{en_text}<end_of_turn> - Alignment (Hi):
<start_of_turn>user\nRewrite this in Hindi: {hi_text}<end_of_turn>\n<start_of_turn>model\n{hi_paraphrase}<end_of_turn> - Alignment (En):
<start_of_turn>user\nRewrite this in English: {en_text}<end_of_turn>\n<start_of_turn>model\n{en_paraphrase}<end_of_turn>
Since Gemma 3 270M is small, it can be fine-tuned very efficiently. We use QLoRA to modify the model's capacity while keeping it lightweight.
Rank (r): 16. A rank of 16-32 is optimal for a 270M model to handle the complexity of bidirectional translation without losing capacity.
-
Alpha: 32.
-
Target Modules:
["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]. Targeting all linear layers is necessary for translation tasks. -
Learning Rate:
$2 \times 10^{-4}$ . This is a stable rate for LoRA on small architectures. -
Epochs: 2–3. Scaling the number of training examples generally follows a log-linear improvement curve.
For deployment on a Snapdragon 695, INT4 quantization is mandatory to optimize memory bandwidth and ensure smooth performance.
Combine the LoRA adapters with the base model to create a standalone translator.
merged_model = model.merge_and_unload()
merged_model.save_pretrained("./merged_bidirectional_translator")Convert the PyTorch model to a LiteRT (TFLite) flatbuffer.
import ai_edge_torch
from ai_edge_torch.quantize import ptq_int4
edge_model = ai_edge_torch.convert(
merged_model,
sample_input, # Use a sample tensor of tokens for tracing
quant_config=ai_edge_torch.QuantConfig(ptq_int4())
)
edge_model.export("gemma3_270m_bidirectional_translator_int4.tflite")Bundle the .tflite model and the tokenizer.model into a single .task file using the MediaPipe LLM Bundler. This final artifact is what your mobile application will load.