#evaluation

3 notes

Few-shot Translation Evaluation

2023-12-22 · 2 min read

#translation #ai

In the context of a translation project, where we aim to utilize only the resources at hand, let’s consider a simplified approach to the evaluation stage. This approach would involve the following steps:

  1. Start by taking the draft translation and identify the translation pairs that are most similar to the paired prediction.
  2. Evaluate the draft translation for any potential misuse of words.
    1. One possible method could be to use a rules-based matcher that identifies any tokens in the source+prediction pair that only occur in EITHER the source OR the target in any of the similar examples.
    2. It could also be beneficial to examine whether any words seem to be mistranslated based on the gold standard examples.
    3. To aid this process, aligning these examples by phonetics, orthography, and apparent semantic similarities could be helpful.
  3. Once the evaluation is complete, pass the task back to the translation bot, providing special notes about any words that appear to be mistranslated.

In essence, the process can be summarized as: [ most similar examples ] —> [ prediction ] —> [ most similar examples ]

The goal is to identify the most similar examples to the prediction, and then use those examples to evaluate the prediction. This process can be repeated iteratively, with the bot learning from the evaluation and then generating a new prediction and a new set of examples that will help the initial bot improve its output.

Permalink →

#translation #ai

Here’s an idea I had for translation quality checking.

The process begins by prompting for a translation using n examples and a directive to complete the translation pairs. The entire prompt, along with the AI’s completion, is then passed to the evaluator. The evaluator’s task is to determine which translation was generated by the AI. This may require fine-tuning a discriminator model on the gold-standard pairs.

The concept, inspired by Generative Adversarial Networks (GANs), involves an evaluator attempting to distinguish AI-drafted translations from a list of real translations. While it doesn’t function exactly like a GAN, the idea is to use a similar principle for evaluation.

The method could potentially be used to refine the prompts or the models themselves, by replacing the evaluator with a separate AI. Initial implementation will use Language Models (LLMs) to gauge the viability of the concept. If successful, this approach could be extended to train GANs for specific languages.

Permalink →

#translation #ai

In the realm of machine translation, achieving the highest level of accuracy is paramount. One approach to improve the precision of translations is through a method called “Token Matching Evaluation”. This method centres around the use of valid tokens, which are words or phrases that have appeared in sample translation pairs.

The Process

  1. Populate a list of valid tokens for back-translation: The first step is to create a list of valid tokens. These tokens are the words or phrases that the model is allowed to use for back-translations.

  2. Back-translate using valid tokens: The model is instructed to back-translate any given word only from the list of valid tokens. For example, using the {{select options=valid_tokens n=6}} syntax, the model selects a specific value from a chunk of 6 valid tokens.

    • If a word hasn’t even been translated once by a human, then any attempt would effectively be indistinguishable from a hallucination. Therefore, it is important that the valid tokens correspond to reality. For instance, if you need to translate the English word ‘Jerusalem’, you would find 5 sentences with ‘Jerusalem’ in them and those sentences would have Abanyom renderings, like ‘Yerusalem’. The only valid renderings must be contained in the sample sentences.

    • Any tokens not in the set of valid tokens are simply retained in their translated form in [square brackets] to indicate they have not been back-translated because they are unknown. This could be facilitated by having an agent send out messages to the translation team via WhatsApp, or using a simple gamified app.

  3. Gather samples based on token coverage: It might be beneficial to gather samples based on whether or not they help cover the missing tokens. Start with an initial set of 5 semantically similar examples, followed by a second set of 5 more sentences that include the English glosses not represented in the first five examples.

By following these steps, the output translation will more likely align with the sample translations provided in the prompt, thereby enhancing the accuracy of the translation.

Permalink →