Home / Blog / Understanding MTQE: How MTQE transforms the machine translation industry

Understanding MTQE: How MTQE transforms the machine translation industry

Localization workflows & operations
Tanja Schöllhammer
Content Marketer

Last updated

7/23/2026

Read time

7 min

Best for

Managers

How MTQE transforms the machine translation industry

Machine translation and AI engines like Google Translate, DeepL, ChatGPT and Claude make neural machine translation accessible to hundreds of millions of users worldwide. They are part of the daily routine of translating entertainment content, providing support during travels, etc. This technology has already become the core of the translation industry as it allows a significant volume of information to be translated in a fraction of the time required for human translation.

Despite its benefits, the quality of machine translation results still needs improvement. Errors can occasionally occur, and the rarer the language pair, the more issues can arise.

Machine translation quality estimation (MTQE) helps organizations predict whether an AI-generated translation is likely to be usable before a human reviews it. Unlike traditional translation quality evaluation metrics such as BLEU or METEOR, MTQE does not require a reference translation.

What is MTQE?

Machine translation quality estimation (MTQE) predicts whether an AI generated translation is likely to be usable before human review. Unlike reference based metrics such as BLEU or METEOR, it does not compare translations against a reference translation.

Organizations use MTQE to reduce post editing effort, prioritize human review, improve translation workflows, and automate quality checks across large volumes of multilingual content.

Let’s look at the prominent cases where automated machine translation quality estimation is required and can significantly assist.

  • Real time translation systems, including live subtitles and multilingual chat applications.

  • High-volume localization projects, where reviewing every translation manually is impractical.

  • Customer support and AI chatbots, where translation quality directly affects customer experience.

  • Website localization, document translation, and ecommerce catalogs, where thousands of pages must be translated efficiently.

  • Cross Language Information Retrieval (CLIR) systems, such as multilingual search engines and digital libraries.

Now, let’s check how MTQE approaches work and how they can cover the cases above.

MTQE vs. MTQEVAL

While MTQE predicts translation quality, MTQEVAL measures translation quality against a known correct translation. The two approaches solve different problems and are frequently combined in enterprise localization workflows.

The MTQE methods are based on predicting the quality of the machine translation results. Under the hood of this system can be neural networks, statistical and machine learning methods, or hybrid solutions. Modern MTQE systems increasingly rely on transformer models and large language models (LLMs) to evaluate translation quality more accurately than earlier statistical approaches. The MTQE’s main task is to evaluate how logically and accurately a translation matches the source text. This approach checks the meaning alignment between the source and the translated text and the accuracy of using key elements such as numbers, names, or terms and can be split into:

  • Evaluate the syntactic accuracy (checks whether the translation complies with the grammatical rules of the target language)

  • Evaluate the semantic accuracy (checks how accurately words and phrases are translated in terms of meaning)

  • Сhecks the usage of the terms.

  • Сhecks the entire context.

In simple words, the MTQE “understands” the meaning of the source text and the translated result and can then provide a rating based on numerical values ​​or categories (e.g., “excellent translation,” “needs improvement,” "poor quality).

MTQE pros:

  • Saved time on reference translation creation - the models can work without high-quality references, which reduces manual translation.

  • Raised speed - the MTQE can be implemented for real-time usage, which is vital for online translation systems.

  • Easy scalability - the increased number of content for translation wouldn’t somehow affect the process as there is no need to prepare a reference translation.

  • Simplifies processing of the big data: MTQE helps quickly determine which translation needs revision and which can be already used. This logic allows companies with a vast content volume to focus on the text segments with low-quality rates instead of proofreading all the content.

MTQE cons:

  • Limited accuracy. Since MTQE does not use reference translation, it may be less accurate than the reference-based methods.

  • Training and interpretation challenges. MTQE models require high-quality training, and sometimes, it can be difficult to understand which aspects of the translation cause a high or low-quality score.


Conversely, the MTQEVAL always requires the reference translation as a base for comparison. For example, let’s overview a few popular approaches - BLEU and METEOR and how they work.

BLEU and METEOR are among the best known automatic machine translation evaluation metrics, although newer neural evaluation metrics such as COMET, which generally show a stronger correlation with human quality judgments than traditional metrics like BLEU.

BLEU (Bilingual Evaluation Understudy) compares the translated text with reference translation based on a list of criteria (analyzes sequences of words/characters, the length (brevity penalty)). The entire BLEU logic is built around “n-gram” - a sequence of n elements.

For instance, let’s assume that we have:

  • Reference translation: “This is a blog about localization”

  • Translation: “This is the blog about localization.”

The 1-grams are - “This” “is” “the” “blog” “about” “localization” (5 matches out of 6.)

The 2-grams are -“This is” “is the” “the blog” “blog about” “about localization” (4 matches out of 5.)

So, BLEU combines all the results of 1-gram, 2-grams, 3-grams, and 4-grams to provide the final assessment (e.g., if a translation contains 20 n-grams and 10 of them match the standard, the accuracy is 50%).

METEOR (Metric for Evaluation of Translation with Explicit Ordering) is an “improved BLEU” because it considers synonyms, roots, and paraphrases in addition to the n-gram. Thus, this approach is more sensitive to the meaning of the translation but still requires the reference translation for work.

MTQEVAL approach: Pros

  • High accuracy. The MTQEVAL methods typically provide a more accurate and objective review of translation quality, as they have an exact example of the translation needed.

  • Flexible metrics overview. Approaches like BLEU and METEOR have transparent criteria, making it easier to see all aspects and allowing for a more in-depth assessment.

MTQEVAL approach: Cons.

  • Expensive and time-consuming. As these methods require reference translation it means that users have to prepare these references and spend additional time and budgets on them.

  • High dependence on the quality of the reference translation. The issues in the reference text will lead to continuous multiplying them in the subsequent translations.

  • Not suitable for real-time processes. Besides preparing the reference translation the process is resource-intensive, especially for large volumes of text.

For convenience, take a look at the table below:

Criteria

MTQE

MTQEVAL

Reference needed

No

Yes

Purpose

Predicting quality

Comparing with an ideal standard

Metrics

Neural, statistical, and hybrid models

BLEU, METEOR, etc

Speed

Fast

Slow

Accuracy

Moderate

High

Typical use case

Translation processes automation

Accurate comparisons

Note: It is a widespread practice when both approaches are combined. Hence, the MTQE can help with a quick quality assessment, while MTQEVAL can be used for a more in-depth analysis.

To sum up

Machine Translation Quality Estimation helps organizations scale multilingual content while reducing unnecessary human review. Combined with glossaries, translation memories, terminology management, and modern AI translation engines, MTQE supports AI translation workflows, software localization, website localization, and multilingual content management while reducing manual review effort.

Getting perfect translation quality without compromises is the number one target for any business. At LingoHub, we understand the business needs and that implementing custom models that evaluate the quality can be unprofitable. That’s why providing the perfect translation from the start is better than improving it afterward. To simplify the process, we offer the following:

  • Machine and AI translation that combine the top engines on the market - Google Translate, DeepL, Amazon Translate, ChatGPT, Mistral, Claude, etc.;

  • Glossary and style guide that affect the machine translation results;

  • Translation memory, which allows reusing of the previous translations to reduce the repetitive tasks;

  • Post-editing score to evaluate the translators’ efforts fair and transparent;

  • And many more.

Book a demo call with our team for more information, or try the 14-day free trial and check everything yourself.


Frequently asked questions

Is MTQE more accurate than BLEU?

MTQE and BLEU solve different problems. MTQE predicts translation quality without a reference translation, while BLEU measures similarity between a machine translation and a reference translation. Organizations often combine both methods in production localization workflows.

Does MTQE replace human translators?

No. MTQE identifies translations that are likely to require review, helping translators focus their effort where it has the greatest impact.

Which industries use MTQE?

MTQE is commonly used in software localization, e-commerce, customer support, gaming, documentation, and enterprise content localization.

Related articles