Fluent isn’t the same as correct

AI translation is fast, cheap, and convincing. That’s the problem. Raw machine translation reads well even when it’s wrong, and at enterprise volume no one can review every segment in every language to find out. Quality you can’t see is quality you can’t govern.

So teams do one of two things. They over-review, and lose the speed and cost savings that made automation worth it. Or they trust the output, and carry a risk they can’t measure. Neither scales. What scales is knowing, segment by segment, how good the translation actually is, and being able to prove it.

Illustration of Phrase Quality Evaluation profiles showing customizable translation quality checks and scoring rules for different content types.

Quality evaluation is the control layer

Generating translation is no longer the hard part. Governing it is. Control and Governance is the part of the platform that turns AI translation from a gamble into a system: a standard to measure against, a score for every segment, your own rules enforced automatically, and a record of all of it.

That layer is what you can’t get from a model API alone, and it’s what makes automation safe to scale. Phrase measures each translation against the MQM 2.0 framework, holds it to the quality profile you define, and routes it accordingly, without a person reading every line. You automate on evidence rather than hope, and you keep the audit trail that regulated and brand-sensitive content depends on.

From score to action:
how Phrase manages translation quality

Phrase QPS dashboard showing automated translation quality scoring and performance metrics.

Measure – Quality Performance Score (QPS)

QPS rates every segment from 0 to 100 against the MQM 2.0 framework, so quality is a consistent number across projects, languages, and engines rather than a subjective call. Track it on dashboards in Phrase Analytics.

Dashboard showing Phrase quality profiles used to configure translation quality standards and automated evaluation settings.

Control – Quality Profiles

Define what “good” means for your business: company, content type, and language-specific checks in one profile, enforced automatically on every job. Quality reflects your real risk, not a generic benchmark.

Act – QPS thresholds

Set a QPS threshold and let it work. Segments above it are approved and locked; only segments below it go to human or AI post-editing. You raise straight-through processing without lowering the bar.

Illustration of Phrase Smart QA checks highlighting automatic quality assurance steps for localized content.

Catch – TMS or Strings rule-based QA

Automated QA checks in the TMS editor and in Phrase Strings catch the objective errors a score should never have to debate: terminology that breaks the termbase, missing or broken tags and placeholders, length overruns in UI.

Frequently asked questions about machine translation quality evaluation

What is machine translation quality evaluation?

It’s the process of measuring how accurate, fluent, and fit-for-purpose a machine translation is, so you can decide whether it can be published as-is or needs human or AI post-editing. Done well, it lets you automate safely instead of reviewing everything by hand.

What is the Quality Performance Score (QPS)?

QPS is an automated quality metric in the Phrase Platform that scores translations from 0 to 100 using the MQM 2.0 framework. It gives you an immediate, consistent quality number for every segment, which is what makes automated routing and reliable comparison possible.

What is MQM and why does Phrase use it?

MQM (Multidimensional Quality Metrics) is an open, industry-standard framework for categorising and weighting translation errors. Building QPS on MQM 2.0 means scores are transparent and defensible against a recognised standard, rather than a proprietary number you have to take on trust.

How is this different from just using a good AI model?

A model generates translation; it doesn’t tell you whether the output is right, and it doesn’t enforce your rules. Quality evaluation adds the control layer: measurement against a standard, your quality profile enforced automatically, routing, and an audit trail. That’s the difference between fluent output and content you can stand behind.

Can quality evaluation be customised to our business?

Yes. Quality Profiles let you define company, content, and language-specific checks and thresholds, so evaluation reflects your real risk and your brand and terminology requirements, not a generic benchmark.

How does this reduce post-editing cost?

You set a quality threshold. Segments that clear it are approved automatically; only segments below it are routed to human or AI post-editing. Review effort concentrates on higher-risk content, which lowers post-editing volume and turnaround time.

Does it work across our whole stack, not just one tool?

Yes. Evaluation and QA run where translation happens across the platform, including the TMS editor and Phrase Strings, so the same standard and controls apply to documents, software, and UI content rather than living in a separate tool.

Suggested Reading

Blog post

The $549 Billion Industry Where Content Is Clinical Infrastructure

A physician choosing a medical device is making a decision with consequences far beyond the brand. After more than 40 product launches across 135 countries, Cordis marketing leader Terrence Wiggins has seen how much trust depends on the information surrounding that device. Now AI is increasing content volumes just as regulation demands greater control.For MedTech companies, content now affects far more than communication. It can determine whether a product reaches a market and how safely it is used.

Webinar

The global content disconnect: Insights from 550 business leaders

In this session, we share findings from a survey of 550 senior business leaders and explore why scaling content is not scaling customer experience. We discuss what the most successful global enterprises are doing differently and why language intelligence is critical to winning customers in every market.

News

New Research Reveals Half of Enterprises Lost Revenue From Disconnected Customer Experiences

Most enterprises are scaling content faster than they can scale customer experience, according to an independent study of 550 senior leaders in nine countries.

The global content disconnect feature image

Resource

The global content disconnect report

Global research across nine countries reveals a growing disconnect between AI investment, content production, and customer experience. Enterprises are expanding faster, investing in AI, and producing more content in more languages for more markets than ever before. Yet the content customers receive often feels generic, inconsistent, and disconnected from the brand behind it.

Abstract graphic of connected icons representing automation, translation, and code syncing together.

Blog post

The localization work you never signed up for (and how continuous localization takes it back)

Long before I’d ever heard the term ‘continuous localization’ I was already very familiar with the problem it’s meant to solve. I spent years working as an iOS developer before I moved into product. At the time, localization wasn’t my main job but it became one of the many things competing for my attention. Some […]