Phrase quality evaluation
Put a number on translation quality, then govern it
Any model can generate translation. Phrase measures it, scoring every segment, applying your rules, and showing the gap between governed output and raw MT.
Fluent isn’t the same as correct
AI translation is fast, cheap, and convincing. That’s the problem. Raw machine translation reads well even when it’s wrong, and at enterprise volume no one can review every segment in every language to find out. Quality you can’t see is quality you can’t govern.
So teams do one of two things. They over-review, and lose the speed and cost savings that made automation worth it. Or they trust the output, and carry a risk they can’t measure. Neither scales. What scales is knowing, segment by segment, how good the translation actually is, and being able to prove it.

Quality evaluation is the control layer
Generating translation is no longer the hard part. Governing it is. Control and Governance is the part of the platform that turns AI translation from a gamble into a system: a standard to measure against, a score for every segment, your own rules enforced automatically, and a record of all of it.
That layer is what you can’t get from a model API alone, and it’s what makes automation safe to scale. Phrase measures each translation against the MQM 2.0 framework, holds it to the quality profile you define, and routes it accordingly, without a person reading every line. You automate on evidence rather than hope, and you keep the audit trail that regulated and brand-sensitive content depends on.
From score to action:
how Phrase manages translation quality

Measure – Quality Performance Score (QPS)
QPS rates every segment from 0 to 100 against the MQM 2.0 framework, so quality is a consistent number across projects, languages, and engines rather than a subjective call. Track it on dashboards in Phrase Analytics.

Control – Quality Profiles
Define what “good” means for your business: company, content type, and language-specific checks in one profile, enforced automatically on every job. Quality reflects your real risk, not a generic benchmark.

Act – QPS thresholds
Set a QPS threshold and let it work. Segments above it are approved and locked; only segments below it go to human or AI post-editing. You raise straight-through processing without lowering the bar.

Catch – TMS or Strings rule-based QA
Automated QA checks in the TMS editor and in Phrase Strings catch the objective errors a score should never have to debate: terminology that breaks the termbase, missing or broken tags and placeholders, length overruns in UI.
See your quality delta
Bring your own content and your own languages. We’ll score it, govern it, and show you the difference between what a model gives you and what you can publish.
Frequently asked questions about machine translation quality evaluation
What is machine translation quality evaluation?
It’s the process of measuring how accurate, fluent, and fit-for-purpose a machine translation is, so you can decide whether it can be published as-is or needs human or AI post-editing. Done well, it lets you automate safely instead of reviewing everything by hand.
What is the Quality Performance Score (QPS)?
QPS is an automated quality metric in the Phrase Platform that scores translations from 0 to 100 using the MQM 2.0 framework. It gives you an immediate, consistent quality number for every segment, which is what makes automated routing and reliable comparison possible.
What is MQM and why does Phrase use it?
MQM (Multidimensional Quality Metrics) is an open, industry-standard framework for categorising and weighting translation errors. Building QPS on MQM 2.0 means scores are transparent and defensible against a recognised standard, rather than a proprietary number you have to take on trust.
How is this different from just using a good AI model?
A model generates translation; it doesn’t tell you whether the output is right, and it doesn’t enforce your rules. Quality evaluation adds the control layer: measurement against a standard, your quality profile enforced automatically, routing, and an audit trail. That’s the difference between fluent output and content you can stand behind.
Can quality evaluation be customised to our business?
Yes. Quality Profiles let you define company, content, and language-specific checks and thresholds, so evaluation reflects your real risk and your brand and terminology requirements, not a generic benchmark.
How does this reduce post-editing cost?
You set a quality threshold. Segments that clear it are approved automatically; only segments below it are routed to human or AI post-editing. Review effort concentrates on higher-risk content, which lowers post-editing volume and turnaround time.
Does it work across our whole stack, not just one tool?
Yes. Evaluation and QA run where translation happens across the platform, including the TMS editor and Phrase Strings, so the same standard and controls apply to documents, software, and UI content rather than living in a separate tool.
























