Every localization decision is a trade-off in disguise. Fully human translation buys quality at the cost of speed and money. Raw machine output buys speed and savings at the cost of risk. Most teams sense this but resolve it badly, choosing one approach for all content instead of matching the approach to the stakes.
This article lays out the competing approaches honestly, names the axes that distinguish them, and ends with a decision rule you can actually apply. The goal is not to crown a winner, because there is none. The goal is to give you a way to decide, content type by content type, which trade-off is the right one to make.
The framing assumes you already accept that machine translation has a role. The open question is where its role ends and human judgment begins.
The Competing Approaches
Fully Human Translation
A professional translator produces and owns the text. This is the gold standard for quality, nuance, and cultural fit. It is also the slowest and most expensive option, and it does not scale to tens of thousands of strings or continuous release cycles. Its place is at the top of the risk tier, where being wrong is costly.
Raw Machine Translation
An engine produces text that ships without human review. It is nearly free and effectively instant, and modern engines make it surprisingly usable for low-stakes content. The risk is that errors are confident and fluent, so they hide. Raw machine output belongs only where a mistake costs little.
Machine Translation Plus Post-Editing
A human reviews and corrects machine drafts rather than translating from scratch. This is the workhorse middle ground: most of the speed of machine, most of the quality of human, at a fraction of the cost of full human translation. It is where the majority of serious content should live.
Transcreation
A fourth approach sits beyond translation entirely. Transcreation rewrites content for cultural impact, used for taglines, campaigns, and brand voice where a literal translation would fall flat. It is the most expensive and least scalable option, and machine translation's role here is limited to briefing the human creator on intent. Transcreation belongs only to the small slice of content where creative impact is the whole point.
The Axes That Matter
Cost of Error
This is the dominant axis. A mistranslated legal disclosure or dosage instruction can cause real harm; a slightly awkward product description cannot. The higher the cost of error, the more human judgment you should buy. This single axis explains most correct localization decisions, as the scenarios in Worked Scenarios Where Machine Translation Earned Its Keep illustrate.
Volume and Cadence
A one-time 40-page site and a continuously shipping product with thousands of strings demand different approaches. High volume and high cadence push you toward machine-heavy workflows simply because human-only cannot keep up. Low volume gives you room to invest human effort.
Visibility
A homepage headline read by every visitor warrants more care than a buried internal changelog. Visibility is a softer axis than cost of error, but it shapes where brand-quality review pays off.
Language Pair Maturity
A quieter axis is how well the engine handles the specific language pair. Machine quality varies widely: high-resource pairs like English-to-Spanish are strong, while lower-resource pairs can be noticeably weaker. The same content might be safe for machine-only treatment in one language and require review in another. This axis means tiering decisions are not purely about content; they interact with which languages you are targeting, and you should sanity-check engine quality per pair before trusting it.
The Decision Rule
Tier by Cost of Error, Then Adjust for Volume
The rule is straightforward. Sort content into three tiers by cost of error. Top tier gets machine drafts with full human review. Middle tier gets machine plus spot-checks. Bottom tier gets machine-only with automated checks. Then adjust: if a tier's volume is so high that human review cannot keep pace, push it down a tier and accept the residual risk consciously rather than by accident.
Make the Trade-Off Explicit
The failure mode is an implicit, uniform choice. The fix is to write down which tier each content type sits in and why. That record turns a vague anxiety into a defensible decision, and it connects directly to the structure in How the TIER Model Structures Localization Work.
Revisit the Tiers as Conditions Change
A tiering decision is not permanent. Engine quality improves, content gains or loses visibility, and a product that starts as an experiment can become brand-critical. Schedule a periodic review of your tier assignments rather than treating the first sort as final. A help center article that warranted spot-checks last year might warrant full review now that it ranks in search and drives signups. The decision rule stays the same; the inputs to it shift, and a tiering that is never revisited slowly drifts out of alignment with reality.
When the Rule Bends
Brand-Defining Content
Sometimes content is low-risk legally but high-stakes for brand. A tagline cannot be wrong even if a literal error costs nothing. Treat brand-defining content as top tier regardless of legal risk.
Regulated Domains
In medical, legal, and financial contexts, regulation may mandate certified human translation regardless of your internal tiering. The decision rule yields to the law, and you should confirm requirements before assuming any machine role. Measuring whether your tiering held is covered in Reading the Signal in Localization Quality Numbers.
A Worked Example of the Rule
Consider a SaaS product entering a new market. The terms of service are top tier because a mistranslated clause creates legal exposure, so they get machine drafts with full legal review. The marketing homepage is also top tier, but for brand reasons, so it gets transcreation. The help center sits in the middle, machine-translated with spot-checks, because errors are correctable and visible but not dangerous. Internal release notes go machine-only. One product, four different approaches, each chosen by running the same content type through the same axes. That is the rule working as intended: not a single verdict on the tools, but a per-content decision that the framework in How the TIER Model Structures Localization Work operationalizes.
The Cost of Choosing Badly
Over-Investing in Low-Stakes Content
When teams apply uniform high effort, the waste shows up as money and time spent reviewing content that nobody would have noticed errors in. Full human translation of internal changelogs is the classic example: real cost, negligible benefit. The opportunity cost is sharper than the direct cost, because every reviewer hour spent on a changelog is an hour not spent on the homepage that actually shapes the brand. Misallocated effort is invisible on the surface because the work looks diligent, but it quietly starves the content that matters.
Under-Investing in High-Stakes Content
The mirror failure is more dangerous. Running legal disclosures or medical instructions through raw machine output to save money can produce confident, fluent, and wrong text that creates real liability. The savings are small and immediate; the downside is large and delayed, which is exactly the asymmetry that makes this error so common. People discount a risk they cannot yet see. The decision rule exists precisely to force this trade-off into the open before the wrong content ships, so that under-investment becomes a visible choice rather than a silent default.
Frequently Asked Questions
Is there a single best approach?
No. The right approach depends on the content. A uniform choice across all content is the most common and most costly mistake. Match the approach to the stakes, volume, and visibility of each type.
What is the most important axis?
Cost of error. It dominates the decision: the more a mistranslation can harm users, the brand, or the business, the more human judgment you should buy for that content.
When is raw machine translation defensible?
When content is high-volume, low-visibility, and low-stakes, such as internal changelogs or generic specs. The test is whether a confident, fluent error would cause meaningful harm.
How does volume change the decision?
High volume and fast cadence can force content down a tier because human review simply cannot keep pace. The key is to make that downgrade a conscious choice with eyes open, not an accident.
What if legal and brand risk disagree?
Take the higher of the two. Brand-defining taglines warrant top-tier care even at zero legal risk, and regulated content warrants certified human translation even when brand stakes are low.
Key Takeaways
- There is no single best approach; the right one depends on the content.
- Cost of error is the dominant axis in nearly every correct decision.
- Tier content into machine-only, machine-plus-spot-check, and machine-draft-plus-review.
- Adjust tiers for volume and cadence, downgrading consciously rather than by accident.
- Brand-defining and regulated content override the default tiering toward more human judgment.