What makes an AI translator good for documents
Author: Yuri Vodostoy — Technical translator and founder of Insight Translation · Updated 2026-08-11
Three questions decide whether a tool can translate a document rather than a sentence. Does it remember the whole job? Does it hold your terms across every file? Does the file come back in the format you sent, or only the text? Sentence quality is where most comparisons stop, and it's the part every modern engine has largely solved. On HardTrans the whole job goes in at once: one pass over all the files builds a project brief and a glossary, and both stay loaded while each file is translated. Pricing is $1.50 per 250-word page, and your first document up to 10 pages is free with no signup.
What “good for documents” actually means
Four things, and none of them is fluency.
- Whole-job memory. A document isn't a list of independent sentences. If the engine translates file 12 without knowing what file 3 called the same assembly, the set drifts, and a reader who meets three names for one part stops trusting the whole package.
- Your terminology, treated as fixed. Your glossary has to hold across the entire job, page 3 and page 300 alike. And you want to see the terms the tool settled on where your glossary was silent, so you can check them.
- The file back, not the text back. Text in a window isn't a deliverable. A DOCX with its tables, numbering, and section structure intact is.
- Throughput that matches the volume. A 300-page set translated at reading speed is no use on a deadline.
How the four options differ
Free web engines are strong per sentence and cost nothing. What they don't do is carry context between requests, so terminology gets settled one file at a time. Our dated side-by-side output across engineering, legal, medical, and scientific passages sits on the Google Translate vs HardTrans examples page.
ChatGPT translates well. The limitation is in the design: it's built around a chat window, not a job — no way to upload the whole set and get the files back, and the formatting is yours to rebuild. A pattern we keep seeing: in-house translation teams use chat AI every day for routine work, and still route multi-file document sets through a document pipeline. Daily users are exactly the people who know where the chat window stops.
Of the two general engines we have run, DeepL is the stronger on document prose: it won the August 5, 2026 passages, and on a long German text three weeks later it held six recurring terms to 8 English forms against Google's 13. Four easier passages had already come back clean from Google and were dropped before the August comparison. DeepL's own failures landed on terms whose only correct form is the one printed in the standard. The passages, the verbatim panels, and the clause references are on the DeepL alternative for documents page.
Specialized document pipelines, HardTrans included, trade the instant free translation for a first pass over the job. That pass costs a few minutes and is what makes page 480 agree with page 12.
| Free web engines | ChatGPT | DeepL | HardTrans | |
|---|---|---|---|---|
| What you hand over | one file, or pasted text | pasted text | one file, or pasted text | the whole job, every file at once |
| Terminology across files | each file on its own | each paste on its own | each file on its own | one brief and glossary, applied to every file |
| What comes back | a translated file or text | text in a chat window | a translated file or text | the same file type, formatting kept |
| Price | free | free tier, then subscription | free tier, then subscription | $1.50 per 250-word page |
DeepL plan structure checked August 5, 2026.
What an AI translator for large documents does first
Upload the whole set at once. Before any translation happens, we read every file, write a project brief describing what the document is and who it's for, and pull out a glossary of the terms that recur across it. Both stay loaded while each file is translated. That is the mechanism behind consistency on volume, and it is measurable: on one German regulation run through all three engines on August 25, 2026, six recurring terms came back in 13 English forms from Google Translate, 8 from DeepL and 6 from us. The panels and the count are on the DeepL alternative for documents page.
Your glossary is respected across the whole job; terms it does not cover go into the glossary we build, downloadable as CSV when the job finishes. Reading it takes a few minutes and tells you exactly which terms were locked.
The speed figures come from 15 multi-file jobs of 100 pages or more, March to July 2026: median 0.34 minutes per page. A 100-page set lands around 34 minutes, a 231-page set ran in 75, a 10-page document finishes in about 5.
Legal and official documents
The translation part works; the certification part does not exist.
On a contract set the brief and glossary from pass 1 stay loaded through every file, so a term defined in the definitions clause keeps the same English form in a schedule 200 pages later. The legal example on our comparison page shows a French contract clause where the term of art came out right with no glossary supplied.
There is no human editor in our loop, and we don't issue certified or sworn translations. If a court, a registry, or an immigration authority requires certification, our output doesn't meet that requirement no matter how well it reads. For anything that gets signed, filed, or submitted under liability, have a qualified specialist review the text and treat the AI output as the draft they start from.
Medical and pharmaceutical files raise the same certification question with higher terminology stakes; that case, SmPC headings included, is covered in the medical document translation guide.
PDF, and what the text layer decides
A PDF either has a text layer or it doesn't, and that single fact decides whether any AI translator can work on it. Export a PDF from Word, InDesign, or a CAD package and the characters are in the file. Photograph a page or scan it, and the file holds a picture of characters.
We translate text-layer PDFs in place: the source text is removed and the translation goes back into the same box. That has a price. The typeface changes wherever we can't resolve the original font, the type is squeezed where the translation runs longer than the source, and the repeat discount does not apply to PDFs at all. On a dense table or a drawing the page will not come back looking like the original, so if the layout is the point, send the DOCX or the XLIFF the PDF was made from. Scanned PDFs with no text layer are not supported, and nothing on our side changes that until we add OCR.
Translate a document and keep the formatting
You get back the file type you sent. DOCX, XLSX, and PPTX keep their tables, numbering, lists, and section structure; a text-layer PDF is translated in place, with the limits above; XLIFF comes back with inline tags intact, covering .xliff, .sdlxliff from Trados, .mqxliff and .mqxlz from memoQ, .mxliff from Phrase, and Smartcat exports. The round trip is in our guide to translating XLIFF online. How the same German source came back from DeepL, Google Translate and us, term by term, is on the DeepL alternative for documents page.
What we don't have: a translation memory, a CAT layer of our own, or any human in the loop. If your process runs on TM matches, HardTrans sits in front of your CAT tool as the pre-translation step. It doesn't replace it.
What it costs
$1.50 per page of 250 words. That covers the pass over the whole job, the glossary work, the translation itself, and the formatting on the way out. There is no subscription and no minimum: you pay for the pages in the job you actually ran.
Common questions
Is ChatGPT good enough for document translation?
For a paragraph you need to understand, yes. For a document you need to deliver, the problem is the shape of the tool, not the quality of its language: no file pipeline for a multi-file set, no memory carried across it, and the formatting left to you. We haven't run a measured quality comparison against it, and we're not going to guess at one.
What's the best free AI translator for documents?
For short documents where terminology carries no risk, the free tiers of the general engines are genuinely useful. They stop being the right tool at volume and under signature. Ours is one document up to 10 pages, no signup, which is enough to check the output on your own file before you decide anything.
Can it handle a 300-page set across many files?
Yes, and that's the case it was built for. Upload the files together so the first pass sees all of them. At our measured median, 300 pages runs under two hours; the largest set in our sample, 231 pages, took 75 minutes.
Which languages?
From any language into 12 targets, among them English, German, French, Spanish, Chinese, Japanese, and Arabic. The source language is detected automatically. A rare pair is worth a test file before you commit a job to it.
Do you keep my documents?
Files are deleted from our servers automatically after 7 days; until then you can download them again. Your text is never used to train models — neither by us nor by Anthropic, the model provider. Anthropic retains API data for a limited period under its policy and does not train on it.
Upload the document that gave you trouble last time. The first 10 pages cost nothing, and there's nothing to sign up for.
Related guides
First document free