Skip to main content

Multilingual AI data work

Multilingual AI evaluation

This is for AI teams and SaaS companies that need careful human judgment on model output in languages other than English. We rate and evaluate responses against your rubric, compare model outputs side by side, and red-team for the failures that only show up in a particular language or culture. Romanian and Russian are covered in-house at native level; other languages go to vetted partner linguists working under the same guidelines.

Price

$35/hour

Response

Reply within 1 business day

After delivery

30-day fix window

Who it's for#who-its-for

  • AI and ML teams that need evaluation data in languages their in-house raters do not speak, from someone who reads the guidelines closely and flags where they break down.
  • SaaS teams shipping an AI assistant in several markets, who need to know whether it answers correctly, and safely, in each language before customers find out the hard way.
  • Evaluation leads who want a reliable second opinion on a vendor's multilingual ratings, with disagreements explained rather than averaged away.
  • Data and annotation vendors that need a senior rater for Romanian or Russian, or a partner who can coordinate other languages under one set of rules.

The work suits teams that value rationale over raw throughput. If you need very large volumes of quick labels, a crowd platform is the better fit; if you need ratings you can defend, with reasons attached, this is. Throughput still matters, and we plan the queue with you so deadlines hold.

What you get#what-you-get

Rating work in the format your pipeline expects, plus the notes that make the data useful:

  • Response rating on your rubric, covering accuracy, helpfulness, fluency, tone, instruction following and safety, with written rationales wherever the rubric asks for them.
  • Side-by-side preference judgments for comparing models, prompts or system instructions.
  • Red-teaming: adversarial prompts written natively in the target language to surface harmful, biased or simply wrong output, including failures that never appear in English.
  • Checks for locale-specific problems: formality and register, names, dates and addresses, idioms, and cultural references that do not carry over.
  • Guideline feedback: where the rubric is ambiguous in a given language, with examples, so your guidelines improve along with the data.
  • Work done in your platform or tooling, or in spreadsheets if that is what your team uses.

Every batch also comes with a note on what we saw beyond the individual scores: recurring failure types, the prompts that exposed them, and cases where the model was right but the rubric would mark it wrong. That note is usually where guideline improvements come from.

How it runs#how-it-runs

  1. Guidelines. You share the rubric, examples and the tool. We read them, ask questions, and rate a calibration set.
  2. Calibration. We compare our ratings with your gold answers or expectations and settle disagreements before volume work starts.
  3. Rating. We work through the queue. Partner linguists handle other languages under the same guidelines, calibrated the same way. Items nobody can rate reliably are flagged rather than guessed.
  4. Reporting. You receive the ratings, the rationales and a short note on patterns and guideline issues worth fixing.
  5. 30-day fix window. For 30 days after delivery, if your QA flags ratings that do not follow the agreed guidelines, we redo them at no charge.

For red-teaming, we agree the risk categories and boundaries with you first, then log every successful attack with the exact prompt, the output and the language-specific reason it worked.

Confidentiality is standard for this work. We sign your NDA, work inside your tools where possible, and never reuse your prompts, outputs or guidelines elsewhere. Partner linguists sign a confidentiality agreement at least as strict before they see anything, as the confidentiality policy sets out.

Pricing#pricing

Rating, evaluation, red-teaming
$35 / hour
about €32.20

Rating, evaluation and red-teaming are billed by the hour at the rate above, whether the task is quick preference judgments or slow adversarial work. We track time per task, so you can see exactly where the hours go and compare them with your own estimates. The calibration set also gives us both a realistic time per item, so the estimate for the full queue rests on evidence rather than guesswork. Invoices break the hours down by language and task type.

What moves the total: the number of languages, the length and difficulty of each item, how much written rationale the rubric requires, and how much calibration a new guideline needs before ratings settle. A clear rubric with good examples is the single biggest saving, because calibration is where ambiguous guidelines cost time.

Prices are in US dollars; the euro figures are approximate conversions for reference. For translation and linguistic QA of your own product content, rather than model output, see localization and LQA.

Frequently asked questions#faq

Which languages can you evaluate?

Any language. Romanian and Russian we rate in-house at native level. Other languages are covered by vetted partner linguists who are native speakers, working under the same guidelines and calibration process.

Do you work in our annotation platform?

Yes, if you give us access. We can also work in spreadsheets or a simple export format and hand back data your pipeline can ingest.

How is red-teaming different from rating?

Rating judges outputs you already have. Red-teaming writes the inputs: prompts designed to make the model fail in ways specific to a language and culture, followed by a record of what went wrong and how to reproduce it.

Is this the same as your localization work?

They share linguists and attention to language, but the task is different. Localization makes content right; evaluation judges whether a model's output is right and explains why.

How do you keep ratings consistent?

Calibration before volume work, written rationales for borderline cases, and regular checks against gold items if you provide them. Where the guidelines are ambiguous, we raise it instead of guessing.

Send us what you have

Send a link to your docs, an export, or a few lines on what needs to change. We reply within 1 business day with questions or a quote.