Who doesmultilingualdata work.
The companies doing multilingual human data work for AI, from the AI Circle × Data Gradient Multilinguality Map, grouped by the kind of work. If you train on Data Gradient for translation, evaluation, speech or preference work, these are the companies that do it.
Localization & Translation
Acolad
Paris, France · GrowthEuropean language services group covering translation, localization, and content services, with AI data services offered alongside.
DataForce / TransPerfect
New York, NY · SubsidiaryTransPerfect's AI training data business: multilingual data collection, annotation, and linguistic validation built on TransPerfect's global translation workforce.
e2f
San Jose, CA · GrowthLanguage data and localization company supplying multilingual training data, linguistic annotation, and evaluation to AI teams.
Keywords Studios
Dublin, Ireland · GrowthGlobal services provider to the games and entertainment industry, including localization, audio, and player support, with an AI data and testing line.
LILT
San Francisco, CA · Series CAI-first translation platform combining adaptive machine translation with human linguists in the loop, and a growing line of multilingual data and evaluation services for model builders.
Lingo.dev
Developer-focused localization tooling that uses LLMs to translate apps and content from the codebase.
Lionbridge
Waltham, MA · GrowthLarge translation and localization provider that also offers AI training data, evaluation, and model testing services across hundreds of languages.

Pangeanic
Valencia, Spain · GrowthTranslation and language technology company offering machine translation, anonymization, and multilingual datasets for AI training.
Phrase
Prague, Czechia · GrowthLocalization platform (formerly Memsource) covering translation management, machine translation orchestration, and quality scoring.

Powerling
FranceLocalization and global content services company (25+ years, 150+ experts, 6 offices, B Corp) that also builds custom speech and audio datasets for AI, including dual-channel conversational recordings and low-resource languages.
RWS TrainAI
Chalfont St Peter, UK · PublicThe AI data services arm of RWS, one of the largest language services companies. Offers multilingual data collection, annotation, and linguistic evaluation drawn from its translator and linguist network.
Smartling
New York, NY · GrowthCloud translation management platform with integrated machine and human translation services.
Translated
Rome, Italy · GrowthLanguage services company behind ModernMT and a large professional translator network, combining adaptive machine translation with human post-editing.

Unbabel
Lisbon, Portugal / San Francisco, CA · Series CLanguage operations platform that pairs machine translation with a community of human editors, plus translation quality estimation used to evaluate MT output.

wxrks
Translation management system with native AI, built to unify cost, quality, workflow, and automation across localization programs.
Benchmarks & Evals
Appen
Sydney, Australia · PublicOne of the longest-running crowd data companies, with roots in linguistic data: collection, annotation, and evaluation across audio, text, image, and video in many languages.
Argos Data
Kraków, Poland · SubsidiaryAI data arm of language services provider Argos Multilingual: multilingual evaluation, annotation, and data collection including lower-resource languages.

Centific
Bellevue, WA · GrowthLarge-scale data services with a multilingual workforce; OneForma is its crowd platform for collection, annotation, and evaluation work.
DATAmundi
Multilingual AI data services covering evaluation, annotation, and language data. Limited public information.
Hume AI
New York, NY · Series BEmotionally intelligent voice AI research lab, known for expressive speech models and large-scale human ratings of vocal and facial expression across cultures.

Innodata
Ridgefield Park, NJ · PublicData engineering, annotation, and model safety evaluation for large AI labs, with a large multilingual delivery workforce.
Invisible Technologies
New York, NY · GrowthManaged AI training data operations with a large trained workforce; a major post-training (RLHF and SFT) vendor for frontier labs, including multilingual programs.

OpenTrain
Talent network of 476,000+ pre-vetted AI trainers and data labelers in 230+ countries for RLHF, LLM evaluation, red-teaming, and annotation, including multilingual evaluation across 118+ languages.
Perle AI
San Francisco, CA · SeedExpert-driven data labeling and evaluation with deep expertise in linguistics, medical, and robotics data, built on verified human contributors with provenance tracking.
Prolific
London, UK · Series AA vetted, demographically balanced participant pool originally built for academic research, now used for human evaluation, preference data, and model testing.
PublicAI
Web3-based train-to-earn network where contributors complete data tasks to improve AI models and earn rewards.

Respondent
Verified research-participant panel (4.3M+ people across 150+ countries) with an AI training line: model evaluation and safety, red-teaming, alignment and preference data, and multilingual, cross-cultural work in 100+ languages.

Scale AI
San Francisco, CA · GrowthLarge generalist data engine covering RLHF, SFT, evaluation, and public leaderboards, with multilingual contributor programs through Outlier and Remotasks.
Surge AI
San Francisco, CAHigh-skill human data for frontier labs: RLHF, SFT, evaluation, and red-teaming with vetted expert annotators, including multilingual programs.
TaskUs
New Braunfels, TX · PublicOutsourced digital services company with an AI services line covering data labeling, RLHF, safety, and red-teaming across many languages.

Toloka
Amsterdam, Netherlands · GrowthGlobal crowd and expert network for data collection, annotation, and evaluation, with multilingual benchmark and red-teaming programs for frontier labs.

Welo Data
New York, NY · SubsidiaryData services arm of localization firm Welocalize, offering annotation, evaluation, and expert networks across 250+ languages.
Speech, Voice & Audio
Besimple AI
Human data and annotation for AI, including speech and audio. Limited public information.
David AI
San Francisco, CAAudio data company building large conversational speech datasets for voice AI and speech model training.
Defined.ai
Seattle, WA · Series BMarketplace and services for ethically sourced AI training data (formerly DefinedCrowd), strongest in speech and multilingual datasets.

Fluffle
Consumer app that pays people to talk with their friends, generating natural conversational speech data.

Liva AI
San Francisco, CA · SeedBuilding socially intelligent AI, with a focus on human voice and expression data. Founded by Ashley Mo and Aoi Otani.
LXT
Toronto, Canada · GrowthAI data company specializing in speech and multilingual data collection, transcription, and evaluation across 1,000+ language locales.
oto
Full-duplex speech data for voice AI: datasets, models, evals, and tools built from real two-channel conversation.

Panels
Speech and voice data collection for AI. Limited public information.
Sigma AI
Madrid, Spain · GrowthHuman data company with long experience in speech and linguistic annotation, transcription, and multilingual data collection.
Silencio AI
Voice and audio data network for AI, sourcing real voices from people in 180+ countries, including languages current voice AI handles poorly.
Sonexis
IndiaSpeech and audio data services. Limited public information.
SpeechData.ai
Speech data collection and transcription for ASR and TTS. Limited public information.

TELUS Digital
Vancouver, Canada · PublicGlobal data collection and annotation at BPO scale, with a large multilingual crowd for speech, text, and cultural data.

The Agentic Data Co.
San Francisco, CA · SeedTraining data for speech models. Founded by Christian Vestergaard and Matias Voldby Drejer.
Uplift AI
Voice AI for underserved languages, building speech models and data for languages such as Urdu.
RLHF, SFT & Annotation
Argilla
Madrid, Spain · SubsidiaryOpen-source collaboration tool where AI engineers and domain experts build and curate datasets for language-model fine-tuning, RLHF, and evaluation, now part of Hugging Face.
Chipo
Human data for AI training. Limited public information.

Concentrix
Newark, CA · PublicGlobal customer-experience BPO with AI data services, including annotation and RLHF delivered in many languages.
Lifewood
Hong KongAI data services company with delivery centers across Asia and Africa, covering annotation, collection, and multilingual data.

Luel
San Francisco, CA · SeedTurns everyday words and actions into usable training data for AI (per its YC profile).

Mundo AI
Vancouver, Canada · SeedHigh-quality multilingual training data for AI models (per its YC profile).
Turing
Palo Alto, CA · Series EPost-training data and services for frontier labs, specializing in coding, reasoning, and SFT/RLHF data from a global network of vetted experts.

Uber AI Solutions
San Francisco, CA · PublicUber's data business: annotation, collection, and human-in-the-loop work drawn from its global driver and courier network plus a dedicated expert workforce across many languages.
UsergyAI
Human feedback and annotation for AI. Limited public information.
Low-Resource & Cultural Data
Cognegica Networks
Language and cultural data services. Limited public information.
Kalasko
Language and cultural data services. Limited public information.
Karya
Bengaluru, India · NonprofitNonprofit data company that pays rural and low-income workers in India to create speech and text data in Indian languages, with workers sharing in data resale.
Nerval AI
Data for underserved languages and regions. Limited public information.
XRI Global
Language technology and data for low-resource languages. Limited public information.
Profiles are short summaries written from public information and may be out of date; check with each company before you apply. Companies are listed under their main category, with every category they serve shown on the card. Missing a company? Suggest it on the AI Circle map.