{"id":3178,"date":"2026-09-10T11:32:19","date_gmt":"2026-09-10T06:02:19","guid":{"rendered":"https:\/\/www.ulatus.com\/translation-blog\/?p=3178"},"modified":"2026-10-07T11:36:53","modified_gmt":"2026-10-07T06:06:53","slug":"what-ai-needs-from-human-now-to-build-a-completed-automated-world","status":"publish","type":"post","link":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/","title":{"rendered":"What AI needs from human now to build a completed automated world?"},"content":{"rendered":"<p>Roughly half of all websites are written in English, a language spoken by under a fifth of the people on Earth. That mismatch hints at a bigger problem. In 2025, Gartner reported that 63% of organizations lack the right data management practices for AI, or aren&#8217;t sure they have them. Models get better every quarter, but the data for AI that only people can supply is still the bottleneck.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_84 counter-flat ez-toc-counter ez-toc-custom ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<span class=\"ez-toc-title-toggle\"><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#What_Does_AI_Need_From_Humans_to_Build_a_Fully_Automated_World\" >What Does AI Need From Humans to Build a Fully Automated World?<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Why_each_of_the_five_still_needs_people\" >Why each of the five still needs people<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Why_%E2%80%9Cdata_for_AI%E2%80%9D_is_not_the_same_as_%E2%80%9Cmore_data%E2%80%9D\" >Why &#8220;data for AI&#8221; is not the same as &#8220;more data&#8221;<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#What_Is_Data_for_AI_and_Which_Parts_Can_Machines_Make_Themselves\" >What Is Data for AI, and Which Parts Can Machines Make Themselves?<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Raw_data_vs_training_data_vs_evaluation_data\" >Raw data vs training data vs evaluation data<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#What_synthetic_data_can_and_cannot_replace\" >What synthetic data can and cannot replace<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Why_Is_Human-Generated_Data_Running_Short\" >Why Is Human-Generated Data Running Short?<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#The_public-text_ceiling\" >The public-text ceiling<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Why_scraping_is_giving_way_to_consented_licensed_data\" >Why scraping is giving way to consented, licensed data<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#How_Do_Humans_Label_Rank_and_Test_What_AI_Learns\" >How Do Humans Label, Rank and Test What AI Learns?<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Annotation_turning_raw_data_into_ground_truth\" >Annotation: turning raw data into ground truth<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#RLHF_teaching_models_what_people_prefer\" >RLHF: teaching models what people prefer<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Evaluation_and_red-teaming_checking_before_automating\" >Evaluation and red-teaming: checking before automating<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Where_Is_the_Biggest_Gap_in_Data_for_AI\" >Where Is the Biggest Gap in Data for AI?<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-15\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#English_dominates_the_data_not_the_planet\" >English dominates the data, not the planet<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-16\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#What_breaks_when_AI_meets_dialects_and_low-resource_languages\" >What breaks when AI meets dialects and low-resource languages<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-17\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Code-switching_culture_and_context_what_native_speakers_add\" >Code-switching, culture and context: what native speakers add<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-18\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#What_Does_Good_Human_Data_for_AI_Look_Like\" >What Does Good Human Data for AI Look Like?<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-19\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#A_five-point_quality_test_consent_coverage_consistency_context_compliance\" >A five-point quality test: consent, coverage, consistency, context, compliance<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-20\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Questions_to_ask_any_data_partner_before_you_sign\" >Questions to ask any data partner before you sign<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-21\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#What_Should_Teams_Do_Now_to_Prepare_Their_Data_for_AI\" >What Should Teams Do Now to Prepare Their Data for AI?<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-22\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#A_four-step_plan_audit_source_label_evaluate\" >A four-step plan: audit, source, label, evaluate<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-23\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Where_to_start_if_you_operate_in_more_than_one_language\" >Where to start if you operate in more than one language<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-24\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#Conclusion_Automation_Runs_on_Human-Supplied_Data\" >Conclusion: Automation Runs on Human-Supplied Data<\/a><\/li><\/ul><\/nav><\/div>\n<h2><span class=\"ez-toc-section\" id=\"What_Does_AI_Need_From_Humans_to_Build_a_Fully_Automated_World\"><\/span>What Does AI Need From Humans to Build a Fully Automated World?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>AI needs five kinds of human input: original human-created data, human labels and judgment, human preference feedback, human evaluation and red-teaming, and multilingual, culturally grounded data. Automation scales up whatever people teach it. In our view, a model cut off from those five inputs tends to stall, drift, or fail the moment it leaves the languages and situations it has already seen.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Why_each_of_the_five_still_needs_people\"><\/span>Why each of the five still needs people<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Original behavior, writing and speech are the raw material, and no model has produced them from scratch. Labels make that material learnable, preference rankings show which answer people actually favor, and evaluation catches failures before real work gets handed over. Multilingual and cultural data is the most under-served of the five, and it&#8217;s also the one that decides whether a system works for most of the world.<\/p>\n<p>Edwin Chen, founder of Surge AI, told Forbes in 2025: &#8220;without us, AGI just won&#8217;t happen.&#8221; He sells human data, so he has a commercial interest. Still, the research below backs up his basic point.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Why_%E2%80%9Cdata_for_AI%E2%80%9D_is_not_the_same_as_%E2%80%9Cmore_data%E2%80%9D\"><\/span>Why &#8220;data for AI&#8221; is not the same as &#8220;more data&#8221;<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Volume drove the last decade. Quality, provenance and coverage will drive this one. Gartner predicted in February 2025 that through 2026, organizations will abandon 60% of AI projects that lack AI-ready data, which tells you where executives expect trouble. If you need help sourcing this kind of material, Ulatus offers <a href=\"https:\/\/www.ulatus.com\/ai-training-data\">AI training data services<\/a> built around human contributors.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"What_Is_Data_for_AI_and_Which_Parts_Can_Machines_Make_Themselves\"><\/span>What Is Data for AI, and Which Parts Can Machines Make Themselves?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>In one sentence, data for AI is any information a model learns from or is measured against, whether it&#8217;s raw, labeled, ranked or tested, and its most valuable parts still start with humans. Machines can generate synthetic data to add to what already exists, like paraphrases, simulated scenarios or extra examples. What they can&#8217;t yet do reliably is come up with new real-world behavior, supply ground-truth judgment, or confirm what&#8217;s culturally appropriate. People remain the source of truth, and machines multiply it.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Raw_data_vs_training_data_vs_evaluation_data\"><\/span>Raw data vs training data vs evaluation data<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Raw data is whatever you collect before anyone shapes it: logs, documents, recordings, images. Training data is the portion a model learns from. Validation data gets held back to tune settings, and test data stays locked away until the end so you get an honest score. Wikipedia&#8217;s page on <a href=\"https:\/\/en.wikipedia.org\/wiki\/Training,_validation,_and_test_data_sets\">training, validation and test data sets<\/a> covers the distinctions well.<\/p>\n<p>Why does the split matter? Because a model graded on data it has already seen looks brilliant, then performs badly in the real world.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"What_synthetic_data_can_and_cannot_replace\"><\/span>What synthetic data can and cannot replace<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Synthetic data is genuinely useful. It fills gaps, balances out rare cases, and can be cheaper than collecting everything by hand. The trouble starts when it replaces real data instead of supplementing it.<\/p>\n<p>Model collapse happens when models train again and again on other models&#8217; output until the rare parts of the original data disappear. A July 2024 Nature paper by Shumailov and colleagues demonstrated this across several kinds of models, including large language models, and described the defects as irreversible.<\/p>\n<p>Follow-up work adds some nuance. Gerstgrasser and colleagues showed in 2024 that collapse depends on whether synthetic data replaces real data or piles up alongside it: when real data stays in the mix, error stays bounded. And a 2025 arXiv preprint, &#8220;A Probabilistic Perspective on Model Collapse,&#8221; finds that collapse is avoided only if the training sample keeps growing at each step. Both results argue for keeping humans in the loop, since people are the main source of fresh, real data.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Why_Is_Human-Generated_Data_Running_Short\"><\/span>Why Is Human-Generated Data Running Short?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Public human-written text is finite. In a 2024 paper, Epoch AI, a research institute, projected that models will train on datasets roughly equal to the entire stock of public human text sometime between 2026 and 2032, if scaling trends continue.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_public-text_ceiling\"><\/span>The public-text ceiling<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Epoch&#8217;s 2025 report, &#8220;Can AI Scaling Continue Through 2030?&#8221;, estimates that the indexed web holds about 500 trillion words of unique text, and that it will grow by about 50% by 2030.<\/p>\n<p>After adjusting for quality, repeated passes and tokenizer efficiency, Epoch estimates that 400 trillion to 20 quadrillion token-equivalents will be available for training by 2030 (tokens are, roughly, chunks of words). That range is enormous. Nobody knows exactly when the well runs dry, but it&#8217;s not bottomless, and Epoch also notes that synthetic data may be needed to ease bottlenecks, which brings the collapse risk right back.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Why_scraping_is_giving_way_to_consented_licensed_data\"><\/span>Why scraping is giving way to consented, licensed data<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Scarcity isn&#8217;t the only pressure on scraped data, because the law is moving too. Under Article 53 of the EU AI Act, providers of general-purpose AI models must publish a sufficiently detailed public summary of their training content and keep a copyright-compliance policy. According to Bird &amp; Bird, the Commission released the mandatory template on 24 July 2025, and the obligations applied from 2 August 2025.<\/p>\n<p>Some high-risk deadlines were later pushed back by the AI Omnibus, formally adopted in mid-2026, but the general-purpose model rules stayed as they were. Data with clear origins is easier to document, which is why <a href=\"https:\/\/www.ulatus.com\/ai-training-data\/data-collection-services\">consent-backed AI data collection<\/a> is increasingly expected rather than optional.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"How_Do_Humans_Label_Rank_and_Test_What_AI_Learns\"><\/span>How Do Humans Label, Rank and Test What AI Learns?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Humans do three jobs that models still can&#8217;t fully do for themselves, even with model-assisted labeling. They annotate data so it becomes ground truth, they rank outputs so models learn what people prefer, and they test systems for errors and risks before deployment. All three are human-in-the-loop work, meaning a person reviews, corrects or approves what a machine produces.<\/p>\n<p>The market is growing, too. The Business Research Company estimates the AI annotation market at USD 1.91 billion in 2025, rising to USD 2.51 billion in 2026.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Annotation_turning_raw_data_into_ground_truth\"><\/span>Annotation: turning raw data into ground truth<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Annotation is where a person marks what a recording, image or sentence contains. Picture a call-center transcript. Someone tags the customer&#8217;s intent, the sentiment, and any account numbers that must be masked. Train a model on thousands of examples like that, and it learns to do much of the tagging with limited supervision. That&#8217;s why <a href=\"https:\/\/www.ulatus.com\/ai-training-data\/data-annotation-services\">human data annotation<\/a> remains the foundation of supervised learning.<\/p>\n<p>So how do you know the labels are any good? Ask for an agreement score. Cohen&#8217;s kappa measures how often two annotators agree beyond chance. A long-standing rule of thumb (Landis and Koch, 1977) treats 0.81 and above as almost perfect agreement, while McHugh (2012) treats anything below 0.60 as inadequate. Those bands give buyers an actual number to ask for.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"RLHF_teaching_models_what_people_prefer\"><\/span>RLHF: teaching models what people prefer<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p><a href=\"https:\/\/en.wikipedia.org\/wiki\/Reinforcement_learning_from_human_feedback\">Reinforcement learning from human feedback (RLHF)<\/a> trains a model using people&#8217;s rankings of its answers. Those rankings train a reward model, which then steers the main model. Imagine a chatbot gives three replies to a billing complaint. A reviewer picks the one that&#8217;s accurate, polite and right for the customer&#8217;s culture, and the model learns to favor that pattern.<\/p>\n<p>Expertise counts here. A tax answer is best judged by a tax professional, and a question about Japanese honorifics by a fluent speaker. Generic crowd work can miss exactly those distinctions.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Evaluation_and_red-teaming_checking_before_automating\"><\/span>Evaluation and red-teaming: checking before automating<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Before a system takes over a workflow, someone has to check it. Evaluators score outputs for accuracy and tone, while red-teamers try to break the model with tricky or harmful prompts. Doing this in more than one language matters, because <a href=\"https:\/\/www.ulatus.com\/ai-training-data\/data-evaluation-services\">multilingual model evaluation by native speakers<\/a> catches problems that English-only testing often misses. And the more work you hand to automation, the more that judgment is worth.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Where_Is_the_Biggest_Gap_in_Data_for_AI\"><\/span>Where Is the Biggest Gap in Data for AI?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Multilingual data is a strong candidate for the biggest gap, because the web that models learn from doesn&#8217;t look like the people who use them. English dominates websites far beyond its share of speakers, and models perform noticeably worse in lower-resource languages.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"English_dominates_the_data_not_the_planet\"><\/span>English dominates the data, not the planet<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>W3Techs reports that English is the content language of 49.5% of websites with a known content language, as of October 2026, followed by Spanish at 6.0% and German at 5.9%. That&#8217;s a share of websites, not of training corpora. Most frontier labs don&#8217;t disclose their language mix, but where it was disclosed for older models, English made up about 93% of GPT-3&#8217;s training data (OpenAI, 2020) and about 90% of Llama 2&#8217;s (Meta, 2023).<\/p>\n<p>Ethnologue&#8217;s 2025 figures put English at about 1.5 billion total speakers, against a world population of about 8.2 billion. So you get half the websites for under a fifth of the people.<\/p>\n<p>Ethnologue also counts about 7,170 living languages in 2025, and many of them have little presence online.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"What_breaks_when_AI_meets_dialects_and_low-resource_languages\"><\/span>What breaks when AI meets dialects and low-resource languages<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Stanford HAI warned in May 2025 that approximately 5 billion people who don&#8217;t speak English are poorly served by today&#8217;s AI tools. Large language models work well for English&#8217;s roughly 1.5 billion speakers, it noted, but they underperform for Vietnamese (97 million speakers) and do far worse for Nahuatl (1.5 million). Researcher Sanmi Koyejo defines low-resource languages as those &#8220;with limited amounts of computer-readable data about them.&#8221;<\/p>\n<p>Benchmarks show what that means in practice. IrokoBench, published at NAACL 2025, tested models on 17 African languages using human-translated tasks. At the time of testing, the best open model, Gemma 2 27B, reached only 63% of the score of the best proprietary model, GPT-4o. So for a team supporting customers in, say, Yoruba or Swahili, that means starting below two-thirds of the top system&#8217;s performance. Translating the test sets into English first helped close the gap for larger English-centric models, which suggests the weakness lies more in native-language data than in the task itself.<\/p>\n<p>The safety picture is messier. A May 2026 arXiv preprint from Stellenbosch University tested multi-turn jailbreak conversations in Afrikaans, Kiswahili, isiXhosa and isiZulu. Harmful response rates reached 60% to 78% in Afrikaans and up to 71% in Kiswahili, against an English baseline of 53% to 84%. The authors found that translation quality, not just how much data a language has, drove jailbreak success. That&#8217;s another argument for native speakers who can judge translation quality.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Code-switching_culture_and_context_what_native_speakers_add\"><\/span>Code-switching, culture and context: what native speakers add<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Real speakers mix languages mid-sentence, use regional slang, and shift register depending on who&#8217;s listening. A model trained mostly on English web text sees comparatively little of that, and Stanford&#8217;s team cautions that models risk flattening cultural diversity into a largely US-centric view.<\/p>\n<p>Native speakers fix this because they know what&#8217;s polite, what&#8217;s risky and what simply sounds wrong. In our view, translators and linguists are an underused supply of AI data, since they work across languages every day. For speech systems, <a href=\"https:\/\/www.ulatus.com\/ai-training-data\/data-transcription-services\">native-speaker transcription for speech models<\/a> is a practical first step: accurate transcripts of real accents and dialects.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"What_Does_Good_Human_Data_for_AI_Look_Like\"><\/span>What Does Good Human Data for AI Look Like?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Good human data passes a five-point test, the 5 Cs: consent, coverage, consistency, context and compliance. It&#8217;s our own checklist, not a regulatory standard, but each C maps to a check buyers can actually verify.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"A_five-point_quality_test_consent_coverage_consistency_context_compliance\"><\/span>A five-point quality test: consent, coverage, consistency, context, compliance<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<ul>\n<li>Contributors agreed to how their data is used, and you can show the record.<\/li>\n<li>The data spans the languages, dialects and situations your users represent.<\/li>\n<li>Labels follow written guidelines, and an agreed kappa target shows annotators apply them the same way.<\/li>\n<li>Annotators understand the domain and culture, not just the words.<\/li>\n<li>Provenance, licensing and copyright policy are documented, which is what the EU training-data summary template pushes providers toward.<\/li>\n<\/ul>\n<h3><span class=\"ez-toc-section\" id=\"Questions_to_ask_any_data_partner_before_you_sign\"><\/span>Questions to ask any data partner before you sign<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Before you sign, ask four things:<\/p>\n<ul>\n<li>How are contributors recruited and vetted, and is consent documented?<\/li>\n<li>Which agreement metric do you report, at what threshold and on what sample size?<\/li>\n<li>Can we see sample QA reports and error logs, not a slide deck?<\/li>\n<li>How do you handle dialects and match reviewers to domains?<\/li>\n<\/ul>\n<p>Vendor numbers deserve some caution. Ulatus says its network includes more than 100,000 vetted contributors across 200-plus languages and 125-plus countries, and it lists inter-annotator agreement reporting as a service. Those figures are self-reported, so ask any vendor, including us, to back them up with audit artifacts.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"What_Should_Teams_Do_Now_to_Prepare_Their_Data_for_AI\"><\/span>What Should Teams Do Now to Prepare Their Data for AI?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Start small and measurable: audit one workflow, source the missing data, label it with agreement checks, and evaluate before scaling. A narrow pilot exposes data gaps cheaply, and it gives you real error rates instead of vendor promises.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"A_four-step_plan_audit_source_label_evaluate\"><\/span>A four-step plan: audit, source, label, evaluate<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<ol>\n<li>List the data your chosen workflow depends on. Mark what&#8217;s original, what&#8217;s synthetic, what has clear consent and which languages it covers.<\/li>\n<li>Fill gaps with consented, documented data from real contributors, prioritizing your customers&#8217; languages.<\/li>\n<li>Annotate with written guidelines, trained reviewers and a kappa target agreed in advance.<\/li>\n<li>Test with held-out data and native-speaker reviewers, and track error rates before widening the rollout.<\/li>\n<\/ol>\n<h3><span class=\"ez-toc-section\" id=\"Where_to_start_if_you_operate_in_more_than_one_language\"><\/span>Where to start if you operate in more than one language<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Pick a single language pair or locale. Then measure model performance directly in that language instead of assuming English results carry over, a caution IrokoBench&#8217;s findings support. Stanford HAI also recommends community-driven data collection with fair ownership for low-resource languages. Once agreement and error rates look healthy, add the next locale.<\/p>\n<p>For content-heavy workflows, human-corrected translations can double as training data, which makes <a href=\"https:\/\/www.ulatus.com\/machine-translation\">machine translation with human post-editing<\/a> a sensible bridge. The machine handles the volume, and linguists catch the nuance.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Conclusion_Automation_Runs_on_Human-Supplied_Data\"><\/span>Conclusion: Automation Runs on Human-Supplied Data<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Automation needs original human data, careful labels, ranked preferences, honest evaluation and, above all, data in the languages people actually speak. Skip any of them and you end up with a model that looks brilliant on its own test data, then fails in the language nobody tested.<\/p>\n<p>So this quarter, audit one workflow and write down its data gaps: what&#8217;s missing, what&#8217;s synthetic, and which languages go untested. If you want a partner for consent-backed multilingual training data, annotation and evaluation, bring <a href=\"https:\/\/www.ulatus.com\/contact-us\">Ulatus<\/a> one workflow and one language, and ask for a sample batch with an agreement report. It&#8217;s the quickest way to find out whether your data for AI is ready.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Roughly half of all websites are written in English, a language spoken by under a fifth of the people on Earth. That mismatch hints at a bigger problem. In 2025, Gartner reported that 63% of organizations lack the right data management practices for AI, or aren&#8217;t sure they have them. Models get better every quarter, [&hellip;]<\/p>\n","protected":false},"author":25,"featured_media":3179,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[464],"tags":[],"class_list":["post-3178","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-for-ai"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>What AI needs from human now to build a completed automated world?<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"What AI needs from human now to build a completed automated world?\" \/>\n<meta property=\"og:description\" content=\"Roughly half of all websites are written in English, a language spoken by under a fifth of the people on Earth. That mismatch hints at a bigger problem. In 2025, Gartner reported that 63% of organizations lack the right data management practices for AI, or aren&#8217;t sure they have them. Models get better every quarter, [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-10T06:02:19+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-10-07T06:06:53+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.ulatus.com\/translation-blog\/wp-content\/uploads\/2026\/10\/banner-what-ai-needs-from-humans-automated-world.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1280\" \/>\n\t<meta property=\"og:image:height\" content=\"720\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Elizabeth Rodrigues\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@ulatus\" \/>\n<meta name=\"twitter:site\" content=\"@ulatus\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Elizabeth Rodrigues\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/\"},\"author\":{\"name\":\"Elizabeth Rodrigues\",\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/#\\\/schema\\\/person\\\/2c88b235c1a21525603db0395518c787\"},\"headline\":\"What AI needs from human now to build a completed automated world?\",\"datePublished\":\"2026-09-10T06:02:19+00:00\",\"dateModified\":\"2026-10-07T06:06:53+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/\"},\"wordCount\":2423,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/banner-what-ai-needs-from-humans-automated-world.png\",\"articleSection\":[\"Data for AI\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/\",\"url\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/\",\"name\":\"What AI needs from human now to build a completed automated world?\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/banner-what-ai-needs-from-humans-automated-world.png\",\"datePublished\":\"2026-09-10T06:02:19+00:00\",\"dateModified\":\"2026-10-07T06:06:53+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/#\\\/schema\\\/person\\\/2c88b235c1a21525603db0395518c787\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#primaryimage\",\"url\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/banner-what-ai-needs-from-humans-automated-world.png\",\"contentUrl\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/banner-what-ai-needs-from-humans-automated-world.png\",\"width\":1280,\"height\":720,\"caption\":\"AI needs from human now to build a completed automated world\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"What AI needs from human now to build a completed automated world?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/#website\",\"url\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/\",\"name\":\"\",\"description\":\"Translation Trends &amp; Insights\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/#\\\/schema\\\/person\\\/2c88b235c1a21525603db0395518c787\",\"name\":\"Elizabeth Rodrigues\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/8cbc606363a4bbd8eda6ccc8c6a750eae05a2de996ee4dc49e0c776598bb1d45?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/8cbc606363a4bbd8eda6ccc8c6a750eae05a2de996ee4dc49e0c776598bb1d45?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/8cbc606363a4bbd8eda6ccc8c6a750eae05a2de996ee4dc49e0c776598bb1d45?s=96&d=mm&r=g\",\"caption\":\"Elizabeth Rodrigues\"},\"url\":\"https:\\\/\\\/www.ulatus.com\\\/translation-blog\\\/author\\\/elizabeth-rodrigues\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"What AI needs from human now to build a completed automated world?","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/","og_locale":"en_US","og_type":"article","og_title":"What AI needs from human now to build a completed automated world?","og_description":"Roughly half of all websites are written in English, a language spoken by under a fifth of the people on Earth. That mismatch hints at a bigger problem. In 2025, Gartner reported that 63% of organizations lack the right data management practices for AI, or aren&#8217;t sure they have them. Models get better every quarter, [&hellip;]","og_url":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/","article_published_time":"2026-09-10T06:02:19+00:00","article_modified_time":"2026-10-07T06:06:53+00:00","og_image":[{"width":1280,"height":720,"url":"https:\/\/www.ulatus.com\/translation-blog\/wp-content\/uploads\/2026\/10\/banner-what-ai-needs-from-humans-automated-world.png","type":"image\/png"}],"author":"Elizabeth Rodrigues","twitter_card":"summary_large_image","twitter_creator":"@ulatus","twitter_site":"@ulatus","twitter_misc":{"Written by":"Elizabeth Rodrigues","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#article","isPartOf":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/"},"author":{"name":"Elizabeth Rodrigues","@id":"https:\/\/www.ulatus.com\/translation-blog\/#\/schema\/person\/2c88b235c1a21525603db0395518c787"},"headline":"What AI needs from human now to build a completed automated world?","datePublished":"2026-09-10T06:02:19+00:00","dateModified":"2026-10-07T06:06:53+00:00","mainEntityOfPage":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/"},"wordCount":2423,"commentCount":0,"image":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#primaryimage"},"thumbnailUrl":"https:\/\/www.ulatus.com\/translation-blog\/wp-content\/uploads\/2026\/10\/banner-what-ai-needs-from-humans-automated-world.png","articleSection":["Data for AI"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/","url":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/","name":"What AI needs from human now to build a completed automated world?","isPartOf":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#primaryimage"},"image":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#primaryimage"},"thumbnailUrl":"https:\/\/www.ulatus.com\/translation-blog\/wp-content\/uploads\/2026\/10\/banner-what-ai-needs-from-humans-automated-world.png","datePublished":"2026-09-10T06:02:19+00:00","dateModified":"2026-10-07T06:06:53+00:00","author":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/#\/schema\/person\/2c88b235c1a21525603db0395518c787"},"breadcrumb":{"@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#primaryimage","url":"https:\/\/www.ulatus.com\/translation-blog\/wp-content\/uploads\/2026\/10\/banner-what-ai-needs-from-humans-automated-world.png","contentUrl":"https:\/\/www.ulatus.com\/translation-blog\/wp-content\/uploads\/2026\/10\/banner-what-ai-needs-from-humans-automated-world.png","width":1280,"height":720,"caption":"AI needs from human now to build a completed automated world"},{"@type":"BreadcrumbList","@id":"https:\/\/www.ulatus.com\/translation-blog\/what-ai-needs-from-human-now-to-build-a-completed-automated-world\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.ulatus.com\/translation-blog\/"},{"@type":"ListItem","position":2,"name":"What AI needs from human now to build a completed automated world?"}]},{"@type":"WebSite","@id":"https:\/\/www.ulatus.com\/translation-blog\/#website","url":"https:\/\/www.ulatus.com\/translation-blog\/","name":"","description":"Translation Trends &amp; Insights","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.ulatus.com\/translation-blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/www.ulatus.com\/translation-blog\/#\/schema\/person\/2c88b235c1a21525603db0395518c787","name":"Elizabeth Rodrigues","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/8cbc606363a4bbd8eda6ccc8c6a750eae05a2de996ee4dc49e0c776598bb1d45?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/8cbc606363a4bbd8eda6ccc8c6a750eae05a2de996ee4dc49e0c776598bb1d45?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/8cbc606363a4bbd8eda6ccc8c6a750eae05a2de996ee4dc49e0c776598bb1d45?s=96&d=mm&r=g","caption":"Elizabeth Rodrigues"},"url":"https:\/\/www.ulatus.com\/translation-blog\/author\/elizabeth-rodrigues\/"}]}},"_links":{"self":[{"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/posts\/3178","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/users\/25"}],"replies":[{"embeddable":true,"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/comments?post=3178"}],"version-history":[{"count":1,"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/posts\/3178\/revisions"}],"predecessor-version":[{"id":3180,"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/posts\/3178\/revisions\/3180"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/media\/3179"}],"wp:attachment":[{"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/media?parent=3178"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/categories?post=3178"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.ulatus.com\/translation-blog\/wp-json\/wp\/v2\/tags?post=3178"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}