Roughly half of all websites are written in English, a language spoken by under a fifth of the people on Earth. That mismatch hints at a bigger problem. In 2025, Gartner reported that 63% of organizations lack the right data management practices for AI, or aren’t sure they have them. Models get better every quarter, but the data for AI that only people can supply is still the bottleneck.
What Does AI Need From Humans to Build a Fully Automated World?
AI needs five kinds of human input: original human-created data, human labels and judgment, human preference feedback, human evaluation and red-teaming, and multilingual, culturally grounded data. Automation scales up whatever people teach it. In our view, a model cut off from those five inputs tends to stall, drift, or fail the moment it leaves the languages and situations it has already seen.
Why each of the five still needs people
Original behavior, writing and speech are the raw material, and no model has produced them from scratch. Labels make that material learnable, preference rankings show which answer people actually favor, and evaluation catches failures before real work gets handed over. Multilingual and cultural data is the most under-served of the five, and it’s also the one that decides whether a system works for most of the world.
Edwin Chen, founder of Surge AI, told Forbes in 2025: “without us, AGI just won’t happen.” He sells human data, so he has a commercial interest. Still, the research below backs up his basic point.
Why “data for AI” is not the same as “more data”
Volume drove the last decade. Quality, provenance and coverage will drive this one. Gartner predicted in February 2025 that through 2026, organizations will abandon 60% of AI projects that lack AI-ready data, which tells you where executives expect trouble. If you need help sourcing this kind of material, Ulatus offers AI training data services built around human contributors.
What Is Data for AI, and Which Parts Can Machines Make Themselves?
In one sentence, data for AI is any information a model learns from or is measured against, whether it’s raw, labeled, ranked or tested, and its most valuable parts still start with humans. Machines can generate synthetic data to add to what already exists, like paraphrases, simulated scenarios or extra examples. What they can’t yet do reliably is come up with new real-world behavior, supply ground-truth judgment, or confirm what’s culturally appropriate. People remain the source of truth, and machines multiply it.
Raw data vs training data vs evaluation data
Raw data is whatever you collect before anyone shapes it: logs, documents, recordings, images. Training data is the portion a model learns from. Validation data gets held back to tune settings, and test data stays locked away until the end so you get an honest score. Wikipedia’s page on training, validation and test data sets covers the distinctions well.
Why does the split matter? Because a model graded on data it has already seen looks brilliant, then performs badly in the real world.
What synthetic data can and cannot replace
Synthetic data is genuinely useful. It fills gaps, balances out rare cases, and can be cheaper than collecting everything by hand. The trouble starts when it replaces real data instead of supplementing it.
Model collapse happens when models train again and again on other models’ output until the rare parts of the original data disappear. A July 2024 Nature paper by Shumailov and colleagues demonstrated this across several kinds of models, including large language models, and described the defects as irreversible.
Follow-up work adds some nuance. Gerstgrasser and colleagues showed in 2024 that collapse depends on whether synthetic data replaces real data or piles up alongside it: when real data stays in the mix, error stays bounded. And a 2025 arXiv preprint, “A Probabilistic Perspective on Model Collapse,” finds that collapse is avoided only if the training sample keeps growing at each step. Both results argue for keeping humans in the loop, since people are the main source of fresh, real data.
Why Is Human-Generated Data Running Short?
Public human-written text is finite. In a 2024 paper, Epoch AI, a research institute, projected that models will train on datasets roughly equal to the entire stock of public human text sometime between 2026 and 2032, if scaling trends continue.
The public-text ceiling
Epoch’s 2025 report, “Can AI Scaling Continue Through 2030?”, estimates that the indexed web holds about 500 trillion words of unique text, and that it will grow by about 50% by 2030.
After adjusting for quality, repeated passes and tokenizer efficiency, Epoch estimates that 400 trillion to 20 quadrillion token-equivalents will be available for training by 2030 (tokens are, roughly, chunks of words). That range is enormous. Nobody knows exactly when the well runs dry, but it’s not bottomless, and Epoch also notes that synthetic data may be needed to ease bottlenecks, which brings the collapse risk right back.
Why scraping is giving way to consented, licensed data
Scarcity isn’t the only pressure on scraped data, because the law is moving too. Under Article 53 of the EU AI Act, providers of general-purpose AI models must publish a sufficiently detailed public summary of their training content and keep a copyright-compliance policy. According to Bird & Bird, the Commission released the mandatory template on 24 July 2025, and the obligations applied from 2 August 2025.
Some high-risk deadlines were later pushed back by the AI Omnibus, formally adopted in mid-2026, but the general-purpose model rules stayed as they were. Data with clear origins is easier to document, which is why consent-backed AI data collection is increasingly expected rather than optional.
How Do Humans Label, Rank and Test What AI Learns?
Humans do three jobs that models still can’t fully do for themselves, even with model-assisted labeling. They annotate data so it becomes ground truth, they rank outputs so models learn what people prefer, and they test systems for errors and risks before deployment. All three are human-in-the-loop work, meaning a person reviews, corrects or approves what a machine produces.
The market is growing, too. The Business Research Company estimates the AI annotation market at USD 1.91 billion in 2025, rising to USD 2.51 billion in 2026.
Annotation: turning raw data into ground truth
Annotation is where a person marks what a recording, image or sentence contains. Picture a call-center transcript. Someone tags the customer’s intent, the sentiment, and any account numbers that must be masked. Train a model on thousands of examples like that, and it learns to do much of the tagging with limited supervision. That’s why human data annotation remains the foundation of supervised learning.
So how do you know the labels are any good? Ask for an agreement score. Cohen’s kappa measures how often two annotators agree beyond chance. A long-standing rule of thumb (Landis and Koch, 1977) treats 0.81 and above as almost perfect agreement, while McHugh (2012) treats anything below 0.60 as inadequate. Those bands give buyers an actual number to ask for.
RLHF: teaching models what people prefer
Reinforcement learning from human feedback (RLHF) trains a model using people’s rankings of its answers. Those rankings train a reward model, which then steers the main model. Imagine a chatbot gives three replies to a billing complaint. A reviewer picks the one that’s accurate, polite and right for the customer’s culture, and the model learns to favor that pattern.
Expertise counts here. A tax answer is best judged by a tax professional, and a question about Japanese honorifics by a fluent speaker. Generic crowd work can miss exactly those distinctions.
Evaluation and red-teaming: checking before automating
Before a system takes over a workflow, someone has to check it. Evaluators score outputs for accuracy and tone, while red-teamers try to break the model with tricky or harmful prompts. Doing this in more than one language matters, because multilingual model evaluation by native speakers catches problems that English-only testing often misses. And the more work you hand to automation, the more that judgment is worth.
Where Is the Biggest Gap in Data for AI?
Multilingual data is a strong candidate for the biggest gap, because the web that models learn from doesn’t look like the people who use them. English dominates websites far beyond its share of speakers, and models perform noticeably worse in lower-resource languages.
English dominates the data, not the planet
W3Techs reports that English is the content language of 49.5% of websites with a known content language, as of October 2026, followed by Spanish at 6.0% and German at 5.9%. That’s a share of websites, not of training corpora. Most frontier labs don’t disclose their language mix, but where it was disclosed for older models, English made up about 93% of GPT-3’s training data (OpenAI, 2020) and about 90% of Llama 2’s (Meta, 2023).
Ethnologue’s 2025 figures put English at about 1.5 billion total speakers, against a world population of about 8.2 billion. So you get half the websites for under a fifth of the people.
Ethnologue also counts about 7,170 living languages in 2025, and many of them have little presence online.
What breaks when AI meets dialects and low-resource languages
Stanford HAI warned in May 2025 that approximately 5 billion people who don’t speak English are poorly served by today’s AI tools. Large language models work well for English’s roughly 1.5 billion speakers, it noted, but they underperform for Vietnamese (97 million speakers) and do far worse for Nahuatl (1.5 million). Researcher Sanmi Koyejo defines low-resource languages as those “with limited amounts of computer-readable data about them.”
Benchmarks show what that means in practice. IrokoBench, published at NAACL 2025, tested models on 17 African languages using human-translated tasks. At the time of testing, the best open model, Gemma 2 27B, reached only 63% of the score of the best proprietary model, GPT-4o. So for a team supporting customers in, say, Yoruba or Swahili, that means starting below two-thirds of the top system’s performance. Translating the test sets into English first helped close the gap for larger English-centric models, which suggests the weakness lies more in native-language data than in the task itself.
The safety picture is messier. A May 2026 arXiv preprint from Stellenbosch University tested multi-turn jailbreak conversations in Afrikaans, Kiswahili, isiXhosa and isiZulu. Harmful response rates reached 60% to 78% in Afrikaans and up to 71% in Kiswahili, against an English baseline of 53% to 84%. The authors found that translation quality, not just how much data a language has, drove jailbreak success. That’s another argument for native speakers who can judge translation quality.
Code-switching, culture and context: what native speakers add
Real speakers mix languages mid-sentence, use regional slang, and shift register depending on who’s listening. A model trained mostly on English web text sees comparatively little of that, and Stanford’s team cautions that models risk flattening cultural diversity into a largely US-centric view.
Native speakers fix this because they know what’s polite, what’s risky and what simply sounds wrong. In our view, translators and linguists are an underused supply of AI data, since they work across languages every day. For speech systems, native-speaker transcription for speech models is a practical first step: accurate transcripts of real accents and dialects.
What Does Good Human Data for AI Look Like?
Good human data passes a five-point test, the 5 Cs: consent, coverage, consistency, context and compliance. It’s our own checklist, not a regulatory standard, but each C maps to a check buyers can actually verify.
A five-point quality test: consent, coverage, consistency, context, compliance
- Contributors agreed to how their data is used, and you can show the record.
- The data spans the languages, dialects and situations your users represent.
- Labels follow written guidelines, and an agreed kappa target shows annotators apply them the same way.
- Annotators understand the domain and culture, not just the words.
- Provenance, licensing and copyright policy are documented, which is what the EU training-data summary template pushes providers toward.
Questions to ask any data partner before you sign
Before you sign, ask four things:
- How are contributors recruited and vetted, and is consent documented?
- Which agreement metric do you report, at what threshold and on what sample size?
- Can we see sample QA reports and error logs, not a slide deck?
- How do you handle dialects and match reviewers to domains?
Vendor numbers deserve some caution. Ulatus says its network includes more than 100,000 vetted contributors across 200-plus languages and 125-plus countries, and it lists inter-annotator agreement reporting as a service. Those figures are self-reported, so ask any vendor, including us, to back them up with audit artifacts.
What Should Teams Do Now to Prepare Their Data for AI?
Start small and measurable: audit one workflow, source the missing data, label it with agreement checks, and evaluate before scaling. A narrow pilot exposes data gaps cheaply, and it gives you real error rates instead of vendor promises.
A four-step plan: audit, source, label, evaluate
- List the data your chosen workflow depends on. Mark what’s original, what’s synthetic, what has clear consent and which languages it covers.
- Fill gaps with consented, documented data from real contributors, prioritizing your customers’ languages.
- Annotate with written guidelines, trained reviewers and a kappa target agreed in advance.
- Test with held-out data and native-speaker reviewers, and track error rates before widening the rollout.
Where to start if you operate in more than one language
Pick a single language pair or locale. Then measure model performance directly in that language instead of assuming English results carry over, a caution IrokoBench’s findings support. Stanford HAI also recommends community-driven data collection with fair ownership for low-resource languages. Once agreement and error rates look healthy, add the next locale.
For content-heavy workflows, human-corrected translations can double as training data, which makes machine translation with human post-editing a sensible bridge. The machine handles the volume, and linguists catch the nuance.
Conclusion: Automation Runs on Human-Supplied Data
Automation needs original human data, careful labels, ranked preferences, honest evaluation and, above all, data in the languages people actually speak. Skip any of them and you end up with a model that looks brilliant on its own test data, then fails in the language nobody tested.
So this quarter, audit one workflow and write down its data gaps: what’s missing, what’s synthetic, and which languages go untested. If you want a partner for consent-backed multilingual training data, annotation and evaluation, bring Ulatus one workflow and one language, and ask for a sample batch with an agreement report. It’s the quickest way to find out whether your data for AI is ready.
