Major AI labs (OpenAI, Anthropic, Google, Meta, etc.) spend hundreds of millions to billions of dollars annually on human data work.
Companies like Scale AI, Surge AI (the company behind DataAnnotation), Outlier, and others employ or contract tens to hundreds of thousands of people worldwide.
Even “synthetic data” approaches (AI generating training data for other AI) still require human oversight, filtering, and validation to avoid model collapse or amplifying errors.
Bottom line: Current AI is better understood as a powerful pattern-matching engine that is steered, corrected, and quality-controlled by large amounts of human judgment. The “AI” part generates fluent output at scale; the human part supplies the signal of what “good” actually looks like. Remove that human layer and the systems rapidly become less reliable, less useful, and more prone to failure.